{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"bf59bc94","cell_type":"markdown","source":"# RSNA Knee Abnormality Detection — 2D CNN Study-Level Baseline\n\n**Notebook:** `RSNA_Knee_2D_CNN_StudyLevel_Baseline`\n\n## Competition objective\n\nDetect twelve clinically important knee abnormalities from multimodal knee MRI studies\n(MRI images + original radiology reports), predicting a **confidence score per target,\nper study**. The leaderboard metric is the **macro-averaged ROC AUC across the 12 targets**.\n\n## What this notebook is (and is not)\n\nThis is a **first-stage baseline notebook**. Its purpose is to:\n\n1. Inspect the actual competition data (not just the written spec) and report what is really there.\n2. Build a correct, leak-free, **study-level** validation split.\n3. Establish a simple, reproducible **image-only** baseline using a small 2D CNN trained from scratch.\n4. Report per-target and macro ROC AUC honestly, including its limitations.\n\nIt intentionally does **not** attempt multimodal fusion, pretrained backbones, 3D modeling,\nensembling, or pseudo-labeling. Those are listed as future work at the end of the notebook.\n\n## Story of this notebook\n\n```\nProblem → Data → EDA → Validation → Preprocessing → Baseline Model → Evaluation → Submission\n```\n\n## A note on official facts vs. assumptions\n\nWherever this notebook states something about the data, it is either:\n- **Verified directly from the provided CSV files** (preferred, and called out explicitly), or\n- Stated as coming from the official competition Context document, or\n- Explicitly flagged as an assumption / design decision, together with the reasoning.\n\nNo DICOM pixel data was available to the notebook author ahead of time (569.76 GB dataset,\nper the competition Context document) — all DICOM-related code below is written to be run\ndirectly in the Kaggle environment against the real files, and is defensive about the fact\nthat it could not be test-executed against real DICOMs beforehand.\n","metadata":{}},{"id":"9a3f7d46","cell_type":"markdown","source":"## 1. Imports","metadata":{}},{"id":"72450832","cell_type":"code","source":"# Standard library\nimport os\nimport gc\nimport time\nimport json\nimport random\nimport warnings\nfrom pathlib import Path\nfrom collections import defaultdict\n\n# Data handling\nimport numpy as np\nimport pandas as pd\n\n# Visualization\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Modeling / evaluation\nfrom sklearn.model_selection import KFold\nfrom sklearn.metrics import roc_auc_score\n\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torch.utils.data import Dataset, DataLoader\n\n# DICOM handling (part of the standard Kaggle Python image, no internet required)\ntry:\n    import pydicom\n    from pydicom.pixel_data_handlers.util import apply_voi_lut\n    PYDICOM_AVAILABLE = True\nexcept ImportError:\n    PYDICOM_AVAILABLE = False\n\nwarnings.filterwarnings(\"ignore\")\nsns.set_style(\"whitegrid\")\n\nprint(\"PyTorch version:\", torch.__version__)\nprint(\"CUDA available:\", torch.cuda.is_available())\nprint(\"pydicom available:\", PYDICOM_AVAILABLE)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:46.896697Z","iopub.execute_input":"2026-08-08T08:06:46.897037Z","iopub.status.idle":"2026-08-08T08:06:46.905491Z","shell.execute_reply.started":"2026-08-08T08:06:46.897008Z","shell.execute_reply":"2026-08-08T08:06:46.90457Z"}},"outputs":[],"execution_count":null},{"id":"5ae26df4","cell_type":"markdown","source":"## 2. Configuration","metadata":{}},{"id":"72f82a55","cell_type":"code","source":"class CFG:\n    # ---- Paths -----------------------------------------------------------\n    # Official Kaggle dataset path (per competition instructions).\n    COMP_DIR = \"/kaggle/input/competitions/rsna-knee-abnormality-detection\"\n\n    TRAIN_CSV = f\"{COMP_DIR}/train.csv\"\n    TRAIN_SERIES_CSV = f\"{COMP_DIR}/train_series.csv\"\n    TEST_CSV = f\"{COMP_DIR}/test.csv\"\n    TEST_SERIES_CSV = f\"{COMP_DIR}/test_series.csv\"\n    SAMPLE_SUB_CSV = f\"{COMP_DIR}/sample_submission.csv\"\n\n    TRAIN_DICOM_DIR = f\"{COMP_DIR}/train_series\"\n    TEST_DICOM_DIR = f\"{COMP_DIR}/test_series\"\n\n    # ---- Targets -----------------------------------------------------------\n    TARGET_COLS = [\n        \"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\",\n        \"Medial OA\", \"Lateral OA\", \"PF OA\", \"Effusion\",\n        \"Synovitis\", \"Baker's\", \"Contusion\", \"Fracture\",\n    ]\n    N_TARGETS = len(TARGET_COLS)\n\n    # ---- Reproducibility -----------------------------------------------------------\n    SEED = 42\n    N_FOLDS = 5\n\n    # ---- Image preprocessing -----------------------------------------------------------\n    IMG_SIZE = 224          # resize side length (square) for the representative slice\n    N_INPUT_CHANNELS = 1    # single grayscale slice as input for this baseline\n\n    # ---- Model / training -----------------------------------------------------------\n    MODEL_NAME = \"RSNA_Knee_2D_CNN_StudyLevel_Baseline\"\n    BATCH_SIZE = 16\n    NUM_EPOCHS = 20\n    LEARNING_RATE = 1e-3\n    WEIGHT_DECAY = 1e-4\n    EARLY_STOP_PATIENCE = 5\n    NUM_WORKERS = 2\n\n    DEVICE = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n\n    OUTPUT_SUBMISSION = \"submission.csv\"\n\n\ncfg = CFG()\nprint(f\"Model name: {cfg.MODEL_NAME}\")\nprint(f\"Device: {cfg.DEVICE}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:46.90744Z","iopub.execute_input":"2026-08-08T08:06:46.907757Z","iopub.status.idle":"2026-08-08T08:06:46.928123Z","shell.execute_reply.started":"2026-08-08T08:06:46.907721Z","shell.execute_reply":"2026-08-08T08:06:46.927212Z"}},"outputs":[],"execution_count":null},{"id":"c1fce5fa","cell_type":"markdown","source":"## 3. Seed setup","metadata":{}},{"id":"cbedccea","cell_type":"code","source":"def set_seed(seed: int) -> None:\n    '''Set random seeds across python/numpy/torch for reproducibility.'''\n    random.seed(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed_all(seed)\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = False\n\n\nset_seed(cfg.SEED)\nprint(f\"Seed set to {cfg.SEED}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:46.929498Z","iopub.execute_input":"2026-08-08T08:06:46.92996Z","iopub.status.idle":"2026-08-08T08:06:46.953394Z","shell.execute_reply.started":"2026-08-08T08:06:46.929933Z","shell.execute_reply":"2026-08-08T08:06:46.95238Z"}},"outputs":[],"execution_count":null},{"id":"b8380398","cell_type":"markdown","source":"## 4. Step 1 — Dataset Inspection\n\nWe load and inspect all five provided CSV files directly. Nothing here is assumed from the\nContext document without being checked against the actual files.\n","metadata":{}},{"id":"ac60ed8d","cell_type":"code","source":"train = pd.read_csv(cfg.TRAIN_CSV)\ntrain_series = pd.read_csv(cfg.TRAIN_SERIES_CSV)\ntest = pd.read_csv(cfg.TEST_CSV)\ntest_series = pd.read_csv(cfg.TEST_SERIES_CSV)\nsample_sub = pd.read_csv(cfg.SAMPLE_SUB_CSV)\n\nprint(\"=== Shapes ===\")\nfor name, df in [(\"train\", train), (\"train_series\", train_series),\n                  (\"test\", test), (\"test_series\", test_series),\n                  (\"sample_submission\", sample_sub)]:\n    print(f\"{name:20s}: {df.shape}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:46.954662Z","iopub.execute_input":"2026-08-08T08:06:46.9551Z","iopub.status.idle":"2026-08-08T08:06:47.402184Z","shell.execute_reply.started":"2026-08-08T08:06:46.955074Z","shell.execute_reply":"2026-08-08T08:06:47.401347Z"}},"outputs":[],"execution_count":null},{"id":"aed364fc","cell_type":"code","source":"print(\"=== train.csv dtypes ===\")\nprint(train.dtypes)\nprint()\nprint(\"=== train.csv columns ===\")\nprint(list(train.columns))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:47.404658Z","iopub.execute_input":"2026-08-08T08:06:47.404957Z","iopub.status.idle":"2026-08-08T08:06:47.411745Z","shell.execute_reply.started":"2026-08-08T08:06:47.404931Z","shell.execute_reply":"2026-08-08T08:06:47.410889Z"}},"outputs":[],"execution_count":null},{"id":"5c7cc31d","cell_type":"code","source":"# ---------------------------------------------------------------------------\n# IMPORTANT DATA DISCREPANCY CHECK\n#\n# The Context document (Section 8.1) states that train.csv contains a\n# `PatientSex` column (\"Male\" or \"Female\"; may be blank). We check this\n# directly against the actual file rather than assuming it is present.\n# ---------------------------------------------------------------------------\nhas_patient_sex = \"PatientSex\" in train.columns\nprint(f\"'PatientSex' column present in the actual train.csv: {has_patient_sex}\")\n\nif not has_patient_sex:\n    print(\n        \"\\nDISCREPANCY: The Context document describes a `PatientSex` column, \"\n        \"but the actual train.csv provided to this notebook does NOT contain it.\\n\"\n        \"This baseline therefore does NOT use PatientSex as a feature. \"\n        \"If a later export of train.csv includes this column, the EDA/preprocessing \"\n        \"code should be revisited.\"\n    )\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:47.413085Z","iopub.execute_input":"2026-08-08T08:06:47.41346Z","iopub.status.idle":"2026-08-08T08:06:47.438381Z","shell.execute_reply.started":"2026-08-08T08:06:47.413433Z","shell.execute_reply":"2026-08-08T08:06:47.437281Z"}},"outputs":[],"execution_count":null},{"id":"1fdd627b","cell_type":"code","source":"print(\"=== Missing values per column (train.csv) ===\")\nprint(train.isnull().sum())\nprint()\nprint(\"=== Missing values per column (train_series.csv) ===\")\nprint(train_series.isnull().sum())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:47.439473Z","iopub.execute_input":"2026-08-08T08:06:47.439788Z","iopub.status.idle":"2026-08-08T08:06:47.463788Z","shell.execute_reply.started":"2026-08-08T08:06:47.439763Z","shell.execute_reply":"2026-08-08T08:06:47.46281Z"}},"outputs":[],"execution_count":null},{"id":"dfe99e40","cell_type":"code","source":"print(\"=== Duplicate rows ===\")\nprint(\"train.csv duplicate rows:\", train.duplicated().sum())\nprint(\"train_series.csv duplicate rows:\", train_series.duplicated().sum())\nprint(\"train_series.csv duplicate (Study, Series) pairs:\",\n      train_series.duplicated(subset=[\"StudyInstanceUID\", \"SeriesInstanceUID\"]).sum())\n\nprint()\nprint(\"=== Unique identifiers ===\")\nprint(\"Unique StudyInstanceUID in train.csv:\", train[\"StudyInstanceUID\"].nunique())\nprint(\"Duplicate StudyInstanceUID in train.csv:\", train[\"StudyInstanceUID\"].duplicated().sum())\nprint(\"Unique StudyInstanceUID in train_series.csv:\", train_series[\"StudyInstanceUID\"].nunique())\nprint(\"Unique SeriesInstanceUID in train_series.csv:\", train_series[\"SeriesInstanceUID\"].nunique())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:47.464977Z","iopub.execute_input":"2026-08-08T08:06:47.465369Z","iopub.status.idle":"2026-08-08T08:06:47.568677Z","shell.execute_reply.started":"2026-08-08T08:06:47.465333Z","shell.execute_reply":"2026-08-08T08:06:47.56779Z"}},"outputs":[],"execution_count":null},{"id":"043b037c","cell_type":"code","source":"# ---------------------------------------------------------------------------\n# Study <-> Series relationship consistency between train.csv and train_series.csv\n# ---------------------------------------------------------------------------\ntrain_studies = set(train[\"StudyInstanceUID\"])\nseries_studies = set(train_series[\"StudyInstanceUID\"])\n\nprint(\"Studies in train.csv but missing from train_series.csv:\", len(train_studies - series_studies))\nprint(\"Studies in train_series.csv but missing from train.csv:\", len(series_studies - train_studies))\n\nseries_per_study = train_series.groupby(\"StudyInstanceUID\").size()\nprint()\nprint(\"=== Series per study — summary statistics ===\")\nprint(series_per_study.describe())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:47.57001Z","iopub.execute_input":"2026-08-08T08:06:47.570384Z","iopub.status.idle":"2026-08-08T08:06:47.595775Z","shell.execute_reply.started":"2026-08-08T08:06:47.570347Z","shell.execute_reply":"2026-08-08T08:06:47.594798Z"}},"outputs":[],"execution_count":null},{"id":"65340a84","cell_type":"code","source":"fig, ax = plt.subplots(figsize=(7, 4))\nsns.histplot(series_per_study, bins=range(series_per_study.min(), series_per_study.max() + 2),\n             ax=ax, color=\"steelblue\")\nax.set_title(\"Number of MRI series per study (train)\")\nax.set_xlabel(\"Series per study\")\nax.set_ylabel(\"Number of studies\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:47.596874Z","iopub.execute_input":"2026-08-08T08:06:47.597165Z","iopub.status.idle":"2026-08-08T08:06:47.876558Z","shell.execute_reply.started":"2026-08-08T08:06:47.597133Z","shell.execute_reply":"2026-08-08T08:06:47.875797Z"}},"outputs":[],"execution_count":null},{"id":"4410cda3","cell_type":"code","source":"print(\"=== test.csv ===\")\nprint(test)\nprint()\nprint(\"=== test_series.csv ===\")\nprint(test_series)\nprint()\nprint(\n    \"Note: the Context document states the example test.csv contains only \"\n    \"3 public study IDs, and that the real (hidden) test set used for scoring \"\n    \"contains approximately 1,300 studies. All inference code below is written \"\n    \"to operate on `test.csv` generically (by iterating over its rows), so it \"\n    \"will work unchanged against the real hidden test set.\"\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:47.877725Z","iopub.execute_input":"2026-08-08T08:06:47.877979Z","iopub.status.idle":"2026-08-08T08:06:47.893355Z","shell.execute_reply.started":"2026-08-08T08:06:47.877956Z","shell.execute_reply":"2026-08-08T08:06:47.892316Z"}},"outputs":[],"execution_count":null},{"id":"e995f434","cell_type":"code","source":"print(\"=== Train vs test schema differences ===\")\nprint(\"train.csv columns:      \", list(train.columns))\nprint(\"test.csv columns:       \", list(test.columns))\nprint(\"train_series columns:   \", list(train_series.columns))\nprint(\"test_series columns:    \", list(test_series.columns))\nprint(\"sample_submission cols: \", list(sample_sub.columns))\n\nassert list(train_series.columns) == list(test_series.columns), \\\n    \"train_series and test_series schemas differ unexpectedly.\"\nassert list(sample_sub.columns) == [\"StudyInstanceUID\"] + cfg.TARGET_COLS, \\\n    \"sample_submission column order differs from CFG.TARGET_COLS — check target ordering.\"\nprint(\"\\nSchema checks passed: test_series matches train_series schema, \"\n      \"and sample_submission columns match the configured target order.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:47.894886Z","iopub.execute_input":"2026-08-08T08:06:47.895348Z","iopub.status.idle":"2026-08-08T08:06:47.906734Z","shell.execute_reply.started":"2026-08-08T08:06:47.895311Z","shell.execute_reply":"2026-08-08T08:06:47.905712Z"}},"outputs":[],"execution_count":null},{"id":"75d21c33","cell_type":"markdown","source":"## 5. Step 3 — Label Analysis\n\nWe analyze the 12 target labels directly from `train.csv`. A study is only counted as\n**labeled** if it has a non-null value for the targets — we explicitly do **not** treat\nmissing/unlabeled targets as negatives.\n","metadata":{}},{"id":"9ee4786c","cell_type":"code","source":"labeled_mask = train[cfg.TARGET_COLS].notnull().any(axis=1)\nn_labeled = labeled_mask.sum()\nn_unlabeled = (~labeled_mask).sum()\n\nprint(f\"Studies with at least one non-null target label: {n_labeled} / {len(train)} \"\n      f\"({n_labeled / len(train):.2%})\")\nprint(f\"Studies with ALL 12 targets missing: {n_unlabeled} / {len(train)} \"\n      f\"({n_unlabeled / len(train):.2%})\")\n\n# Check whether labeling is \"all-or-nothing\" (all 12 present) or partial per study.\nall_present = train.loc[labeled_mask, cfg.TARGET_COLS].notnull().all(axis=1)\nprint(f\"\\nAmong labeled studies, rows with ALL 12 targets present: \"\n      f\"{all_present.sum()} / {n_labeled}\")\nprint(f\"Among labeled studies, rows with SOME but not all targets present: \"\n      f\"{(~all_present).sum()} / {n_labeled}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:47.908466Z","iopub.execute_input":"2026-08-08T08:06:47.908785Z","iopub.status.idle":"2026-08-08T08:06:47.93096Z","shell.execute_reply.started":"2026-08-08T08:06:47.908757Z","shell.execute_reply":"2026-08-08T08:06:47.929823Z"}},"outputs":[],"execution_count":null},{"id":"f535c4dc","cell_type":"code","source":"if all_present.all():\n    print(\n        \"FINDING: Labeling in this dataset export is all-or-nothing at the study level — \"\n        \"a study either has all 12 targets populated, or all 12 are missing. There is no \"\n        \"partial per-target labeling observed in the actual data.\\n\"\n        \"This is consistent with, but more specific than, the Context document's statement \"\n        \"that 'only a small subset of training studies has per-condition labels.'\"\n    )\nelse:\n    print(\n        \"FINDING: Labeling is NOT strictly all-or-nothing — some studies have a mix of \"\n        \"populated and missing targets. Per-target missingness must be handled independently \"\n        \"downstream (e.g. per-target loss masking).\"\n    )\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:47.935314Z","iopub.execute_input":"2026-08-08T08:06:47.935648Z","iopub.status.idle":"2026-08-08T08:06:47.95288Z","shell.execute_reply.started":"2026-08-08T08:06:47.935592Z","shell.execute_reply":"2026-08-08T08:06:47.95206Z"}},"outputs":[],"execution_count":null},{"id":"06a48d3a","cell_type":"code","source":"labeled_df = train.loc[labeled_mask].reset_index(drop=True)\n\nrows = []\nfor t in cfg.TARGET_COLS:\n    col = train[t]\n    pos = (col == 1).sum()\n    neg = (col == 0).sum()\n    miss = col.isnull().sum()\n    total_known = pos + neg\n    pos_rate = pos / total_known if total_known > 0 else np.nan\n    rows.append({\"target\": t, \"positive\": pos, \"negative\": neg,\n                 \"missing\": miss, \"positive_rate_among_labeled\": pos_rate})\n\nlabel_stats = pd.DataFrame(rows).set_index(\"target\")\nprint(label_stats)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:47.954039Z","iopub.execute_input":"2026-08-08T08:06:47.954338Z","iopub.status.idle":"2026-08-08T08:06:47.980748Z","shell.execute_reply.started":"2026-08-08T08:06:47.954313Z","shell.execute_reply":"2026-08-08T08:06:47.979731Z"}},"outputs":[],"execution_count":null},{"id":"7951c882","cell_type":"code","source":"fig, axes = plt.subplots(1, 2, figsize=(14, 5))\n\nlabel_stats[[\"positive\", \"negative\"]].plot(kind=\"bar\", stacked=True, ax=axes[0],\n                                            color=[\"indianred\", \"steelblue\"])\naxes[0].set_title(f\"Positive vs. negative counts per target (n={n_labeled} labeled studies)\")\naxes[0].set_ylabel(\"Count\")\naxes[0].tick_params(axis=\"x\", rotation=60)\n\nlabel_stats[\"positive_rate_among_labeled\"].plot(kind=\"bar\", ax=axes[1], color=\"darkorange\")\naxes[1].set_title(\"Positive rate among labeled studies\")\naxes[1].set_ylabel(\"Positive rate\")\naxes[1].set_ylim(0, 1)\naxes[1].tick_params(axis=\"x\", rotation=60)\n\nplt.tight_layout()\nplt.show()\n\nprint(\n    \"\\nOBSERVATION: The labeled subset is very small \"\n    f\"({n_labeled} studies) relative to the full training set ({len(train)} studies), \"\n    \"and every target is imbalanced (minority positive rate roughly 15-45%). \"\n    \"This has direct consequences for validation design (Section 7 below) and for how \"\n    \"much confidence we can place in the baseline's AUC estimates.\"\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:47.981961Z","iopub.execute_input":"2026-08-08T08:06:47.982472Z","iopub.status.idle":"2026-08-08T08:06:48.645954Z","shell.execute_reply.started":"2026-08-08T08:06:47.982436Z","shell.execute_reply":"2026-08-08T08:06:48.645138Z"}},"outputs":[],"execution_count":null},{"id":"7a51c8b4","cell_type":"markdown","source":"## 6. Step 4 — Series Analysis\n\nWe analyze series-level metadata (`Fluid_Sensitive`, `Fat_Suppression`, `Anatomical_Plane`)\nacross `train_series.csv`, and how it relates to studies.\n","metadata":{}},{"id":"e7c0c274","cell_type":"code","source":"print(\"=== Anatomical_Plane distribution (series-level) ===\")\nprint(train_series[\"Anatomical_Plane\"].value_counts(dropna=False))\nprint()\nprint(\"=== Fluid_Sensitive distribution (series-level) ===\")\nprint(train_series[\"Fluid_Sensitive\"].value_counts(dropna=False))\nprint()\nprint(\"=== Fat_Suppression distribution (series-level) ===\")\nprint(train_series[\"Fat_Suppression\"].value_counts(dropna=False))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:48.647097Z","iopub.execute_input":"2026-08-08T08:06:48.647384Z","iopub.status.idle":"2026-08-08T08:06:48.657582Z","shell.execute_reply.started":"2026-08-08T08:06:48.647359Z","shell.execute_reply":"2026-08-08T08:06:48.65673Z"}},"outputs":[],"execution_count":null},{"id":"b374688c","cell_type":"code","source":"fluid_fs_match = (train_series[\"Fluid_Sensitive\"] == train_series[\"Fat_Suppression\"]).mean()\nprint(f\"Fraction of series where Fluid_Sensitive == Fat_Suppression: {fluid_fs_match:.4f}\")\nif fluid_fs_match == 1.0:\n    print(\n        \"OBSERVATION: In this dataset export, Fluid_Sensitive and Fat_Suppression are \"\n        \"identical for every series. They should be treated as redundant for series \"\n        \"selection purposes (using one is equivalent to using both), though this should be \"\n        \"re-verified if the dataset export changes.\"\n    )\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:48.658846Z","iopub.execute_input":"2026-08-08T08:06:48.660098Z","iopub.status.idle":"2026-08-08T08:06:48.678584Z","shell.execute_reply.started":"2026-08-08T08:06:48.66006Z","shell.execute_reply":"2026-08-08T08:06:48.677671Z"}},"outputs":[],"execution_count":null},{"id":"2a8e987e","cell_type":"code","source":"combo_counts = (\n    train_series.groupby([\"Anatomical_Plane\", \"Fluid_Sensitive\", \"Fat_Suppression\"])\n    .size()\n    .rename(\"n_series\")\n    .reset_index()\n)\nprint(combo_counts)\n\nfig, ax = plt.subplots(figsize=(8, 5))\npivot = train_series.pivot_table(index=\"Anatomical_Plane\", columns=\"Fluid_Sensitive\",\n                                  values=\"SeriesInstanceUID\", aggfunc=\"count\", fill_value=0)\npivot.plot(kind=\"bar\", ax=ax, color=[\"slategray\", \"seagreen\"])\nax.set_title(\"Series count by Anatomical Plane x Fluid_Sensitive\")\nax.set_ylabel(\"Number of series\")\nax.legend(title=\"Fluid_Sensitive\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:48.67992Z","iopub.execute_input":"2026-08-08T08:06:48.680549Z","iopub.status.idle":"2026-08-08T08:06:48.909706Z","shell.execute_reply.started":"2026-08-08T08:06:48.68052Z","shell.execute_reply":"2026-08-08T08:06:48.908857Z"}},"outputs":[],"execution_count":null},{"id":"ac0be197","cell_type":"code","source":"# ---------------------------------------------------------------------------\n# Coverage check for a candidate series-selection rule:\n#   \"Sagittal plane AND Fluid_Sensitive == 1\"\n# This is the combination most EDA-favored for knee MRI protocols in this\n# dataset (see combo_counts above), so we check how often it is actually\n# available per study before committing to it as the primary selection rule.\n# ---------------------------------------------------------------------------\ndef has_combo(g, plane, fluid_sensitive):\n    return ((g[\"Anatomical_Plane\"] == plane) & (g[\"Fluid_Sensitive\"] == fluid_sensitive)).any()\n\nsag_fs_coverage = train_series.groupby(\"StudyInstanceUID\").apply(\n    lambda g: has_combo(g, \"Sagittal\", 1)\n)\nsag_only_coverage = train_series.groupby(\"StudyInstanceUID\").apply(\n    lambda g: (g[\"Anatomical_Plane\"] == \"Sagittal\").any()\n)\n\nprint(f\"Studies with >=1 Sagittal + Fluid_Sensitive series: \"\n      f\"{sag_fs_coverage.sum()} / {len(sag_fs_coverage)} ({sag_fs_coverage.mean():.2%})\")\nprint(f\"Studies with >=1 Sagittal series (any fluid sensitivity): \"\n      f\"{sag_only_coverage.sum()} / {len(sag_only_coverage)} ({sag_only_coverage.mean():.2%})\")\n\n# Coverage restricted to the 58 labeled studies used for training the baseline.\nlabeled_study_ids = labeled_df[\"StudyInstanceUID\"]\nlab_series = train_series[train_series[\"StudyInstanceUID\"].isin(labeled_study_ids)]\nlab_sag_fs = lab_series.groupby(\"StudyInstanceUID\").apply(lambda g: has_combo(g, \"Sagittal\", 1))\nprint(f\"\\nAmong the {len(labeled_study_ids)} LABELED studies: \"\n      f\"{lab_sag_fs.sum()} / {len(lab_sag_fs)} have >=1 Sagittal + Fluid_Sensitive series.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:48.91075Z","iopub.execute_input":"2026-08-08T08:06:48.911021Z","iopub.status.idle":"2026-08-08T08:06:50.673516Z","shell.execute_reply.started":"2026-08-08T08:06:48.910996Z","shell.execute_reply":"2026-08-08T08:06:50.672669Z"}},"outputs":[],"execution_count":null},{"id":"dbdca8e9","cell_type":"markdown","source":"### Decision: series-selection rule for the image baseline\n\n**Decision.** For each study, select one representative series using this priority order:\n\n1. A `Sagittal` series with `Fluid_Sensitive == 1` (pick the one with the most slices if\n   several qualify, as a proxy for a more complete acquisition).\n2. Else, any `Sagittal` series (most slices).\n3. Else, any series at all (most slices) — this should not occur in `train_series.csv`\n   given every study has >= 1 Sagittal series (checked above), but the fallback keeps the\n   pipeline robust for `test_series.csv` / the hidden test set.\n\n**Why.** Sagittal fluid-sensitive (e.g. PD/T2/STIR-type) sequences are the most common single\nacquisition type in this dataset (see `combo_counts` above) and are the sequence radiologists\ntypically rely on most for assessing ligaments, menisci and effusion — the majority of our 12\ntargets. The EDA above shows this combination is available for the large majority of studies\n(~94% overall, ~97% of the labeled subset), so it is a reasonable single-series choice for a\n*first* baseline.\n\n**Alternatives considered.** (a) Use every series per study with 3D/multi-view fusion — more\ninformation but much higher engineering and runtime cost, deferred to a later iteration.\n(b) Use `Fat_Suppression` instead of/in addition to `Fluid_Sensitive` — shown above to be\nredundant with `Fluid_Sensitive` in this export, so it adds no extra signal for series\nselection.\n\n**Trade-offs.** Discarding the other ~4-5 series per study throws away information (e.g.\nAxial/Coronal views, which matter for some targets such as `Lateral Meniscus` or `PF OA`).\nThis is an explicit simplification for a first baseline, not a claim that it is optimal.\n\n**Effect on runtime.** Using a single 2D slice per study instead of full volumes keeps\npreprocessing and training extremely cheap (a few hundred small grayscale images total),\nwhich matters given the 9-hour CPU/GPU runtime limit and the Efficiency Prize.\n\n**What happens when the preferred series is missing.** Handled by the fallback chain above;\nthis is logged per-study so it is visible how often the fallback triggers (see the DICOM\nloading section).\n","metadata":{}},{"id":"d5cc82e8","cell_type":"markdown","source":"## 7. Step 5 — Report Analysis\n\nWe analyze the free-text `Report` column. Reports are the original radiology reports and,\nper the Context document, may be multilingual and may contain sensitive information — so we\nonly print short, truncated snippets for illustration.\n","metadata":{}},{"id":"ca4b4792","cell_type":"code","source":"report_missing = train[\"Report\"].isnull().sum()\nreport_empty = (train[\"Report\"].fillna(\"\").str.strip() == \"\").sum()\nprint(f\"Missing (NaN) reports: {report_missing} / {len(train)}\")\nprint(f\"Empty-string reports: {report_empty} / {len(train)}\")\n\nreport_len_chars = train[\"Report\"].fillna(\"\").str.len()\nreport_len_words = train[\"Report\"].fillna(\"\").str.split().str.len()\n\nprint(\"\\n=== Report length in characters — summary ===\")\nprint(report_len_chars.describe())\nprint(\"\\n=== Report length in words — summary ===\")\nprint(report_len_words.describe())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:50.674632Z","iopub.execute_input":"2026-08-08T08:06:50.67497Z","iopub.status.idle":"2026-08-08T08:06:50.783213Z","shell.execute_reply.started":"2026-08-08T08:06:50.674934Z","shell.execute_reply":"2026-08-08T08:06:50.782377Z"}},"outputs":[],"execution_count":null},{"id":"43fad0bf","cell_type":"code","source":"fig, axes = plt.subplots(1, 2, figsize=(13, 4))\nsns.histplot(report_len_chars, bins=40, ax=axes[0], color=\"teal\")\naxes[0].set_title(\"Report length (characters)\")\nsns.histplot(report_len_words, bins=40, ax=axes[1], color=\"purple\")\naxes[1].set_title(\"Report length (words)\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:50.784316Z","iopub.execute_input":"2026-08-08T08:06:50.78458Z","iopub.status.idle":"2026-08-08T08:06:51.228802Z","shell.execute_reply.started":"2026-08-08T08:06:50.784557Z","shell.execute_reply":"2026-08-08T08:06:51.227946Z"}},"outputs":[],"execution_count":null},{"id":"3753e9f9","cell_type":"code","source":"# ---------------------------------------------------------------------------\n# Lightweight, heuristic language detection.\n#\n# No internet access is available in the scoring environment, so we cannot rely\n# on downloading a language-ID model. This is a *rough* keyword-based heuristic\n# over a handful of languages that were visually observed in a small sample of\n# reports during EDA — it is NOT a validated language identifier and should be\n# treated as descriptive only.\n# ---------------------------------------------------------------------------\nLANGUAGE_HINTS = {\n    \"english\": [\"the\", \"and\", \"with\", \"findings\", \"impression\"],\n    \"spanish\": [\"de la\", \"el \", \"hallazgos\", \"rodilla\", \"impresi\"],\n    \"german\": [\"und\", \"der\", \"die\", \"kniegelenk\", \"befund\"],\n    \"dutch\": [\"van\", \"het\", \"knie\", \"bevindingen\"],\n    \"french\": [\"les \", \"des \", \"articulaire\", \"constatations\"],\n    \"turkish\": [\"diz\", \"bulgular\", \"eklem\"],\n    \"greek\": [\"και\", \"εξέταση\", \"ευρήματα\"],\n}\n\ndef guess_language(text: str) -> str:\n    text_low = str(text).lower()\n    scores = {lang: sum(text_low.count(h) for h in hints) for lang, hints in LANGUAGE_HINTS.items()}\n    best_lang = max(scores, key=scores.get)\n    return best_lang if scores[best_lang] > 0 else \"unknown/other\"\n\nsample_for_lang = train[\"Report\"].dropna().sample(min(300, len(train)), random_state=cfg.SEED)\nlang_counts = sample_for_lang.apply(guess_language).value_counts()\nprint(\"Rough language-hint distribution over a random sample of 300 reports:\")\nprint(lang_counts)\nprint(\n    \"\\nThis confirms the Context document's statement that reports may be multilingual — \"\n    \"at least English, Spanish, German, Dutch, French, Turkish and Greek hints were observed. \"\n    \"The first image baseline (Step 7) does not use report text at all; a text/multimodal \"\n    \"model is left as future work (see final section).\"\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:51.230411Z","iopub.execute_input":"2026-08-08T08:06:51.230774Z","iopub.status.idle":"2026-08-08T08:06:51.263657Z","shell.execute_reply.started":"2026-08-08T08:06:51.23073Z","shell.execute_reply":"2026-08-08T08:06:51.262637Z"}},"outputs":[],"execution_count":null},{"id":"37c3a510","cell_type":"code","source":"print(\"=== Example reports (truncated to 120 characters, first 3 only) ===\")\nfor i, r in enumerate(train[\"Report\"].dropna().head(3)):\n    print(f\"[{i}] {r[:120]!r} ...\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:51.264918Z","iopub.execute_input":"2026-08-08T08:06:51.265242Z","iopub.status.idle":"2026-08-08T08:06:51.273526Z","shell.execute_reply.started":"2026-08-08T08:06:51.265207Z","shell.execute_reply":"2026-08-08T08:06:51.272702Z"}},"outputs":[],"execution_count":null},{"id":"bd8f0b9c","cell_type":"markdown","source":"## 8. Step 6 — DICOM Inspection\n\nThe full DICOM dataset (~569.76 GB, ~819,640 files per the Context document) is far too\nlarge to load into this notebook's memory. This section inspects a **small representative\nsample** of actual DICOM files directly from the Kaggle dataset path, without loading full\nvolumes or the full dataset.\n\nThis code needs the real files to run and could not be executed ahead of time — it is\nwritten defensively (existence checks, try/except per file) so a single unreadable file\ndoes not crash the whole notebook.\n","metadata":{}},{"id":"c49445d7","cell_type":"code","source":"def list_series_dir(study_uid: str, series_uid: str, dicom_root: str) -> list:\n    '''Return sorted list of .dcm file paths for a given study/series.'''\n    series_dir = Path(dicom_root) / study_uid / series_uid\n    if not series_dir.exists():\n        return []\n    return sorted(series_dir.glob(\"*.dcm\"))\n\n\ndef inspect_dicom_file(path) -> dict:\n    '''Read a single DICOM file's header (and pixel array) and report key properties.'''\n    info = {\"path\": str(path), \"readable\": False}\n    if not PYDICOM_AVAILABLE:\n        info[\"error\"] = \"pydicom not available in this environment\"\n        return info\n    try:\n        ds = pydicom.dcmread(str(path))\n        info[\"readable\"] = True\n        info[\"rows\"] = getattr(ds, \"Rows\", None)\n        info[\"columns\"] = getattr(ds, \"Columns\", None)\n        info[\"bits_allocated\"] = getattr(ds, \"BitsAllocated\", None)\n        info[\"photometric_interpretation\"] = getattr(ds, \"PhotometricInterpretation\", None)\n        info[\"pixel_spacing\"] = getattr(ds, \"PixelSpacing\", None)\n        info[\"slice_thickness\"] = getattr(ds, \"SliceThickness\", None)\n        info[\"transfer_syntax\"] = str(getattr(ds.file_meta, \"TransferSyntaxUID\", None)) \\\n            if hasattr(ds, \"file_meta\") else None\n        info[\"image_orientation_patient\"] = getattr(ds, \"ImageOrientationPatient\", None)\n        try:\n            arr = ds.pixel_array\n            info[\"pixel_array_shape\"] = arr.shape\n            info[\"intensity_min\"] = float(arr.min())\n            info[\"intensity_max\"] = float(arr.max())\n            info[\"intensity_mean\"] = float(arr.mean())\n        except Exception as e:\n            info[\"pixel_read_error\"] = str(e)\n    except Exception as e:\n        info[\"error\"] = str(e)\n    return info\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:51.274549Z","iopub.execute_input":"2026-08-08T08:06:51.275463Z","iopub.status.idle":"2026-08-08T08:06:51.297942Z","shell.execute_reply.started":"2026-08-08T08:06:51.275433Z","shell.execute_reply":"2026-08-08T08:06:51.29666Z"}},"outputs":[],"execution_count":null},{"id":"ce347a80","cell_type":"code","source":"# Sample a handful of studies/series to inspect (defensive: this cell will simply\n# report \"0 studies found\" if the DICOM directory is not mounted, e.g. when this\n# notebook is run outside the Kaggle competition environment).\nn_studies_to_inspect = 3\nn_series_per_study = 2\n\nsample_study_ids = train_series[\"StudyInstanceUID\"].drop_duplicates().head(n_studies_to_inspect).tolist()\nprint(f\"Inspecting DICOMs for {len(sample_study_ids)} sample studies \"\n      f\"under: {cfg.TRAIN_DICOM_DIR}\")\n\ndicom_inspection_rows = []\nfor study_uid in sample_study_ids:\n    study_series = train_series[train_series[\"StudyInstanceUID\"] == study_uid]\n    for series_uid in study_series[\"SeriesInstanceUID\"].head(n_series_per_study):\n        files = list_series_dir(study_uid, series_uid, cfg.TRAIN_DICOM_DIR)\n        n_slices = len(files)\n        row = {\"study\": study_uid, \"series\": series_uid, \"n_slices_found\": n_slices}\n        if n_slices > 0:\n            mid_file = files[n_slices // 2]\n            row.update(inspect_dicom_file(mid_file))\n        dicom_inspection_rows.append(row)\n\ndicom_inspection_df = pd.DataFrame(dicom_inspection_rows)\nprint(dicom_inspection_df)\n\nif dicom_inspection_df[\"n_slices_found\"].sum() == 0:\n    print(\n        \"\\nNOTE: No DICOM files were found under the expected path. This is expected if this \"\n        \"notebook is being reviewed outside the Kaggle competition environment (the DICOM \"\n        \"files are not attached). When run inside the actual Kaggle Notebook editor with the \"\n        \"competition data attached, this cell will populate real header/pixel statistics.\"\n    )\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:51.299078Z","iopub.execute_input":"2026-08-08T08:06:51.299351Z","iopub.status.idle":"2026-08-08T08:06:51.49427Z","shell.execute_reply.started":"2026-08-08T08:06:51.299328Z","shell.execute_reply":"2026-08-08T08:06:51.493447Z"}},"outputs":[],"execution_count":null},{"id":"ff589b94","cell_type":"code","source":"# Visualize representative slices, if any were found above.\nfound_rows = [r for r in dicom_inspection_rows if r.get(\"readable\")]\nif found_rows:\n    n_show = min(4, len(found_rows))\n    fig, axes = plt.subplots(1, n_show, figsize=(4 * n_show, 4))\n    if n_show == 1:\n        axes = [axes]\n    for ax, row in zip(axes, found_rows[:n_show]):\n        path = row[\"path\"]\n        ds = pydicom.dcmread(path)\n        arr = ds.pixel_array.astype(np.float32)\n        ax.imshow(arr, cmap=\"gray\")\n        ax.set_title(f\"{row['series'][-6:]}\\n{arr.shape}\", fontsize=8)\n        ax.axis(\"off\")\n    plt.tight_layout()\n    plt.show()\nelse:\n    print(\"No readable DICOM files available to visualize in this environment.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:51.495347Z","iopub.execute_input":"2026-08-08T08:06:51.495721Z","iopub.status.idle":"2026-08-08T08:06:52.295876Z","shell.execute_reply.started":"2026-08-08T08:06:51.495686Z","shell.execute_reply":"2026-08-08T08:06:52.294927Z"}},"outputs":[],"execution_count":null},{"id":"2916bae1","cell_type":"code","source":"# Slice count distribution across the same sample of series (cheap: only counts\n# files per directory, does not read pixel data). This is used to sanity-check the\n# Context document's claim of a typical 20-45 slice range (median ~30).\nslice_count_rows = []\nsample_series_ids = train_series[[\"StudyInstanceUID\", \"SeriesInstanceUID\"]].drop_duplicates().head(200)\nfor _, r in sample_series_ids.iterrows():\n    files = list_series_dir(r[\"StudyInstanceUID\"], r[\"SeriesInstanceUID\"], cfg.TRAIN_DICOM_DIR)\n    slice_count_rows.append(len(files))\n\nslice_counts = pd.Series(slice_count_rows)\nif slice_counts.sum() > 0:\n    print(slice_counts[slice_counts > 0].describe())\nelse:\n    print(\n        \"No DICOM directories found in this environment — this cell will report real slice \"\n        \"counts once run inside Kaggle with the competition data attached.\"\n    )\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:52.296966Z","iopub.execute_input":"2026-08-08T08:06:52.297215Z","iopub.status.idle":"2026-08-08T08:06:54.287193Z","shell.execute_reply.started":"2026-08-08T08:06:52.297185Z","shell.execute_reply":"2026-08-08T08:06:54.286349Z"}},"outputs":[],"execution_count":null},{"id":"fe899261","cell_type":"markdown","source":"## 9. Step 2 — Study-Level Validation Strategy\n\n**Decision.** Split at the `StudyInstanceUID` level, never at the series/slice level, so\nno slice or series from the same study can appear in both train and validation. Since one\nrow of `train.csv` already corresponds to exactly one study (verified above: no duplicate\n`StudyInstanceUID`), a per-row split is automatically a per-study split — but we implement\nit with `sklearn`'s `KFold` over the **labeled study IDs explicitly** (rather than a single\n80/20 holdout) for the following reason:\n\n**Why K-fold instead of a single static holdout.** EDA above shows only **58 studies** have\nany label at all. A single, e.g., 80/20 holdout would validate on ~12 studies — extremely\nnoisy for a 12-target macro AUC (a single flipped prediction can move a fold's AUC by a lot).\nUsing 5-fold cross-validation and collecting **out-of-fold (OOF) predictions** lets every one\nof the 58 labeled studies be scored exactly once as validation data, while never being seen\nin training for that fold — this uses the scarce labels more efficiently without introducing\nleakage, and gives one macro AUC computed over all 58 OOF predictions instead of five very\nsmall, high-variance numbers.\n\n**Trade-off.** This is still a small-sample estimate (n=58) with real uncertainty — the\nmacro AUC reported below should be read as indicative, not a precise leaderboard estimate.\nFuture iterations should consider repeated CV, semi-supervised use of the 4,349 unlabeled\nstudies, or waiting for more labels to appear in later dataset updates.\n\n**Effect on runtime.** Training 5 small folds on ~46-47 studies each is inexpensive (seconds\nto low minutes even on CPU), so this does not threaten the 9-hour runtime budget.\n","metadata":{}},{"id":"d1837d86","cell_type":"code","source":"kf = KFold(n_splits=cfg.N_FOLDS, shuffle=True, random_state=cfg.SEED)\n\nlabeled_df = labeled_df.reset_index(drop=True)\nlabeled_df[\"fold\"] = -1\nfor fold, (_, val_idx) in enumerate(kf.split(labeled_df)):\n    labeled_df.loc[val_idx, \"fold\"] = fold\n\nprint(labeled_df[\"fold\"].value_counts().sort_index())\nassert (labeled_df[\"fold\"] == -1).sum() == 0, \"Every labeled study must be assigned to a fold.\"\n\n# Sanity check: no StudyInstanceUID appears in more than one fold (guaranteed here since\n# each study is a single row, but asserted explicitly to document the leakage-prevention\n# guarantee requested by the task instructions).\nassert labeled_df[\"StudyInstanceUID\"].duplicated().sum() == 0\nprint(\"\\nLeakage check passed: every StudyInstanceUID is assigned to exactly one fold.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:54.288256Z","iopub.execute_input":"2026-08-08T08:06:54.28867Z","iopub.status.idle":"2026-08-08T08:06:54.302097Z","shell.execute_reply.started":"2026-08-08T08:06:54.288626Z","shell.execute_reply":"2026-08-08T08:06:54.301211Z"}},"outputs":[],"execution_count":null},{"id":"3d02c919","cell_type":"markdown","source":"## 10. Preprocessing pipeline: loading the representative slice","metadata":{}},{"id":"e5dcb9db","cell_type":"code","source":"def select_representative_series(study_uid: str, series_df: pd.DataFrame) -> dict:\n    '''\n    Select one representative series for a study using the EDA-derived priority rule:\n      1. Sagittal + Fluid_Sensitive == 1 (most slices if multiple)\n      2. Any Sagittal series (most slices)\n      3. Any series at all (most slices)\n    Returns a dict with the chosen SeriesInstanceUID and which rule matched, or None if\n    the study has no series at all.\n    '''\n    study_series = series_df[series_df[\"StudyInstanceUID\"] == study_uid]\n    if len(study_series) == 0:\n        return None\n\n    def pick_most_slices(candidates, dicom_root):\n        best_row, best_n = None, -1\n        for _, row in candidates.iterrows():\n            n = len(list_series_dir(study_uid, row[\"SeriesInstanceUID\"], dicom_root))\n            if n > best_n:\n                best_row, best_n = row, n\n        return best_row, best_n\n\n    dicom_root = cfg.TRAIN_DICOM_DIR if study_uid in set(train[\"StudyInstanceUID\"]) else cfg.TEST_DICOM_DIR\n\n    rule1 = study_series[(study_series[\"Anatomical_Plane\"] == \"Sagittal\") &\n                          (study_series[\"Fluid_Sensitive\"] == 1)]\n    if len(rule1) > 0:\n        row, n_slices = pick_most_slices(rule1, dicom_root)\n        return {\"series_uid\": row[\"SeriesInstanceUID\"], \"rule\": \"sagittal_fluid_sensitive\",\n                \"n_slices\": n_slices}\n\n    rule2 = study_series[study_series[\"Anatomical_Plane\"] == \"Sagittal\"]\n    if len(rule2) > 0:\n        row, n_slices = pick_most_slices(rule2, dicom_root)\n        return {\"series_uid\": row[\"SeriesInstanceUID\"], \"rule\": \"sagittal_any\", \"n_slices\": n_slices}\n\n    row, n_slices = pick_most_slices(study_series, dicom_root)\n    return {\"series_uid\": row[\"SeriesInstanceUID\"], \"rule\": \"any_series\", \"n_slices\": n_slices}\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:54.30312Z","iopub.execute_input":"2026-08-08T08:06:54.303406Z","iopub.status.idle":"2026-08-08T08:06:54.325158Z","shell.execute_reply.started":"2026-08-08T08:06:54.303381Z","shell.execute_reply":"2026-08-08T08:06:54.324191Z"}},"outputs":[],"execution_count":null},{"id":"e8cbbf81","cell_type":"code","source":"def load_middle_slice(study_uid: str, series_uid: str, dicom_root: str, img_size: int) -> np.ndarray:\n    '''\n    Load, normalize and resize the middle slice of a given series.\n    Returns a (img_size, img_size) float32 array in [0, 1], or a zero array of the same\n    shape if the slice could not be read (e.g. unsupported transfer syntax, missing file) —\n    this keeps the training/inference loop robust to individual corrupt/unsupported DICOMs,\n    which the Context document explicitly warns can occur (mixed transfer syntaxes: JPEG\n    Lossless, JPEG 2000, Explicit/Implicit VR Little Endian).\n    '''\n    files = list_series_dir(study_uid, series_uid, dicom_root)\n    if len(files) == 0:\n        return np.zeros((img_size, img_size), dtype=np.float32), False\n\n    mid_path = files[len(files) // 2]\n    try:\n        ds = pydicom.dcmread(str(mid_path))\n        arr = ds.pixel_array.astype(np.float32)\n\n        # Apply rescale slope/intercept if present (standard DICOM linear transform).\n        slope = float(getattr(ds, \"RescaleSlope\", 1.0))\n        intercept = float(getattr(ds, \"RescaleIntercept\", 0.0))\n        arr = arr * slope + intercept\n\n        # Min-max normalize per-slice to [0, 1] — robust to the varying intensity\n        # characteristics across scanners/sequences noted in the Context document.\n        arr_min, arr_max = arr.min(), arr.max()\n        if arr_max > arr_min:\n            arr = (arr - arr_min) / (arr_max - arr_min)\n        else:\n            arr = np.zeros_like(arr)\n\n        # Resize to a fixed square size using simple PyTorch interpolation (avoids adding\n        # an extra image-processing dependency that might be unavailable offline).\n        t = torch.from_numpy(arr).unsqueeze(0).unsqueeze(0)\n        t = F.interpolate(t, size=(img_size, img_size), mode=\"bilinear\", align_corners=False)\n        arr_resized = t.squeeze(0).squeeze(0).numpy()\n\n        return arr_resized.astype(np.float32), True\n    except Exception as e:\n        return np.zeros((img_size, img_size), dtype=np.float32), False\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:54.326541Z","iopub.execute_input":"2026-08-08T08:06:54.327056Z","iopub.status.idle":"2026-08-08T08:06:54.349249Z","shell.execute_reply.started":"2026-08-08T08:06:54.327019Z","shell.execute_reply":"2026-08-08T08:06:54.348357Z"}},"outputs":[],"execution_count":null},{"id":"876609df","cell_type":"code","source":"def build_study_index(study_ids, series_df, dicom_root) -> pd.DataFrame:\n    '''Pre-compute the selected series (and selection rule) for a list of studies.'''\n    rows = []\n    for study_uid in study_ids:\n        sel = select_representative_series(study_uid, series_df)\n        if sel is None:\n            rows.append({\"StudyInstanceUID\": study_uid, \"series_uid\": None,\n                         \"rule\": \"no_series_found\", \"n_slices\": 0})\n        else:\n            rows.append({\"StudyInstanceUID\": study_uid, **sel})\n    return pd.DataFrame(rows)\n\n\ntrain_study_index = build_study_index(labeled_df[\"StudyInstanceUID\"], train_series, cfg.TRAIN_DICOM_DIR)\nprint(train_study_index[\"rule\"].value_counts())\ntrain_study_index.head()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:54.350396Z","iopub.execute_input":"2026-08-08T08:06:54.350989Z","iopub.status.idle":"2026-08-08T08:06:55.198324Z","shell.execute_reply.started":"2026-08-08T08:06:54.350957Z","shell.execute_reply":"2026-08-08T08:06:55.197459Z"}},"outputs":[],"execution_count":null},{"id":"0700f634","cell_type":"markdown","source":"## 11. Dataset / DataLoader","metadata":{}},{"id":"c7c91512","cell_type":"code","source":"class KneeStudyDataset(Dataset):\n    '''\n    One sample = one study. Loads the middle slice of the study's representative series\n    on-the-fly (not pre-loaded into RAM), and returns the 12 target labels when available.\n    '''\n\n    def __init__(self, study_index: pd.DataFrame, dicom_root: str, img_size: int,\n                 labels_df: pd.DataFrame = None, target_cols=None):\n        self.study_index = study_index.reset_index(drop=True)\n        self.dicom_root = dicom_root\n        self.img_size = img_size\n        self.target_cols = target_cols or cfg.TARGET_COLS\n\n        if labels_df is not None:\n            self.labels = (\n                labels_df.set_index(\"StudyInstanceUID\")[self.target_cols]\n                .reindex(self.study_index[\"StudyInstanceUID\"])\n                .fillna(0.0)  # only used at training time, where every row here IS labeled\n                .values.astype(np.float32)\n            )\n        else:\n            self.labels = None\n\n    def __len__(self):\n        return len(self.study_index)\n\n    def __getitem__(self, idx):\n        row = self.study_index.iloc[idx]\n        if row[\"series_uid\"] is None:\n            img = np.zeros((self.img_size, self.img_size), dtype=np.float32)\n        else:\n            img, _ = load_middle_slice(row[\"StudyInstanceUID\"], row[\"series_uid\"],\n                                        self.dicom_root, self.img_size)\n        img_t = torch.from_numpy(img).unsqueeze(0)  # (1, H, W)\n\n        if self.labels is not None:\n            label_t = torch.from_numpy(self.labels[idx])\n            return img_t, label_t\n        return img_t, row[\"StudyInstanceUID\"]\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:55.199543Z","iopub.execute_input":"2026-08-08T08:06:55.200334Z","iopub.status.idle":"2026-08-08T08:06:55.209504Z","shell.execute_reply.started":"2026-08-08T08:06:55.200306Z","shell.execute_reply":"2026-08-08T08:06:55.208398Z"}},"outputs":[],"execution_count":null},{"id":"58f1fd77","cell_type":"markdown","source":"## 12. Model architecture — small 2D CNN trained from scratch\n\n**Decision.** Use a small convolutional network trained **from random initialization**,\nrather than a pretrained ImageNet backbone (e.g. ResNet/EfficientNet).\n\n**Why.** The competition requires `Internet = Off` for scoring, and no pretrained-weights\ndataset is guaranteed to be attached to this notebook. Rather than assume a specific weights\ndataset will be pre-attached (which the instructions explicitly say not to assume), the\nbaseline is designed to run correctly with **zero external dependencies** beyond the\ncompetition data itself. This keeps \"Step 7: Simple Image Baseline\" reliably executable.\n\n**Alternative considered.** Attach a Kaggle \"Models\" dataset with offline ImageNet weights\nand fine-tune a torchvision backbone — likely to perform much better, and is recommended as\nthe natural next iteration once a specific weights dataset is chosen and attached.\n\n**Trade-off.** Training from scratch on ~46-47 labeled studies per fold is a very small\nsample for a CNN; we compensate with a small parameter count, dropout and simple\naugmentation (random flip), and set expectations accordingly (see the evaluation section).\n","metadata":{}},{"id":"8d4d25f8","cell_type":"code","source":"class SimpleKneeCNN(nn.Module):\n    '''Small 2D CNN: a few conv blocks + global average pooling + a linear head\n    producing one logit per target (multi-label, sigmoid applied via the loss function).'''\n\n    def __init__(self, n_targets: int, in_channels: int = 1):\n        super().__init__()\n        self.features = nn.Sequential(\n            nn.Conv2d(in_channels, 16, kernel_size=3, padding=1), nn.BatchNorm2d(16), nn.ReLU(),\n            nn.MaxPool2d(2),  # 224 -> 112\n\n            nn.Conv2d(16, 32, kernel_size=3, padding=1), nn.BatchNorm2d(32), nn.ReLU(),\n            nn.MaxPool2d(2),  # 112 -> 56\n\n            nn.Conv2d(32, 64, kernel_size=3, padding=1), nn.BatchNorm2d(64), nn.ReLU(),\n            nn.MaxPool2d(2),  # 56 -> 28\n\n            nn.Conv2d(64, 128, kernel_size=3, padding=1), nn.BatchNorm2d(128), nn.ReLU(),\n            nn.AdaptiveAvgPool2d(1),  # -> (128, 1, 1)\n        )\n        self.dropout = nn.Dropout(0.4)\n        self.head = nn.Linear(128, n_targets)\n\n    def forward(self, x):\n        x = self.features(x)\n        x = x.flatten(1)\n        x = self.dropout(x)\n        return self.head(x)  # raw logits\n\n\nn_params = sum(p.numel() for p in SimpleKneeCNN(cfg.N_TARGETS).parameters())\nprint(f\"Model: {cfg.MODEL_NAME}\")\nprint(f\"Total trainable parameters: {n_params:,}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:55.211486Z","iopub.execute_input":"2026-08-08T08:06:55.211982Z","iopub.status.idle":"2026-08-08T08:06:55.285097Z","shell.execute_reply.started":"2026-08-08T08:06:55.211954Z","shell.execute_reply":"2026-08-08T08:06:55.284168Z"}},"outputs":[],"execution_count":null},{"id":"e8284ac7","cell_type":"markdown","source":"## 13. Training and evaluation functions","metadata":{}},{"id":"3ac5657d","cell_type":"code","source":"def train_one_fold(fold: int, train_idx_df: pd.DataFrame, val_idx_df: pd.DataFrame) -> dict:\n    '''Train on one fold's training studies, return the model and OOF logits for its\n    validation studies.'''\n    train_study_ids = train_idx_df[\"StudyInstanceUID\"].tolist()\n    val_study_ids = val_idx_df[\"StudyInstanceUID\"].tolist()\n\n    train_index = train_study_index[train_study_index[\"StudyInstanceUID\"].isin(train_study_ids)]\n    val_index = train_study_index[train_study_index[\"StudyInstanceUID\"].isin(val_study_ids)]\n\n    train_ds = KneeStudyDataset(train_index, cfg.TRAIN_DICOM_DIR, cfg.IMG_SIZE,\n                                 labels_df=labeled_df, target_cols=cfg.TARGET_COLS)\n    val_ds = KneeStudyDataset(val_index, cfg.TRAIN_DICOM_DIR, cfg.IMG_SIZE,\n                               labels_df=labeled_df, target_cols=cfg.TARGET_COLS)\n\n    train_loader = DataLoader(train_ds, batch_size=cfg.BATCH_SIZE, shuffle=True,\n                               num_workers=cfg.NUM_WORKERS, drop_last=False)\n    val_loader = DataLoader(val_ds, batch_size=cfg.BATCH_SIZE, shuffle=False,\n                             num_workers=cfg.NUM_WORKERS)\n\n    model = SimpleKneeCNN(cfg.N_TARGETS).to(cfg.DEVICE)\n    optimizer = torch.optim.AdamW(model.parameters(), lr=cfg.LEARNING_RATE,\n                                   weight_decay=cfg.WEIGHT_DECAY)\n    criterion = nn.BCEWithLogitsLoss()\n\n    best_val_loss = float(\"inf\")\n    best_state = None\n    epochs_no_improve = 0\n\n    for epoch in range(cfg.NUM_EPOCHS):\n        model.train()\n        train_loss = 0.0\n        for imgs, labels in train_loader:\n            imgs, labels = imgs.to(cfg.DEVICE), labels.to(cfg.DEVICE)\n            # cheap augmentation: random horizontal flip\n            if random.random() < 0.5:\n                imgs = torch.flip(imgs, dims=[3])\n\n            optimizer.zero_grad()\n            logits = model(imgs)\n            loss = criterion(logits, labels)\n            loss.backward()\n            optimizer.step()\n            train_loss += loss.item() * imgs.size(0)\n        train_loss /= len(train_ds)\n\n        model.eval()\n        val_loss = 0.0\n        with torch.no_grad():\n            for imgs, labels in val_loader:\n                imgs, labels = imgs.to(cfg.DEVICE), labels.to(cfg.DEVICE)\n                logits = model(imgs)\n                loss = criterion(logits, labels)\n                val_loss += loss.item() * imgs.size(0)\n        val_loss /= max(len(val_ds), 1)\n\n        if val_loss < best_val_loss:\n            best_val_loss = val_loss\n            best_state = {k: v.cpu().clone() for k, v in model.state_dict().items()}\n            epochs_no_improve = 0\n        else:\n            epochs_no_improve += 1\n\n        if epoch % 5 == 0 or epoch == cfg.NUM_EPOCHS - 1:\n            print(f\"  Fold {fold} | epoch {epoch:2d} | train_loss={train_loss:.4f} \"\n                  f\"| val_loss={val_loss:.4f}\")\n\n        if epochs_no_improve >= cfg.EARLY_STOP_PATIENCE:\n            print(f\"  Fold {fold}: early stopping at epoch {epoch}\")\n            break\n\n    model.load_state_dict(best_state)\n\n    # Produce OOF logits for this fold's validation studies.\n    model.eval()\n    oof_logits, oof_study_ids = [], []\n    with torch.no_grad():\n        for imgs, study_ids in DataLoader(\n            KneeStudyDataset(val_index, cfg.TRAIN_DICOM_DIR, cfg.IMG_SIZE, labels_df=None),\n            batch_size=cfg.BATCH_SIZE, shuffle=False,\n        ):\n            imgs = imgs.to(cfg.DEVICE)\n            logits = model(imgs).cpu().numpy()\n            oof_logits.append(logits)\n            oof_study_ids.extend(study_ids)\n\n    oof_logits = np.concatenate(oof_logits, axis=0) if oof_logits else np.zeros((0, cfg.N_TARGETS))\n    return {\"model\": model, \"oof_logits\": oof_logits, \"oof_study_ids\": oof_study_ids,\n            \"best_val_loss\": best_val_loss}\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:55.286221Z","iopub.execute_input":"2026-08-08T08:06:55.287115Z","iopub.status.idle":"2026-08-08T08:06:55.3027Z","shell.execute_reply.started":"2026-08-08T08:06:55.287086Z","shell.execute_reply":"2026-08-08T08:06:55.301775Z"}},"outputs":[],"execution_count":null},{"id":"facc3e09","cell_type":"code","source":"def compute_auc_report(y_true: np.ndarray, y_pred: np.ndarray, target_cols: list) -> pd.DataFrame:\n    '''Per-target ROC AUC + macro AUC. Targets with a single class present in y_true are\n    reported as NaN (AUC is undefined in that case) and excluded from the macro average.'''\n    rows = []\n    aucs = []\n    for i, t in enumerate(target_cols):\n        yt, yp = y_true[:, i], y_pred[:, i]\n        if len(np.unique(yt)) < 2:\n            rows.append({\"target\": t, \"auc\": np.nan, \"n_positive\": int(yt.sum()), \"n\": len(yt)})\n            continue\n        auc = roc_auc_score(yt, yp)\n        rows.append({\"target\": t, \"auc\": auc, \"n_positive\": int(yt.sum()), \"n\": len(yt)})\n        aucs.append(auc)\n\n    report = pd.DataFrame(rows).set_index(\"target\")\n    macro_auc = float(np.mean(aucs)) if aucs else float(\"nan\")\n    return report, macro_auc\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:55.303866Z","iopub.execute_input":"2026-08-08T08:06:55.304295Z","iopub.status.idle":"2026-08-08T08:06:55.325596Z","shell.execute_reply.started":"2026-08-08T08:06:55.304259Z","shell.execute_reply":"2026-08-08T08:06:55.324747Z"}},"outputs":[],"execution_count":null},{"id":"98422ef8","cell_type":"markdown","source":"## 14. Run 5-fold cross-validation and compute OOF macro AUC","metadata":{}},{"id":"8ccdca75","cell_type":"code","source":"t_start = time.time()\n\nall_oof_logits = np.zeros((len(labeled_df), cfg.N_TARGETS), dtype=np.float32)\nall_oof_filled = np.zeros(len(labeled_df), dtype=bool)\nfold_models = []\n\nstudy_id_to_row = {uid: i for i, uid in enumerate(labeled_df[\"StudyInstanceUID\"])}\n\nfor fold in range(cfg.N_FOLDS):\n    print(f\"\\n=== Fold {fold} ===\")\n    train_idx_df = labeled_df[labeled_df[\"fold\"] != fold]\n    val_idx_df = labeled_df[labeled_df[\"fold\"] == fold]\n    print(f\"  train studies: {len(train_idx_df)} | val studies: {len(val_idx_df)}\")\n\n    result = train_one_fold(fold, train_idx_df, val_idx_df)\n    fold_models.append(result[\"model\"])\n\n    for sid, logits in zip(result[\"oof_study_ids\"], result[\"oof_logits\"]):\n        row_i = study_id_to_row[sid]\n        all_oof_logits[row_i] = logits\n        all_oof_filled[row_i] = True\n\n    del result\n    gc.collect()\n    if torch.cuda.is_available():\n        torch.cuda.empty_cache()\n\ntrain_time_sec = time.time() - t_start\nprint(f\"\\nTotal cross-validation training time: {train_time_sec:.1f}s \"\n      f\"({train_time_sec / 60:.2f} min)\")\nassert all_oof_filled.all(), \"Every labeled study should receive exactly one OOF prediction.\"\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T08:06:55.326894Z","iopub.execute_input":"2026-08-08T08:06:55.32723Z"}},"outputs":[],"execution_count":null},{"id":"a51b3196","cell_type":"code","source":"oof_probs = 1 / (1 + np.exp(-all_oof_logits))  # sigmoid\ny_true = labeled_df[cfg.TARGET_COLS].values.astype(np.float32)\n\nauc_report, macro_auc = compute_auc_report(y_true, oof_probs, cfg.TARGET_COLS)\nprint(auc_report)\nprint(f\"\\nMacro-averaged ROC AUC (OOF, {len(labeled_df)} labeled studies): {macro_auc:.4f}\")\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"42bc6aa9","cell_type":"code","source":"fig, ax = plt.subplots(figsize=(9, 5))\nauc_report[\"auc\"].plot(kind=\"bar\", ax=ax, color=\"darkslateblue\")\nax.axhline(0.5, color=\"red\", linestyle=\"--\", label=\"random (AUC=0.5)\")\nax.axhline(macro_auc, color=\"green\", linestyle=\"--\", label=f\"macro AUC = {macro_auc:.3f}\")\nax.set_title(f\"{cfg.MODEL_NAME} — per-target OOF ROC AUC (n={len(labeled_df)} labeled studies)\")\nax.set_ylabel(\"ROC AUC\")\nax.set_ylim(0, 1)\nax.legend()\nplt.xticks(rotation=60)\nplt.tight_layout()\nplt.show()\n\nprint(\n    \"\\nHONEST CAVEAT: With only 58 labeled studies split across 5 folds, these AUC values \"\n    \"have high variance and should be interpreted as a sanity-checked, correct baseline \"\n    \"pipeline rather than a precise estimate of leaderboard performance. The main purpose of \"\n    \"this first notebook is a working, leak-free, reproducible pipeline — not a competitive \"\n    \"score.\"\n)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"3365d9c6","cell_type":"markdown","source":"## 15. Runtime and resource usage summary","metadata":{}},{"id":"08521236","cell_type":"code","source":"print(\"=== Runtime summary ===\")\nprint(f\"Cross-validation training time: {train_time_sec:.1f} s ({train_time_sec/60:.2f} min)\")\nprint(f\"Device used: {cfg.DEVICE}\")\nprint(f\"Model parameter count: {n_params:,}\")\nprint(f\"Image size: {cfg.IMG_SIZE}x{cfg.IMG_SIZE}, single channel\")\nprint(f\"Labeled studies used for training/validation: {len(labeled_df)}\")\nprint(\n    \"\\nGiven the very small labeled dataset and single-slice-per-study design, both training \"\n    \"and inference are expected to comfortably fit within the competition's 9-hour CPU/GPU \"\n    \"runtime limit even on CPU-only hardware, leaving substantial runtime budget for future, \"\n    \"more expensive iterations (e.g. multi-series 3D modeling).\"\n)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"5b9b3fae","cell_type":"markdown","source":"## 16. Inference on the test set\n\nWe apply the same series-selection rule and slice-loading pipeline to every study in\n`test.csv`, and average the sigmoid probabilities from the 5 fold models (a simple, cheap\nform of ensembling that also reduces variance from the very small training set).\n","metadata":{}},{"id":"bb82ff5f","cell_type":"code","source":"t_inf_start = time.time()\n\ntest_study_index = build_study_index(test[\"StudyInstanceUID\"], test_series, cfg.TEST_DICOM_DIR)\nprint(test_study_index[\"rule\"].value_counts())\n\ntest_ds = KneeStudyDataset(test_study_index, cfg.TEST_DICOM_DIR, cfg.IMG_SIZE, labels_df=None)\ntest_loader = DataLoader(test_ds, batch_size=cfg.BATCH_SIZE, shuffle=False,\n                          num_workers=cfg.NUM_WORKERS)\n\nall_test_probs = np.zeros((len(test_ds), cfg.N_TARGETS), dtype=np.float32)\ntest_study_ids_ordered = []\n\nfor model in fold_models:\n    model.eval()\n\nfold_prob_accum = np.zeros((len(test_ds), cfg.N_TARGETS), dtype=np.float32)\nwith torch.no_grad():\n    offset = 0\n    for imgs, study_ids in test_loader:\n        imgs = imgs.to(cfg.DEVICE)\n        batch_probs = np.zeros((imgs.size(0), cfg.N_TARGETS), dtype=np.float32)\n        for model in fold_models:\n            logits = model(imgs).cpu().numpy()\n            batch_probs += 1 / (1 + np.exp(-logits))\n        batch_probs /= len(fold_models)\n\n        fold_prob_accum[offset: offset + imgs.size(0)] = batch_probs\n        test_study_ids_ordered.extend(study_ids)\n        offset += imgs.size(0)\n\ninference_time_sec = time.time() - t_inf_start\nprint(f\"\\nInference completed for {len(test_study_ids_ordered)} test studies \"\n      f\"in {inference_time_sec:.2f}s.\")\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"404dbc3d","cell_type":"markdown","source":"## 17. Build and validate `submission.csv`","metadata":{}},{"id":"0fc7dd73","cell_type":"code","source":"pred_df = pd.DataFrame(fold_prob_accum, columns=cfg.TARGET_COLS)\npred_df.insert(0, \"StudyInstanceUID\", test_study_ids_ordered)\n\n# Ensure every StudyInstanceUID from test.csv is present, in test.csv's original order.\nsubmission = test[[\"StudyInstanceUID\"]].merge(pred_df, on=\"StudyInstanceUID\", how=\"left\")\n\n# ---------------------------------------------------------------------------\n# Submission validation checks (per task instructions, Section 14)\n# ---------------------------------------------------------------------------\nassert list(submission.columns) == [\"StudyInstanceUID\"] + cfg.TARGET_COLS, \\\n    f\"Column order mismatch: {list(submission.columns)}\"\nassert list(submission.columns) == list(sample_sub.columns), \\\n    \"Submission columns do not match sample_submission.csv columns/order.\"\nassert len(submission) == len(test), \\\n    f\"Row count mismatch: submission has {len(submission)} rows, test.csv has {len(test)}.\"\nassert submission[\"StudyInstanceUID\"].is_unique, \"Duplicate StudyInstanceUID in submission.\"\nassert set(submission[\"StudyInstanceUID\"]) == set(test[\"StudyInstanceUID\"]), \\\n    \"Submission StudyInstanceUID set does not exactly match test.csv.\"\n\ntarget_values = submission[cfg.TARGET_COLS].values\nassert np.isfinite(target_values).all(), \"Non-finite (NaN/inf) values found in predictions.\"\nassert np.issubdtype(target_values.dtype, np.floating), \"Prediction columns must be numeric.\"\nassert (target_values >= 0).all() and (target_values <= 1).all(), \\\n    \"Prediction values must be valid probabilities in [0, 1].\"\n\nprint(\"All submission validation checks passed.\")\nprint(submission.head())\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"1538cdae","cell_type":"code","source":"submission.to_csv(cfg.OUTPUT_SUBMISSION, index=False)\nprint(f\"Saved {cfg.OUTPUT_SUBMISSION} with shape {submission.shape}\")\n\n# Final structural comparison against sample_submission.csv\nprint(\"\\nColumn-by-column comparison with sample_submission.csv:\")\nprint(pd.DataFrame({\n    \"submission_dtype\": submission.dtypes.astype(str),\n    \"sample_submission_columns\": pd.Series(sample_sub.columns),\n}))\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"4a53d445","cell_type":"markdown","source":"## 18. Summary, limitations and next steps\n\n### What this baseline established\n- Verified the real CSV schemas directly (and found one discrepancy vs. the Context\n  document: no `PatientSex` column in the actual `train.csv`).\n- Found that only **58 of 4,407** training studies (~1.3%) carry any label, and that\n  labeling is all-or-nothing at the study level in this export.\n- Built a leak-free, **study-level** 5-fold cross-validation split, using out-of-fold\n  predictions to make the most of the very small labeled set.\n- Derived an EDA-backed series-selection rule (Sagittal + Fluid-Sensitive, with fallbacks)\n  instead of assuming one upfront.\n- Trained a small, from-scratch 2D CNN baseline on a single representative slice per study,\n  with per-target and macro ROC AUC reported honestly, including its statistical limitations.\n- Produced and validated a schema-correct `submission.csv`.\n\n### Known limitations\n- Only a single 2D slice per study is used; the volumetric and multi-series structure of\n  the data (Study → Series → Slice) is not otherwise exploited.\n- The model is trained from scratch on ~58 labeled studies total — very likely underfit\n  relative to what is achievable with more data or pretrained weights.\n- Report text (multilingual radiology reports) is not used at all in this first baseline.\n- The 4,349 unlabeled training studies are not used.\n- DICOM inspection code (Section 8) could not be executed against real pixel data ahead of\n  time and should be re-checked the first time this notebook is actually run in Kaggle.\n\n### Recommended next iterations (not implemented here, per the \"do not overfit the first\nnotebook\" principle)\n1. Attach an offline pretrained-weights dataset and fine-tune a stronger 2D or 3D backbone.\n2. Use multiple series/slices per study (e.g. 2.5D or 3D CNN) instead of a single slice.\n3. Add a text branch over the (multilingual) `Report` column and experiment with\n   image+text fusion.\n4. Explore semi-supervised or weak-label extraction from reports for the 4,349 unlabeled\n   studies.\n5. Replace single-holdout-style thinking entirely with repeated/stratified CV as more\n   labels become available.\n6. Revisit the series-selection rule (currently a single best series) once more image\n   signal is incorporated.\n7. Add test-time augmentation and proper ensembling once a stronger single model exists.\n","metadata":{}}]}