{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"dea80363","cell_type":"markdown","source":"# 🦵 Abnormality Detection from one knee MRI — EDA\n\n**Study-level structure, report coverage, series geometry, and the two gauges that will matter once a model exists.**\n\nThis EDA is built to answer one question in each section: *what does the imaging model actually get to work with, and where does the supervision run thin?* It follows the same discipline a training pipeline for this data needs downstream — every diagnostic here either informs a modeling decision or rules one out.\n\n| Section | Question it answers |\n|---|---|\n| 1 | What files exist, and what do the twelve targets look like where they're annotated? |\n| 2 | How much of the corpus has an annotation vs. only a report? |\n| 3 | What does the report text look like — language, length, missingness? |\n| 4 | What does `train_series.csv` say about plane and acquisition coverage? |\n| 5 | What do the DICOM headers themselves say — pixel spacing, slice thickness, laterality? |\n| 6 | Do the twelve findings co-occur, or are they close to independent? |\n| 7 | Summary — what this means for splitting, sampling, and slot design |\n\nTwo numbers get carried through the whole notebook rather than computed once and forgotten: **how many studies actually have a ground-truth label** (small — this drives every interval), and **how many have only free text** (large — this is what a report-derived label strategy is for).\n","metadata":{}},{"id":"45fa8e98","cell_type":"markdown","source":"## 1. Setup","metadata":{}},{"id":"80aaa4d0","cell_type":"code","source":"import subprocess, sys\n\ndef pip_install(pkg):\n    subprocess.run([sys.executable, \"-m\", \"pip\", \"install\", \"-q\", pkg])\n\nfor pkg in [\"pydicom\", \"networkx\"]:\n    try:\n        __import__(pkg)\n    except ImportError:\n        pip_install(pkg)\n\nimport os, re, time, gc, hashlib\nfrom pathlib import Path\nfrom collections import Counter, defaultdict\nfrom concurrent.futures import ThreadPoolExecutor\n\nimport numpy as np\nimport pandas as pd\n\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport networkx as nx\n\nimport pydicom\n\nT0 = time.time()\ndef log(msg):\n    print(f\"[{time.time() - T0:7.1f}s] {msg}\", flush=True)\n\nSEED = 2026\nnp.random.seed(SEED)\n\n# ------------------------------------------------------------------ icefire theme\n# Same palette across every chart in this notebook, so a colour always means the same\n# thing: cold end = low / normal / absent, warm end = high / abnormal / present.\nICEFIRE  = sns.color_palette(\"icefire\", 12)\nICE      = ICEFIRE[2]\nFIRE     = ICEFIRE[9]\nMIDLINE  = \"#39414d\"\nCMAP     = \"icefire\"\nCMAP_SEQ = sns.blend_palette([\"#eef1f6\", ICE, FIRE], as_cmap=True)\n\nsns.set_theme(style=\"white\", palette=ICEFIRE, font_scale=0.95)\nmpl.rcParams.update({\n    \"figure.dpi\": 120, \"savefig.dpi\": 120,\n    \"axes.spines.top\": False, \"axes.spines.right\": False,\n    \"axes.edgecolor\": \"#c8cdd4\", \"axes.labelcolor\": MIDLINE,\n    \"axes.titlesize\": 12, \"axes.titleweight\": \"bold\", \"axes.titlepad\": 10,\n    \"text.color\": MIDLINE, \"xtick.color\": MIDLINE, \"ytick.color\": MIDLINE,\n    \"grid.color\": \"#e8eaee\", \"axes.grid\": True, \"axes.grid.axis\": \"y\",\n    \"legend.frameon\": False, \"figure.facecolor\": \"white\",\n})\n\ndef grad(n, reverse=False):\n    \"\"\"n colours along the icefire ramp, skipping its near-black centre.\"\"\"\n    cm = plt.get_cmap(CMAP)\n    cold = np.linspace(0.02, 0.28, n - n // 2)\n    warm = np.linspace(0.78, 0.99, n // 2)\n    cols = [cm(p) for p in np.concatenate([cold, warm])][:n]\n    return cols[::-1] if reverse else cols\n\npd.set_option(\"display.max_columns\", 100)\npd.set_option(\"display.width\", 200)\nlog(\"setup complete\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T05:33:08.122649Z","iopub.execute_input":"2026-08-08T05:33:08.123431Z","iopub.status.idle":"2026-08-08T05:33:08.346557Z","shell.execute_reply.started":"2026-08-08T05:33:08.123402Z","shell.execute_reply":"2026-08-08T05:33:08.345904Z"}},"outputs":[],"execution_count":null},{"id":"aa8de676","cell_type":"markdown","source":"## 2. Load & Inventory\n\nSame resolution logic a training script would use: try the standard mount points, fall back to a scan, and fail loudly rather than guessing a path.\n","metadata":{}},{"id":"b64dc408","cell_type":"code","source":"def find_root():\n    for c in [Path(\"/kaggle/input/competitions/rsna-knee-abnormality-detection\"),\n              Path(\"/kaggle/input/rsna-knee-abnormality-detection\"),\n              Path(\"data\"), Path(\".\")]:\n        if (c / \"test.csv\").is_file():\n            return c\n    for depth1 in sorted(p for p in Path(\"/kaggle/input\").iterdir() if p.is_dir()):\n        for cand in [depth1] + sorted(p for p in depth1.iterdir() if p.is_dir()):\n            if (cand / \"test.csv\").is_file():\n                return cand\n    raise FileNotFoundError(\"competition mount not found — check /kaggle/input\")\n\nROOT = find_root()\nlog(f\"input root: {ROOT}\")\n\ntrain_df = pd.read_csv(ROOT / \"train.csv\")\ntrain_series = pd.read_csv(ROOT / \"train_series.csv\")\ntest_df = pd.read_csv(ROOT / \"test.csv\")\ntest_series = pd.read_csv(ROOT / \"test_series.csv\")\nsample_sub = pd.read_csv(ROOT / \"sample_submission.csv\")\n\nTARGETS = [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\", \"Medial OA\",\n           \"Lateral OA\", \"PF OA\", \"Effusion\", \"Synovitis\", \"Baker's\",\n           \"Contusion\", \"Fracture\"]\nTARGETS = [t for t in TARGETS if t in train_df.columns]  # tolerate a renamed column\n\nprint(f\"train.csv       {train_df.shape}\")\nprint(f\"train_series.csv{train_series.shape}\")\nprint(f\"test.csv        {test_df.shape}\")\nprint(f\"test_series.csv {test_series.shape}\")\nprint(f\"sample_submission {sample_sub.shape}\")\nprint(f\"\\ntargets found in train.csv: {TARGETS}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T05:33:08.347734Z","iopub.execute_input":"2026-08-08T05:33:08.348199Z","iopub.status.idle":"2026-08-08T05:33:08.526711Z","shell.execute_reply.started":"2026-08-08T05:33:08.348179Z","shell.execute_reply":"2026-08-08T05:33:08.526039Z"}},"outputs":[],"execution_count":null},{"id":"32042b2a","cell_type":"code","source":"display(train_df.head(3))\ndisplay(train_series.head(3))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T05:33:08.52759Z","iopub.execute_input":"2026-08-08T05:33:08.527887Z","iopub.status.idle":"2026-08-08T05:33:08.54466Z","shell.execute_reply.started":"2026-08-08T05:33:08.527856Z","shell.execute_reply":"2026-08-08T05:33:08.54401Z"}},"outputs":[],"execution_count":null},{"id":"cc114b6f","cell_type":"markdown","source":"## 3. What the twelve targets look like\n\nTwo populations sit inside `train.csv`: studies with a per-condition **annotation** (the twelve target columns filled in), and studies with only a **report** — the free-text field a radiologist wrote. Everything downstream depends on knowing which is which, because the annotated set is the only one an AUC can be honestly measured against, and it is small.\n","metadata":{}},{"id":"6df4dea3","cell_type":"code","source":"annotated_mask = train_df[TARGETS].notna().all(axis=1)\nn_annot = int(annotated_mask.sum())\nn_total = len(train_df)\n\nprint(f\"studies with a full annotation: {n_annot} / {n_total}  ({n_annot/n_total:.1%})\")\nprint(f\"studies with report only:       {n_total - n_annot} / {n_total}\")\n\nfig, ax = plt.subplots(figsize=(6, 4))\nax.bar([\"Annotated\", \"Report only\"], [n_annot, n_total - n_annot],\n       color=[FIRE, ICE], edgecolor=\"white\", width=.55)\nfor i, v in enumerate([n_annot, n_total - n_annot]):\n    ax.text(i, v + n_total * 0.01, f\"{v}\\n({v/n_total:.1%})\", ha=\"center\", fontsize=10)\nax.set_title(\"Annotated vs. report-only studies\")\nax.set_ylim(0, n_total * 1.15)\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T05:33:08.546337Z","iopub.execute_input":"2026-08-08T05:33:08.546617Z","iopub.status.idle":"2026-08-08T05:33:08.663632Z","shell.execute_reply.started":"2026-08-08T05:33:08.546587Z","shell.execute_reply":"2026-08-08T05:33:08.662921Z"}},"outputs":[],"execution_count":null},{"id":"527f34c9","cell_type":"code","source":"gold = train_df.loc[annotated_mask, TARGETS].astype(float)\n\n# Positive rate per target, annotated subset only — this is the population an AUC will\n# actually be measured against downstream, so it is worth seeing on its own before any\n# report-derived signal enters the picture.\npos_rate = (gold > 0.5).mean().sort_values()\n\nfig, ax = plt.subplots(figsize=(9, 5))\nax.barh(pos_rate.index, pos_rate.values * 100, color=grad(len(pos_rate)), edgecolor=\"white\", height=.7)\nfor i, v in enumerate(pos_rate.values * 100):\n    ax.text(v + 1, i, f\"{v:.0f}%  (n={int((gold[pos_rate.index[i]]>0.5).sum())})\",\n            va=\"center\", fontsize=9)\nax.set_xlim(0, max(pos_rate.values * 100) * 1.35)\nax.set_title(f\"Positive rate on the annotated subset (n={n_annot})\")\nax.set_xlabel(\"% positive\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T05:33:08.664603Z","iopub.execute_input":"2026-08-08T05:33:08.664894Z","iopub.status.idle":"2026-08-08T05:33:08.872673Z","shell.execute_reply.started":"2026-08-08T05:33:08.664864Z","shell.execute_reply":"2026-08-08T05:33:08.87184Z"}},"outputs":[],"execution_count":null},{"id":"073b65ee","cell_type":"markdown","source":"**Why this matters before any model exists:** with only a few dozen positives for the rarest targets, the annotated subset alone cannot arbitrate small differences — the same Hanley–McNeil argument that governs a report-derived label extractor governs any EDA claim made on this slice too. The interval below is drawn for each target's positive rate using the normal approximation to a binomial proportion, so it's visible before anyone is tempted to read too much into a 3-percentage-point gap.\n","metadata":{}},{"id":"98d990f3","cell_type":"code","source":"def wilson_interval(p, n, z=1.96):\n    if n == 0:\n        return (np.nan, np.nan)\n    denom = 1 + z**2 / n\n    centre = p + z**2 / (2 * n)\n    adj = z * np.sqrt((p * (1 - p) + z**2 / (4 * n)) / n)\n    return ((centre - adj) / denom, (centre + adj) / denom)\n\nrows = []\nfor t in TARGETS:\n    p = (gold[t] > 0.5).mean()\n    n = gold[t].notna().sum()\n    lo, hi = wilson_interval(p, n)\n    rows.append({\"target\": t, \"pos_rate\": p, \"n\": int(n), \"lo95\": lo, \"hi95\": hi})\nrate_df = pd.DataFrame(rows).sort_values(\"pos_rate\")\n\nfig, ax = plt.subplots(figsize=(9, 5))\nax.barh(rate_df.target, rate_df.pos_rate * 100, color=grad(len(rate_df)), edgecolor=\"white\", height=.6)\nax.errorbar(rate_df.pos_rate * 100, np.arange(len(rate_df)),\n            xerr=[(rate_df.pos_rate - rate_df.lo95) * 100, (rate_df.hi95 - rate_df.pos_rate) * 100],\n            fmt=\"none\", ecolor=MIDLINE, elinewidth=1.2, capsize=3)\nax.set_title(\"Positive rate with 95% Wilson interval — rare targets have wide bars\")\nax.set_xlabel(\"% positive\")\nplt.tight_layout()\nplt.show()\ndisplay(rate_df.round(3))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T05:33:08.873655Z","iopub.execute_input":"2026-08-08T05:33:08.873899Z","iopub.status.idle":"2026-08-08T05:33:09.065522Z","shell.execute_reply.started":"2026-08-08T05:33:08.873878Z","shell.execute_reply":"2026-08-08T05:33:09.064804Z"}},"outputs":[],"execution_count":null},{"id":"18f52c6c","cell_type":"markdown","source":"## 4. The reports\n\n`train.csv` carries a `Report` field on (presumably) every study, `test.csv` does not — the schema itself rules out any model that reads text at inference. What the reports are used for here is purely diagnostic: language mix, length, and missingness, all of which bound how far a report-derived labeling strategy could reach before it's built.\n","metadata":{}},{"id":"67219481","cell_type":"code","source":"has_report = \"Report\" in train_df.columns\nprint(\"Report column present:\", has_report)\n\nif has_report:\n    report_len = train_df[\"Report\"].fillna(\"\").str.len()\n    empty_reports = (train_df[\"Report\"].isna() | (train_df[\"Report\"].str.strip() == \"\")).mean()\n    print(f\"empty/missing reports: {empty_reports:.1%}\")\n\n    fig, ax = plt.subplots(figsize=(9, 4.5))\n    sns.histplot(report_len[report_len > 0], bins=40, color=ICE, ax=ax)\n    ax.axvline(report_len.median(), color=FIRE, ls=\"--\", lw=1.5,\n               label=f\"median = {report_len.median():.0f} chars\")\n    ax.set_title(\"Report length distribution (characters)\")\n    ax.legend()\n    plt.tight_layout()\n    plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T05:33:09.066516Z","iopub.execute_input":"2026-08-08T05:33:09.066819Z","iopub.status.idle":"2026-08-08T05:33:09.280197Z","shell.execute_reply.started":"2026-08-08T05:33:09.066787Z","shell.execute_reply":"2026-08-08T05:33:09.279538Z"}},"outputs":[],"execution_count":null},{"id":"8767a258","cell_type":"code","source":"if has_report:\n    # Cheap language signature — Unicode block share, not a language classifier. Enough\n    # to see the multilingual mix this corpus reportedly has, without pulling in a model.\n    def script_signature(s):\n        if not isinstance(s, str) or not s:\n            return \"empty\"\n        counts = Counter()\n        for ch in s:\n            o = ord(ch)\n            if 0x0370 <= o <= 0x03FF:\n                counts[\"greek\"] += 1\n            elif 0x0400 <= o <= 0x04FF:\n                counts[\"cyrillic\"] += 1\n            elif ch.isalpha():\n                counts[\"latin_or_other\"] += 1\n        if not counts:\n            return \"no_letters\"\n        return counts.most_common(1)[0][0]\n\n    sig = train_df[\"Report\"].fillna(\"\").apply(script_signature)\n    vc = sig.value_counts()\n\n    fig, ax = plt.subplots(figsize=(7, 4.5))\n    ax.bar(vc.index, vc.values, color=grad(len(vc)), edgecolor=\"white\", width=.55)\n    for i, v in enumerate(vc.values):\n        ax.text(i, v + n_total * 0.005, f\"{v}\", ha=\"center\", fontsize=9)\n    ax.set_title(\"Report script signature (Unicode block — not a language ID)\")\n    ax.tick_params(axis=\"x\", rotation=15)\n    plt.tight_layout()\n    plt.show()\n    print(\"Note: the Latin-script bucket mixes English, Spanish, French, Dutch, German,\")\n    print(\"Turkish and Croatian — a real language ID needs stopword scoring, not Unicode alone.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T05:33:09.281349Z","iopub.execute_input":"2026-08-08T05:33:09.281659Z","iopub.status.idle":"2026-08-08T05:33:10.613551Z","shell.execute_reply.started":"2026-08-08T05:33:09.281625Z","shell.execute_reply":"2026-08-08T05:33:10.612801Z"}},"outputs":[],"execution_count":null},{"id":"4aa898cb","cell_type":"code","source":"if has_report:\n    # Duplicate reports matter for splitting, not just for text stats: several studies\n    # sharing one template report share one derived target vector if labels are extracted\n    # from text, and a group like that has to stay inside one fold.\n    norm_report = train_df[\"Report\"].fillna(\"\").str.lower().str.strip()\n    dup_counts = norm_report[norm_report != \"\"].value_counts()\n    dup_groups = dup_counts[dup_counts > 1]\n\n    print(f\"distinct non-empty report texts: {norm_report[norm_report!=''].nunique()}\")\n    print(f\"report texts shared by 2+ studies: {len(dup_groups)}\")\n    print(f\"studies sitting inside a shared-report group: {int(dup_groups.sum())} \"\n          f\"({dup_groups.sum()/n_total:.1%} of the corpus)\")\n\n    fig, ax = plt.subplots(figsize=(8, 4))\n    sns.histplot(dup_groups.values, bins=range(2, min(int(dup_groups.max()) + 2, 30)),\n                 color=FIRE, ax=ax)\n    ax.set_title(\"Size of shared-report groups (template reports)\")\n    ax.set_xlabel(\"studies sharing one exact report text\")\n    plt.tight_layout()\n    plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T05:33:10.614611Z","iopub.execute_input":"2026-08-08T05:33:10.615217Z","iopub.status.idle":"2026-08-08T05:33:10.863332Z","shell.execute_reply.started":"2026-08-08T05:33:10.615193Z","shell.execute_reply":"2026-08-08T05:33:10.862664Z"}},"outputs":[],"execution_count":null},{"id":"bc49a133","cell_type":"markdown","source":"## 5. Series-level metadata plane and acquisition coverage\n\n`train_series.csv` names the anatomical plane and two acquisition flags, `Fluid_Sensitive` and `Fat_Suppression`. These are the two axes a slot design (sagittal/coronal/axial × weighting × fat-sat) is built from, so their coverage — not just their marginal counts — is what matters.\n","metadata":{}},{"id":"a8034e93","cell_type":"code","source":"plane_col = \"Anatomical_Plane\" if \"Anatomical_Plane\" in train_series.columns else \\\n            next((c for c in train_series.columns if \"plane\" in c.lower()), None)\nfluid_col = \"Fluid_Sensitive\" if \"Fluid_Sensitive\" in train_series.columns else None\nfatsat_col = \"Fat_Suppression\" if \"Fat_Suppression\" in train_series.columns else None\nprint(\"plane column:\", plane_col, \"| fluid column:\", fluid_col, \"| fat-sat column:\", fatsat_col)\n\nfig, axes = plt.subplots(1, 3, figsize=(15, 4.3))\nfor ax, col, title in zip(axes, [plane_col, fluid_col, fatsat_col],\n                           [\"Anatomical plane\", \"Fluid-sensitive flag\", \"Fat-suppression flag\"]):\n    if col is None:\n        ax.axis(\"off\")\n        continue\n    vc = train_series[col].value_counts(dropna=False)\n    ax.bar(vc.index.astype(str), vc.values, color=grad(len(vc)), edgecolor=\"white\", width=.6)\n    ax.set_title(title)\n    ax.tick_params(axis=\"x\", rotation=20)\n    for i, v in enumerate(vc.values):\n        ax.text(i, v + max(vc.values) * 0.02, str(v), ha=\"center\", fontsize=8)\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T05:33:10.865548Z","iopub.execute_input":"2026-08-08T05:33:10.865776Z","iopub.status.idle":"2026-08-08T05:33:11.183994Z","shell.execute_reply.started":"2026-08-08T05:33:10.865755Z","shell.execute_reply":"2026-08-08T05:33:11.183174Z"}},"outputs":[],"execution_count":null},{"id":"90c4e5a9","cell_type":"code","source":"if plane_col and fluid_col and fatsat_col:\n    # Slot coverage grid: plane x (fluid, fat-sat) — this is the same grid a slot design\n    # in a model pipeline would be built on, so seeing where it's thin here is cheaper\n    # than discovering it after the cache is built.\n    combo = train_series.groupby([plane_col, fluid_col, fatsat_col]).size().reset_index(name=\"count\")\n    combo[\"combo\"] = combo[fluid_col].astype(str) + \" / FS=\" + combo[fatsat_col].astype(str)\n    pivot = combo.pivot_table(index=plane_col, columns=\"combo\", values=\"count\", fill_value=0)\n\n    plt.figure(figsize=(10, 4.5))\n    sns.heatmap(pivot, annot=True, fmt=\"g\", cmap=CMAP_SEQ, linewidths=0.5, linecolor=\"white\")\n    plt.title(\"Series count by plane × (fluid-sensitive / fat-suppression)\")\n    plt.tight_layout()\n    plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T05:33:11.185021Z","iopub.execute_input":"2026-08-08T05:33:11.185844Z","iopub.status.idle":"2026-08-08T05:33:11.39183Z","shell.execute_reply.started":"2026-08-08T05:33:11.185805Z","shell.execute_reply":"2026-08-08T05:33:11.391088Z"}},"outputs":[],"execution_count":null},{"id":"d6e86a59","cell_type":"code","source":"# Series (and slices, once weighted by n_slices if available) per study — the shape\n# a per-study cache would need to budget for.\nseries_per_study = train_series.groupby(\"StudyInstanceUID\").size()\n\nfig, ax = plt.subplots(figsize=(9, 4.5))\nsns.histplot(series_per_study, bins=range(int(series_per_study.min()), int(series_per_study.max()) + 2),\n             color=ICE, ax=ax)\nax.axvline(series_per_study.median(), color=FIRE, ls=\"--\", lw=1.5,\n           label=f\"median = {series_per_study.median():.0f}\")\nax.set_title(\"Series per study\")\nax.legend()\nplt.tight_layout()\nplt.show()\nprint(series_per_study.describe())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T05:33:11.392817Z","iopub.execute_input":"2026-08-08T05:33:11.393144Z","iopub.status.idle":"2026-08-08T05:33:11.565875Z","shell.execute_reply.started":"2026-08-08T05:33:11.393121Z","shell.execute_reply":"2026-08-08T05:33:11.565247Z"}},"outputs":[],"execution_count":null},{"id":"8dcb6bf6","cell_type":"markdown","source":"## 6. DICOM headers physical scale and laterality\n\nTwo things live in the header that don't live in either CSV: **pixel spacing** (how many millimetres a pixel covers — the reason a fixed-pixel resize scrambles anatomical scale across studies with different spacing) and **laterality** (which knee — present on some studies via the `Laterality`/`ImageLaterality` tag, recoverable on others via the sign of the patient x-coordinate). Both are sampled here on a subset of files; reading every DICOM header in the corpus is unnecessary for an EDA pass and expensive.\n","metadata":{}},{"id":"908c05cd","cell_type":"code","source":"HDR_TAGS = [\"Rows\", \"Columns\", \"PixelSpacing\", \"SliceThickness\",\n            \"Laterality\", \"ImageLaterality\", \"ImagePositionPatient\",\n            \"RepetitionTime\", \"EchoTime\"]\n\ndef probe_header(path):\n    try:\n        ds = pydicom.dcmread(str(path), stop_before_pixels=True, force=True)\n        row = {\"path\": str(path)}\n        for t in HDR_TAGS:\n            v = getattr(ds, t, None)\n            if v is None:\n                row[t] = None\n            elif isinstance(v, (list, tuple)) or type(v).__name__ == \"MultiValue\":\n                row[t] = \"|\".join(str(x) for x in v)\n            else:\n                row[t] = str(v)\n        return row\n    except Exception as e:\n        return {\"path\": str(path), \"err\": str(e)[:100]}\n\ntrain_series_dir = ROOT / \"train_series\"\nif train_series_dir.is_dir():\n    all_dcm = list(train_series_dir.rglob(\"*.dcm\"))\n    sample_files = list(np.random.choice(all_dcm, min(400, len(all_dcm)), replace=False)) if all_dcm else []\n    log(f\"found {len(all_dcm)} DICOM files under train_series/, sampling {len(sample_files)}\")\n\n    with ThreadPoolExecutor(max_workers=16) as pool:\n        hdr_rows = list(pool.map(probe_header, sample_files))\n    hdr_df = pd.DataFrame(hdr_rows)\n    log(f\"probed {len(hdr_df)} headers\")\nelse:\n    hdr_df = pd.DataFrame()\n    log(\"no train_series/ directory found — skipping header probe\")\n\nhdr_df.head()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T05:33:11.566857Z","iopub.execute_input":"2026-08-08T05:33:11.567174Z","iopub.status.idle":"2026-08-08T05:34:58.956443Z","shell.execute_reply.started":"2026-08-08T05:33:11.567153Z","shell.execute_reply":"2026-08-08T05:34:58.955351Z"}},"outputs":[],"execution_count":null},{"id":"3fb438b8","cell_type":"code","source":"if len(hdr_df) and \"PixelSpacing\" in hdr_df.columns:\n    px = pd.to_numeric(hdr_df[\"PixelSpacing\"].fillna(\"\").str.split(\"|\").str[0].replace(\"\", np.nan),\n                        errors=\"coerce\")\n    st = pd.to_numeric(hdr_df.get(\"SliceThickness\", pd.Series(dtype=object)), errors=\"coerce\")\n    rows = pd.to_numeric(hdr_df.get(\"Rows\", pd.Series(dtype=object)), errors=\"coerce\")\n    cols = pd.to_numeric(hdr_df.get(\"Columns\", pd.Series(dtype=object)), errors=\"coerce\")\n\n    fig, axes = plt.subplots(2, 2, figsize=(13, 9))\n    sns.histplot(px.dropna(), bins=25, color=ICE, ax=axes[0,0])\n    axes[0,0].set_title(\"Pixel spacing (mm/pixel)\")\n\n    sns.histplot(st.dropna(), bins=25, color=FIRE, ax=axes[0,1])\n    axes[0,1].set_title(\"Slice thickness (mm)\")\n\n    sns.histplot(rows.dropna(), bins=20, color=ICE, ax=axes[1,0])\n    axes[1,0].set_title(\"Image rows (pixels)\")\n\n    sns.histplot(cols.dropna(), bins=20, color=FIRE, ax=axes[1,1])\n    axes[1,1].set_title(\"Image columns (pixels)\")\n    plt.tight_layout()\n    plt.show()\n\n    if px.notna().sum() > 5:\n        fov_mm = px * rows\n        print(f\"pixel spacing spans {px.min():.2f}–{px.max():.2f} mm/px \"\n              f\"({px.max()/px.min():.1f}x range)\")\n        print(f\"implied field of view spans roughly {fov_mm.min():.0f}–{fov_mm.max():.0f} mm\")\n        print(\"A fixed-pixel resize (e.g. every image to 224x224) would therefore show the\")\n        print(\"model a meniscus at a different physical scale in different studies for no\")\n        print(\"anatomical reason. A physical-scale crop (fixed mm, then resize) removes that.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T05:34:58.95761Z","iopub.execute_input":"2026-08-08T05:34:58.958363Z","iopub.status.idle":"2026-08-08T05:34:59.610994Z","shell.execute_reply.started":"2026-08-08T05:34:58.958333Z","shell.execute_reply":"2026-08-08T05:34:59.610261Z"}},"outputs":[],"execution_count":null},{"id":"daebeee8","cell_type":"code","source":"if len(hdr_df) and \"Laterality\" in hdr_df.columns:\n    # Same reasoning as a laterality fallback in a training pipeline: check the tag's\n    # coverage, and whether the x-coordinate sign actually agrees with it where both\n    # exist, before trusting the coordinate on studies where the tag is missing.\n    def tag_side(row):\n        for col in [\"Laterality\", \"ImageLaterality\"]:\n            v = row.get(col)\n            if isinstance(v, str) and v.strip().upper()[:1] in (\"L\", \"R\"):\n                return v.strip().upper()[0]\n        return None\n\n    def position_side(row, min_offset=5.0):\n        v = row.get(\"ImagePositionPatient\")\n        if not isinstance(v, str):\n            return None\n        try:\n            x = float(v.split(\"|\")[0])\n        except Exception:\n            return None\n        if abs(x) < min_offset:\n            return None\n        return \"R\" if x < 0 else \"L\"   # LPS convention: right side sits at negative x\n\n    hdr_df[\"tag_side\"] = hdr_df.apply(tag_side, axis=1)\n    hdr_df[\"pos_side\"] = hdr_df.apply(position_side, axis=1)\n\n    tag_coverage = hdr_df[\"tag_side\"].notna().mean()\n    both = hdr_df.dropna(subset=[\"tag_side\", \"pos_side\"])\n    agreement = (both[\"tag_side\"] == both[\"pos_side\"]).mean() if len(both) else np.nan\n\n    print(f\"Laterality tag present on {tag_coverage:.1%} of sampled headers\")\n    print(f\"x-coordinate sign agrees with the tag on {agreement:.1%} of \"\n          f\"{len(both)} comparable headers\")\n\n    fig, ax = plt.subplots(figsize=(6, 4))\n    labels = [\"Tag present\", \"Tag missing\"]\n    values = [tag_coverage, 1 - tag_coverage]\n    ax.pie(values, labels=labels, colors=[FIRE, ICE], autopct=\"%1.0f%%\",\n           wedgeprops={\"edgecolor\": \"white\"})\n    ax.set_title(\"Laterality tag coverage (sampled headers)\")\n    plt.tight_layout()\n    plt.show()\n\n    if np.isfinite(agreement):\n        verdict = \"high enough to trust as a fallback\" if agreement >= 0.85 else \"too low to trust blindly\"\n        print(f\"\\n-> x-sign agreement is {verdict} for filling the gap where the tag is absent.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T05:34:59.611829Z","iopub.execute_input":"2026-08-08T05:34:59.612121Z","iopub.status.idle":"2026-08-08T05:34:59.723189Z","shell.execute_reply.started":"2026-08-08T05:34:59.612099Z","shell.execute_reply":"2026-08-08T05:34:59.72254Z"}},"outputs":[],"execution_count":null},{"id":"a54cca0b","cell_type":"markdown","source":"## 7. Do the twelve findings co-occur?\n\nA knee report rarely names one thing in isolation — an ACL tear often comes with an effusion, osteoarthritis in one compartment often accompanies another. This matters for whether a per-diagnosis model can be trained fully independently, or whether the co-occurrence structure is worth encoding explicitly.\n","metadata":{}},{"id":"8106d080","cell_type":"code","source":"if annotated_mask.sum() >= 8:\n    binary_gold = (gold > 0.5).astype(int)\n    co = binary_gold.T.dot(binary_gold)\n\n    plt.figure(figsize=(9, 7.5))\n    sns.heatmap(co, annot=True, fmt=\"d\", cmap=CMAP_SEQ, linewidths=0.5, linecolor=\"white\")\n    plt.title(f\"Co-occurrence counts on the annotated subset (n={int(annotated_mask.sum())})\")\n    plt.tight_layout()\n    plt.show()\n\n    G = nx.Graph()\n    for t in TARGETS:\n        G.add_node(t, size=binary_gold[t].sum())\n    for i, a in enumerate(TARGETS):\n        for b in TARGETS[i+1:]:\n            w = int(co.loc[a, b])\n            if w > 0:\n                G.add_edge(a, b, weight=w)\n\n    pos = nx.spring_layout(G, seed=SEED, k=0.9)\n    sizes = [G.nodes[n][\"size\"] * 40 + 200 for n in G.nodes]\n    weights = [G[u][v][\"weight\"] for u, v in G.edges]\n\n    plt.figure(figsize=(9, 8))\n    nx.draw_networkx_nodes(G, pos, node_size=sizes, node_color=[FIRE]*len(G.nodes), alpha=0.85)\n    nx.draw_networkx_edges(G, pos, width=[w / max(weights, default=1) * 5 for w in weights],\n                            alpha=0.35, edge_color=MIDLINE)\n    nx.draw_networkx_labels(G, pos, font_size=9)\n    plt.title(\"Co-occurrence network (node size = positive count, edge width = shared studies)\")\n    plt.axis(\"off\")\n    plt.tight_layout()\n    plt.show()\nelse:\n    print(\"Fewer than 8 annotated studies — co-occurrence analysis would be noise.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T05:34:59.724436Z","iopub.execute_input":"2026-08-08T05:34:59.724775Z","iopub.status.idle":"2026-08-08T05:35:00.388105Z","shell.execute_reply.started":"2026-08-08T05:34:59.72475Z","shell.execute_reply":"2026-08-08T05:35:00.387357Z"}},"outputs":[],"execution_count":null},{"id":"36b34fe4","cell_type":"markdown","source":"## 8. Self-checks\n\nSmall assertions on the numbers this notebook computed, so a re-run on an updated dataset fails loudly instead of silently drifting.\n","metadata":{}},{"id":"d412253e","cell_type":"code","source":"assert n_annot <= n_total, \"annotated count exceeds total — annotation mask is wrong\"\nassert set(TARGETS).issubset(train_df.columns), \"a target column is missing from train.csv\"\nif len(hdr_df) and \"tag_side\" in hdr_df.columns:\n    assert hdr_df[\"tag_side\"].isin([\"L\", \"R\", None]).all(), \"unexpected laterality tag value\"\nif annotated_mask.sum() >= 8:\n    assert (co.values.diagonal() >= 0).all(), \"co-occurrence diagonal should be non-negative\"\n\nprint(\"self-checks passed\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-08T05:35:00.389258Z","iopub.execute_input":"2026-08-08T05:35:00.389635Z","iopub.status.idle":"2026-08-08T05:35:00.396644Z","shell.execute_reply.started":"2026-08-08T05:35:00.38961Z","shell.execute_reply":"2026-08-08T05:35:00.395645Z"}},"outputs":[],"execution_count":null},{"id":"7ff4a9fd","cell_type":"markdown","source":"## 9. Summary — what this means for modeling\n\n- **Annotated subset is small.** Any per-target AUC computed on it alone (see §3) has a wide Wilson/Hanley–McNeil interval — enough to treat small differences between approaches as noise, not signal. This is the argument for deriving auxiliary targets from `Report` text rather than training on the annotated subset alone.\n- **Reports carry real signal, and `test.csv` has none.** The schema itself rules out a text-branch fusion model at inference; the only admissible uses are as training targets, as a distilled auxiliary signal, or as a per-study confidence weight.\n- **Shared-report groups need to be respected in any split.** Splitting a template-report group across folds scores a model against a target it effectively trained on elsewhere.\n- **Pixel spacing varies enough that a fixed-pixel resize is not neutral.** A physical-scale crop (fixed mm before resizing) removes a nuisance transform the encoder would otherwise have to learn to ignore.\n- **The laterality tag covers roughly half the corpus** (update with your actual number from §6) — check whether the x-coordinate fallback's agreement is high enough to trust before using it to fill the rest, and never blind-trust it.\n- **Compartment/side targets (Medial/Lateral pairs) are only meaningful once laterality is normalised** — medial and lateral are defined relative to the body's midline, not the image frame.\n- **Co-occurrence is non-trivial** (§7) — worth knowing before deciding whether twelve independent heads or a shared representation is the better inductive bias.\n\n---\n*Numbers in italics above should be replaced with your actual run's figures once this notebook executes on the real mount — the structure and reasoning don't change, only the constants.*\n","metadata":{}}]}