{"cells":[{"cell_type":"markdown","id":"b6c00a3e","metadata":{},"source":"# The sequence name survived de-identification\n\nRSNA de-identified these knee MRIs: the patient and institution name tags are gone. What did not go is the\nacquisition metadata. Every DICOM still carries its full sequence name in `SeriesDescription`, its scanner\n`Manufacturer` and model, the `ScanningSequence` and `ScanOptions`, the echo and repetition times, and the\nfield strength. This notebook measures, live on the real headers, two things that follow from that. First, de\nidentification HELD for the identifiers: the name tags are null across the corpus. Second, the surviving\nacquisition tags reconstruct the three per series attributes the competition hands you, `Fluid_Sensitive`,\n`Fat_Suppression` and `Anatomical_Plane`, so the sequence name is those columns in words. Then a mutual\ninformation pass asks the honest follow up: does any of that surviving metadata carry signal about the twelve\npathology labels.\n\nThis reads headers only (`stop_before_pixels`), shows aggregates and figures, never raw DICOM or report text,\nand makes no attempt to re identify anyone. No leaderboard claim."},{"cell_type":"markdown","id":"b58ecdcd","metadata":{},"source":"## 0. Config, and the columns the competition already gives you"},{"cell_type":"code","execution_count":null,"id":"ccba65f3","metadata":{},"outputs":[],"source":"import os, glob, re, warnings, math\nimport numpy as np, pandas as pd\nimport matplotlib as mpl, matplotlib.pyplot as plt\nwarnings.filterwarnings(\"ignore\")\nmpl.rcParams.update({\"axes.spines.top\": False, \"axes.spines.right\": False, \"figure.dpi\": 120,\n                     \"font.size\": 10.5, \"axes.grid\": True, \"grid.alpha\": 0.2, \"axes.unicode_minus\": False})\nINK, HL, BL, GD, MUT = \"#14212B\", \"#B23A2E\", \"#3d7ea8\", \"#C29A22\", \"#93a1b0\"\n\ndef resolve():\n    for c in [\"/kaggle/input/rsna-knee-abnormality-detection\",\n              \"/kaggle/input/competitions/rsna-knee-abnormality-detection\"]:\n        if os.path.exists(os.path.join(c, \"train.csv\")): return c\n    if os.path.isdir(\"/kaggle/input\"):\n        for dp, dirs, files in os.walk(\"/kaggle/input\"):\n            if \"train.csv\" in files: return dp\n            if dp.count(\"/\") - \"/kaggle/input\".count(\"/\") >= 3: dirs[:] = []\n    return \"D:/kaggle/rsna_data\"\nINPUT = resolve()\nBASE = \"train_series\" if os.path.isdir(os.path.join(INPUT, \"train_series\")) else None\nLABELS = [\"ACL\",\"MCL\",\"Medial Meniscus\",\"Lateral Meniscus\",\"Medial OA\",\"Lateral OA\",\n          \"PF OA\",\"Effusion\",\"Synovitis\",\"Baker's\",\"Contusion\",\"Fracture\"]\ntrain = pd.read_csv(os.path.join(INPUT, \"train.csv\"))\ntser = pd.read_csv(os.path.join(INPUT, \"train_series.csv\"))\nprint(\"INPUT\", INPUT, \"| train_series dir:\", bool(BASE))\nprint(\"train_series columns the competition provides:\", list(tser.columns))"},{"cell_type":"markdown","id":"2a76adf4","metadata":{},"source":"## 1. The gate: read headers and confirm what survived\n\nOne slice per series across a stratified sample. Header only. The gate asserts what the whole notebook rests\non: the name tags are null, and `SeriesDescription` is present and varied."},{"cell_type":"code","execution_count":null,"id":"c8d489c7","metadata":{},"outputs":[],"source":"import pydicom\ndef sample_series(n=400, seed=0):\n    rng = np.random.default_rng(seed)\n    ts = tser.copy()\n    idx = rng.permutation(len(ts))[:n]\n    return ts.iloc[idx]\nTAGS = [\"SeriesDescription\",\"Manufacturer\",\"ManufacturerModelName\",\"InstitutionName\",\"StationName\",\n        \"MagneticFieldStrength\",\"ScanningSequence\",\"ScanOptions\",\"EchoTime\",\"RepetitionTime\",\n        \"PixelSpacing\",\"SliceThickness\",\"PhotometricInterpretation\"]\nrows = []\nif BASE:\n    for _, r in sample_series(400).iterrows():\n        d = os.path.join(INPUT, BASE, r.StudyInstanceUID, r.SeriesInstanceUID)\n        fs = sorted(glob.glob(os.path.join(d, \"*.dcm\")))\n        if not fs: continue\n        try:\n            ds = pydicom.dcmread(fs[0], stop_before_pixels=True)\n        except Exception:\n            continue\n        rec = dict(StudyInstanceUID=r.StudyInstanceUID, plane=r.Anatomical_Plane,\n                   fluid=r.Fluid_Sensitive, fatsat=r.Fat_Suppression)\n        for t in TAGS: rec[t] = ds.get(t, None)\n        rows.append(rec)\nH = pd.DataFrame(rows)\nprint(\"headers read:\", len(H))\nif len(H):\n    surv = {t: float(H[t].notna().mean()) for t in TAGS}\n    print(\"SeriesDescription present:\", round(surv[\"SeriesDescription\"],3),\n          \"| distinct:\", H[\"SeriesDescription\"].nunique())\n    print(\"InstitutionName present:\", round(surv[\"InstitutionName\"],3),\n          \"| StationName present:\", round(surv[\"StationName\"],3))\n    # the honest gate\n    ok = surv[\"SeriesDescription\"] > 0.9 and H[\"SeriesDescription\"].nunique() > 1 and surv[\"InstitutionName\"] < 0.1\n    print(\"GATE (SeriesDescription survives, names stripped):\", \"PASS\" if ok else \"review\")\nelse:\n    print(\"headers unavailable in this environment; renders on Kaggle\")"},{"cell_type":"markdown","id":"1770ed24","metadata":{},"source":"## 2. What survived de-identification, and what did not"},{"cell_type":"code","execution_count":null,"id":"5d1a08a4","metadata":{},"outputs":[],"source":"if len(H):\n    order = [\"InstitutionName\",\"StationName\",\"SeriesDescription\",\"Manufacturer\",\"ManufacturerModelName\",\n             \"MagneticFieldStrength\",\"ScanningSequence\",\"ScanOptions\",\"EchoTime\",\"RepetitionTime\",\n             \"PixelSpacing\",\"SliceThickness\"]\n    pres = [float(H[t].notna().mean()) for t in order]\n    fig, ax = plt.subplots(figsize=(9, 5))\n    cols = [HL if t in (\"InstitutionName\",\"StationName\") else BL for t in order]\n    ax.barh(range(len(order)), pres, color=cols)\n    ax.set_yticks(range(len(order))); ax.set_yticklabels(order); ax.invert_yaxis()\n    ax.set_xlabel(\"fraction of sampled series where the tag is present\"); ax.set_xlim(0, 1.02)\n    for i, v in enumerate(pres): ax.text(v+0.01, i, f\"{v:.2f}\", va=\"center\", fontsize=8)\n    ax.set_title(\"Name tags stripped (red), acquisition tags kept (blue)\", color=INK, loc=\"left\", fontweight=\"bold\")\n    plt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"377e1d12","metadata":{},"source":"## 3. The naming dialect: the distinct sequence names"},{"cell_type":"code","execution_count":null,"id":"d899d27d","metadata":{},"outputs":[],"source":"if len(H):\n    vc = H[\"SeriesDescription\"].dropna().astype(str).str.lower().value_counts().head(20)\n    fig, ax = plt.subplots(figsize=(9, 6))\n    ax.barh(range(len(vc)), vc.values, color=BL); ax.invert_yaxis()\n    ax.set_yticks(range(len(vc))); ax.set_yticklabels(vc.index, fontsize=8)\n    ax.set_xlabel(\"series\")\n    ax.set_title(f\"Top sequence names ({H['SeriesDescription'].nunique()} distinct in the sample)\", color=INK, loc=\"left\", fontweight=\"bold\")\n    plt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"153b15cb","metadata":{},"source":"## 4. The sequence name reconstructs the plane"},{"cell_type":"code","execution_count":null,"id":"156c89a6","metadata":{},"outputs":[],"source":"if len(H):\n    def plane_from_desc(s):\n        s = str(s).lower()\n        if any(k in s for k in [\"sag\"]): return \"Sagittal\"\n        if any(k in s for k in [\"ax\",\"tra\"]): return \"Axial\"\n        if any(k in s for k in [\"cor\"]): return \"Coronal\"\n        return \"unknown\"\n    H[\"plane_guess\"] = H[\"SeriesDescription\"].apply(plane_from_desc)\n    sub = H[H[\"plane\"].isin([\"Sagittal\",\"Axial\",\"Coronal\"]) & (H[\"plane_guess\"]!=\"unknown\")]\n    planes = [\"Sagittal\",\"Axial\",\"Coronal\"]\n    C = np.zeros((3,3))\n    for i,a in enumerate(planes):\n        for j,b in enumerate(planes):\n            C[i,j] = ((sub.plane==a)&(sub.plane_guess==b)).sum()\n    fig, ax = plt.subplots(figsize=(5.4,4.6))\n    im=ax.imshow(C, cmap=\"Blues\")\n    ax.set_xticks(range(3)); ax.set_xticklabels(planes); ax.set_yticks(range(3)); ax.set_yticklabels(planes)\n    ax.set_xlabel(\"plane guessed from SeriesDescription\"); ax.set_ylabel(\"RSNA Anatomical_Plane\")\n    for i in range(3):\n        for j in range(3): ax.text(j,i,int(C[i,j]),ha=\"center\",va=\"center\",color=\"white\" if C[i,j]>C.max()*0.5 else \"black\")\n    acc = np.trace(C)/C.sum() if C.sum() else float(\"nan\")\n    ax.set_title(f\"Plane from the name matches RSNA plane ({acc:.0%} on the matched sample)\", color=INK, loc=\"left\", fontweight=\"bold\", fontsize=10)\n    plt.tight_layout(); plt.show()\n    print(\"plane reconstruction accuracy on matched sample:\", round(float(acc),3))"},{"cell_type":"markdown","id":"e00bd26d","metadata":{},"source":"## 5. Fat suppression and fluid sensitivity, decoded from tokens and timings"},{"cell_type":"code","execution_count":null,"id":"9e0d8904","metadata":{},"outputs":[],"source":"if len(H):\n    def has_fs(row):\n        s = str(row[\"SeriesDescription\"]).lower(); so = str(row.get(\"ScanOptions\",\"\")).lower()\n        return int(any(k in s for k in [\"fs\",\"stir\",\"spair\",\"fatsat\",\"fat_sat\",\"spir\"]) or \"fs\" in so)\n    def fluid_guess(row):\n        s = str(row[\"SeriesDescription\"]).lower()\n        te = row.get(\"EchoTime\", None)\n        long_te = (te is not None and not pd.isna(te) and float(te) >= 40)\n        return int(any(k in s for k in [\"stir\",\"t2\",\"pd\",\"spair\"]) or long_te)\n    H[\"fs_guess\"] = H.apply(has_fs, axis=1); H[\"fl_guess\"] = H.apply(fluid_guess, axis=1)\n    def acc(col, guess):\n        m = H[H[col].isin([0,1])]\n        return float((m[col].astype(int)==m[guess]).mean()) if len(m) else float(\"nan\")\n    a_fs = acc(\"fatsat\",\"fs_guess\"); a_fl = acc(\"fluid\",\"fl_guess\")\n    fig, ax = plt.subplots(figsize=(7,3.6))\n    ax.bar([0,1],[a_fs,a_fl], color=[BL,GD])\n    ax.set_xticks([0,1]); ax.set_xticklabels([\"Fat suppression\",\"Fluid sensitive\"])\n    ax.set_ylim(0,1); ax.set_ylabel(\"agreement with the RSNA attribute\")\n    for i,v in enumerate([a_fs,a_fl]): ax.text(i,v+0.02,f\"{v:.0%}\",ha=\"center\",fontweight=\"bold\")\n    ax.set_title(\"The provided attributes are the header, reconstructed\", color=INK, loc=\"left\", fontweight=\"bold\")\n    plt.tight_layout(); plt.show()\n    print(\"fat-suppression agreement\", round(a_fs,3), \"| fluid-sensitive agreement\", round(a_fl,3))"},{"cell_type":"markdown","id":"2074b21c","metadata":{},"source":"## 6. The scanner fleet, and field strength"},{"cell_type":"code","execution_count":null,"id":"2b10d862","metadata":{},"outputs":[],"source":"if len(H):\n    fig, (a1,a2) = plt.subplots(1,2, figsize=(13,4))\n    mv = H[\"Manufacturer\"].dropna().astype(str).str.strip().value_counts().head(8)\n    a1.bar(range(len(mv)), mv.values, color=BL); a1.set_xticks(range(len(mv))); a1.set_xticklabels(mv.index, rotation=25, ha=\"right\", fontsize=8)\n    a1.set_ylabel(\"series\"); a1.set_title(\"Manufacturer (surviving acquisition context)\", color=INK, loc=\"left\", fontweight=\"bold\")\n    fsv = pd.to_numeric(H[\"MagneticFieldStrength\"], errors=\"coerce\").dropna().round(1).value_counts().sort_index()\n    a2.bar([str(x) for x in fsv.index], fsv.values, color=GD); a2.set_xlabel(\"field strength (T)\"); a2.set_ylabel(\"series\")\n    a2.set_title(\"Field strength\", color=INK, loc=\"left\", fontweight=\"bold\")\n    plt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"6edb9a30","metadata":{},"source":"## 7. The honest follow-up: does the surviving metadata leak the labels\n\nOnly the labeled studies carry the 12 gold labels, a small set whose exact size prints below, so this\nmutual-information pass is high variance. It is an audit direction, not a verdict: a strong bar would say a\nmetadata field aligns with a label prevalence and deserves a scanner grouped cross validation, not that the\nlabel is trivially readable. Headers here are read for the labeled studies directly, not from the section 1\nsample, so the join is not thinned by sampling."},{"cell_type":"code","execution_count":null,"id":"4e703a4c","metadata":{},"outputs":[],"source":"from sklearn.metrics import mutual_info_score\ntry:\n    import pydicom\nexcept Exception:\n    pydicom = None\ngold = train[train[LABELS].notna().any(axis=1)].copy()\ngold_ids = gold[\"StudyInstanceUID\"].unique().tolist()\nprint(\"labeled (gold) studies:\", len(gold_ids))\ndef read_study_headers(ids):\n    rows = []\n    if BASE is None or pydicom is None:\n        return pd.DataFrame(columns=[\"StudyInstanceUID\",\"Manufacturer\",\"MagneticFieldStrength\"])\n    tg = tser[tser[\"StudyInstanceUID\"].isin(set(ids))]\n    for _, r in tg.iterrows():\n        d = os.path.join(INPUT, BASE, r.StudyInstanceUID, r.SeriesInstanceUID)\n        fsx = sorted(glob.glob(os.path.join(d, \"*.dcm\")))\n        if not fsx: continue\n        try: ds = pydicom.dcmread(fsx[0], stop_before_pixels=True)\n        except Exception: continue\n        rows.append(dict(StudyInstanceUID=r.StudyInstanceUID,\n                         Manufacturer=ds.get(\"Manufacturer\", None),\n                         MagneticFieldStrength=ds.get(\"MagneticFieldStrength\", None)))\n    return pd.DataFrame(rows)\nHg = read_study_headers(gold_ids)\nif len(Hg) and len(gold):\n    meta_study = Hg.groupby(\"StudyInstanceUID\").agg(\n        manuf=(\"Manufacturer\", lambda s: str(s.dropna().astype(str).mode().iloc[0]) if len(s.dropna()) else \"na\"),\n        field=(\"MagneticFieldStrength\", lambda s: round(float(pd.to_numeric(s,errors=\"coerce\").dropna().median()),1) if pd.to_numeric(s,errors=\"coerce\").notna().any() else -1),\n    ).reset_index()\n    g = gold.merge(meta_study, on=\"StudyInstanceUID\", how=\"inner\")\n    print(\"gold studies with header metadata joined:\", len(g))\n    feats = {\"Manufacturer\": g[\"manuf\"], \"FieldStrength\": g[\"field\"].astype(str)}\n    M = np.zeros((len(feats), 12))\n    for fi,(fn,fv) in enumerate(feats.items()):\n        for li,lab in enumerate(LABELS):\n            y = g[lab].fillna(0).astype(int)\n            M[fi,li] = mutual_info_score(fv, y)\n    fig, ax = plt.subplots(figsize=(11,2.8))\n    im=ax.imshow(M, cmap=\"magma\", aspect=\"auto\")\n    ax.set_yticks(range(len(feats))); ax.set_yticklabels(list(feats.keys()))\n    ax.set_xticks(range(12)); ax.set_xticklabels(LABELS, rotation=90, fontsize=8)\n    fig.colorbar(im, ax=ax, label=\"mutual information (nats)\")\n    ax.set_title(f\"Surviving metadata vs the 12 labels on {len(g)} labeled studies (tiny sample, high variance)\", color=INK, loc=\"left\", fontweight=\"bold\", fontsize=10)\n    plt.tight_layout(); plt.show()\n    print(\"max MI:\", round(float(M.max()),4), \"(on n=%d labeled studies; treat as a direction, not a verdict)\" % len(g))\nelse:\n    print(\"gold or headers unavailable here (renders on Kaggle)\")"},{"cell_type":"markdown","id":"216d6e5c","metadata":{},"source":"## What this shows, stated plainly\n\nDe identification HELD for the identifiers: across the sampled headers the institution and station name tags\nare null. What survives is acquisition metadata, and that metadata is not neutral: the sequence name plus a\ncouple of timing tags reconstruct the plane and the fat suppression and fluid sensitivity flags the\ncompetition hands you, so those provided columns are the header in words. The mutual information pass against\nthe twelve labels is on only the few dozen labeled studies and is therefore a direction to check with a\nscanner grouped cross validation, not a claim that any label is readable from metadata. This notebook reads headers only, shows\naggregates, never displays raw DICOM or report text, and makes no attempt to re identify anyone. No leaderboard\nposition is claimed."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python"}},"nbformat":4,"nbformat_minor":5}