{"cells":[{"cell_type":"markdown","id":"264f9a35","metadata":{},"source":"# The RSNA knee labels are not independent\n\nThe RSNA knee task gives twelve pathology labels, and the default way to model a multi label problem is twelve\nindependent heads. On the labeled studies that assumption is wrong: the labels co-occur in a structured way,\nand several of the strongest pairs are the associations a knee radiologist would expect.\n\nThe honest complication is that the labels are sparse, so the labeled set is small. This notebook does not\npretend otherwise. Instead of leaning on any single fragile correlation it tests the whole structure at once\nwith a permutation: each label column is shuffled independently, which destroys co-occurrence but keeps each\nlabel's prevalence, and the observed aggregate correlation is compared to that null. The observed structure\nlands far outside chance. The individual pairs are then shown for interpretation, with the small sample stated\nnext to them.\n\nEvery number and the permutation p value is computed live and printed. These are associations on the labeled\nstudies, not a causal or clinical claim. No leaderboard position is claimed."},{"cell_type":"markdown","id":"ff14a370","metadata":{},"source":"## 0. Config and the RSNA loader"},{"cell_type":"code","execution_count":null,"id":"1f0389a1","metadata":{},"outputs":[],"source":"import os, glob, warnings\nimport numpy as np, pandas as pd\nimport matplotlib as mpl, matplotlib.pyplot as plt\nwarnings.filterwarnings(\"ignore\")\nmpl.rcParams.update({\"axes.spines.top\": False, \"axes.spines.right\": False, \"figure.dpi\": 120,\n                     \"font.size\": 10.5, \"axes.grid\": True, \"grid.alpha\": 0.2, \"axes.unicode_minus\": False})\nINK=\"#1b1b2f\"; BL=\"#2f6df6\"; GD=\"#e8a63c\"; HL=\"#d1495b\"; GR=\"#8a8a99\"\n\ndef find_input():\n    for c in [\"/kaggle/input/rsna-knee-abnormality-detection\",\n              \"/kaggle/input/competitions/rsna-knee-abnormality-detection\"]:\n        if os.path.exists(os.path.join(c, \"train.csv\")): return c\n    hits = glob.glob(\"/kaggle/input/**/train.csv\", recursive=True)\n    if hits: return os.path.dirname(hits[0])\n    return \"D:/kaggle/rsna_data\"\nINPUT = find_input()\nLABELS = [\"ACL\",\"MCL\",\"Medial Meniscus\",\"Lateral Meniscus\",\"Medial OA\",\"Lateral OA\",\n          \"PF OA\",\"Effusion\",\"Synovitis\",\"Baker's\",\"Contusion\",\"Fracture\"]\nprint(\"INPUT:\", INPUT)"},{"cell_type":"markdown","id":"d58cddbc","metadata":{},"source":"## 1. The labeled studies, and how sparse they are\n\nLoad the training table and keep the studies that carry at least one of the twelve labels. The gate prints how\nmany that is, so the small sample is visible from the start, and asserts all twelve label columns are present."},{"cell_type":"code","execution_count":null,"id":"f013e4d6","metadata":{},"outputs":[],"source":"train = pd.read_csv(os.path.join(INPUT, \"train.csv\"))\nassert all(c in train.columns for c in LABELS), \"missing label columns\"\nlab = train[train[LABELS].notna().any(axis=1)].copy()\nB = (lab[LABELS] == 1).astype(int)\nn, k = B.shape\nprint(\"total training rows:\", len(train), \"| labeled studies:\", n, \"| labels:\", k)\nprev = B.mean().sort_values(ascending=False)\nprint(\"GATE: 12 labels present; labeled studies n =\", n)\nprint(\"positive prevalence (on labeled studies):\")\nfor lb, p in prev.items(): print(f\"  {lb:18s} {p:.3f}\")"},{"cell_type":"markdown","id":"7dc08d21","metadata":{},"source":"## 2. Prevalence, so the correlations are read in context\n\nHow common each label is among the labeled studies. Effusion and Synovitis are frequent; MCL and Lateral OA\nare rarest. Correlations between rare labels are the noisiest, which is why the aggregate permutation test\nmatters more than any single cell."},{"cell_type":"code","execution_count":null,"id":"51a3b3e5","metadata":{},"outputs":[],"source":"if n:\n    fig, ax = plt.subplots(figsize=(9.5,4.4))\n    ax.barh(range(len(prev)), prev.values, color=BL); ax.invert_yaxis()\n    ax.set_yticks(range(len(prev))); ax.set_yticklabels(prev.index, fontsize=9)\n    for i, v in enumerate(prev.values): ax.text(v+0.005, i, f\"{v:.2f}\", va=\"center\", fontsize=8)\n    ax.set_xlabel(\"positive rate among labeled studies\"); ax.set_xlim(0, min(1.0, prev.max()+0.12))\n    ax.set_title(f\"Label prevalence on the {n} labeled studies\", color=INK, loc=\"left\", fontweight=\"bold\", fontsize=11)\n    plt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"d572b40d","metadata":{},"source":"## 3. The co-occurrence structure\n\nThe pairwise correlation of the twelve labels across the labeled studies. Warm cells are labels that tend to\nappear together. The matrix is not diagonal: there is real off-diagonal structure, which the next cell tests\nagainst chance."},{"cell_type":"code","execution_count":null,"id":"f2738086","metadata":{},"outputs":[],"source":"if n:\n    C = np.corrcoef(B.values.T)\n    fig, ax = plt.subplots(figsize=(7.8,6.6))\n    im = ax.imshow(C, cmap=\"RdBu_r\", vmin=-0.5, vmax=0.5)\n    ax.set_xticks(range(k)); ax.set_xticklabels(LABELS, rotation=90, fontsize=8)\n    ax.set_yticks(range(k)); ax.set_yticklabels(LABELS, fontsize=8)\n    fig.colorbar(im, ax=ax, fraction=0.046, label=\"correlation\")\n    ax.set_title(\"Label co-occurrence correlation\", color=INK, loc=\"left\", fontweight=\"bold\", fontsize=11)\n    plt.tight_layout(); plt.show()\n    iu = np.triu_indices(k, 1)\n    offdiag = C[iu][~np.isnan(C[iu])]\n    print(\"mean absolute off-diagonal correlation:\", round(float(np.mean(np.abs(offdiag))),4))\n    print(\"max pair correlation:\", round(float(np.nanmax(offdiag)),3))"},{"cell_type":"markdown","id":"80e51b78","metadata":{},"source":"## 4. The permutation test: the structure is not chance\n\nThe rigorous core. Shuffle each label column independently five thousand times; this keeps every label's\nprevalence but destroys any co-occurrence, giving the distribution of aggregate correlation expected under\nindependence. The observed value sits far to the right of that null, so the co-occurrence structure is real\ndespite the small sample."},{"cell_type":"code","execution_count":null,"id":"7f09e217","metadata":{},"outputs":[],"source":"if n:\n    def mean_abs_offdiag(mat):\n        c = np.corrcoef(mat.T); v = c[np.triu_indices(k,1)]; v = v[~np.isnan(v)]\n        return float(np.mean(np.abs(v)))\n    obs = mean_abs_offdiag(B.values)\n    rng = np.random.default_rng(0)\n    null = np.empty(5000)\n    Bv = B.values\n    for t in range(5000):\n        bs = Bv.copy()\n        for j in range(k): bs[:, j] = rng.permutation(bs[:, j])\n        null[t] = mean_abs_offdiag(bs)\n    pval = (np.sum(null >= obs) + 1) / (len(null) + 1)\n    z = (obs - null.mean()) / null.std()\n    fig, ax = plt.subplots(figsize=(9.5,4.4))\n    ax.hist(null, bins=50, color=GR, alpha=0.85, label=\"null (labels shuffled)\")\n    ax.axvline(obs, color=HL, lw=2.5, label=f\"observed {obs:.3f}\")\n    ax.set_xlabel(\"mean absolute off-diagonal correlation\"); ax.set_ylabel(\"permutations\")\n    ax.legend(); ax.set_title(f\"Observed structure is {z:.1f} sigma beyond chance (permutation p = {pval:.4g})\",\n                              color=INK, loc=\"left\", fontweight=\"bold\", fontsize=11)\n    plt.tight_layout(); plt.show()\n    print(\"observed mean|corr|:\", round(obs,4), \"| null mean:\", round(float(null.mean()),4),\n          \"| 95th pct:\", round(float(np.percentile(null,95)),4))\n    print(\"permutation p (structure > chance):\", pval, \"| z vs null:\", round(float(z),2))"},{"cell_type":"markdown","id":"f7f4bc49","metadata":{},"source":"## 5. The strongest pairs, by lift over independence\n\nFor each label pair, lift is the joint positive rate divided by what independence would predict, the product of\nthe two prevalences. A lift above one means the two labels appear together more than chance. The top pairs are\nosteoarthritis with a Baker's cyst, one compartment's osteoarthritis with another's, meniscus damage with\nosteoarthritis, and a bone contusion with a fracture."},{"cell_type":"code","execution_count":null,"id":"65bef6d9","metadata":{},"outputs":[],"source":"if n:\n    p = B.mean().values\n    rows = []\n    for i in range(k):\n        for j in range(i+1, k):\n            joint = float((B.iloc[:,i] & B.iloc[:,j]).mean())\n            exp = p[i]*p[j]\n            lift = joint/exp if exp > 0 else np.nan\n            rows.append((LABELS[i], LABELS[j], joint, exp, lift))\n    L = pd.DataFrame(rows, columns=[\"a\",\"b\",\"joint\",\"expected\",\"lift\"]).dropna()\n    L = L[L[\"joint\"] > 0].sort_values(\"lift\", ascending=False).head(10)\n    fig, ax = plt.subplots(figsize=(10,4.6))\n    lbl = [f\"{a} + {b}\" for a,b in zip(L[\"a\"], L[\"b\"])]\n    ax.barh(range(len(L)), L[\"lift\"].values, color=BL); ax.invert_yaxis()\n    ax.axvline(1.0, color=GR, lw=1.5, ls=\"--\")\n    ax.set_yticks(range(len(L))); ax.set_yticklabels(lbl, fontsize=8)\n    for i, v in enumerate(L[\"lift\"].values): ax.text(v+0.02, i, f\"{v:.2f}x\", va=\"center\", fontsize=8)\n    ax.set_xlabel(\"lift over independence (joint rate / product of prevalences)\")\n    ax.set_title(\"Label pairs that co-occur more than chance (top 10 by lift)\", color=INK, loc=\"left\", fontweight=\"bold\", fontsize=10.5)\n    plt.tight_layout(); plt.show()\n    print(L.assign(joint=L['joint'].round(3), expected=L['expected'].round(3), lift=L['lift'].round(2)).to_string(index=False))"},{"cell_type":"markdown","id":"12d3d7f8","metadata":{},"source":"## What this shows, stated plainly\n\nOn the labeled RSNA knee studies the twelve pathology labels do not behave independently. The co-occurrence\nmatrix has clear off-diagonal structure, and a permutation that keeps each label's prevalence but destroys\nco-occurrence puts the observed aggregate correlation many standard deviations beyond chance, so the structure\nis real even though the labeled sample is small. The strongest pairs by lift are osteoarthritis with a Baker's\ncyst, meniscus damage with osteoarthritis, and a contusion with a fracture. The practical reading is that a\nmodel with twelve independent heads leaves signal on the table: the label correlations are informative and a\njoint or correlation aware head can use them. The sample is small and these are associations on the labeled\nstudies, not a causal or clinical claim. All values are computed live. No leaderboard position is claimed.\n\nCompanions on the same competition:\n\n- [The sequence name survived de-identification](https://www.kaggle.com/code/busyaprime/the-sequence-name-survived-de-identification)\n- [Osteoarthritis is almost never written as OA](https://www.kaggle.com/code/busyaprime/osteoarthritis-is-almost-never-written-as-oa)\n- [Read the report then the knee silver baseline](https://www.kaggle.com/code/busyaprime/read-the-report-then-the-knee-silver-baseline)"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python"}},"nbformat":4,"nbformat_minor":5}