{"cells":[{"cell_type":"markdown","metadata":{},"source":"# Two targets get skipped. Only one should be.\n\n`rsna-tonylica-repro`'s production pipeline (the `0.920`-scoring public repro this lane already\nverified, cites Mattia Angeli's `0.917` work) hardcodes one line most forks never look at:\n\n```python\n_RAD_EXCLUDE = (\"Baker's\", \"Fracture\")\n```\n\nThose two of the competition's twelve targets never get the RadImageNet-blend treatment every\nother target gets — the DINO baseline's raw prediction ships untouched for them. That's a\nreasonable-sounding engineering call (maybe those two don't benefit, or a second architecture adds\nnoise on a rare class). This notebook checks it against real OOF data instead of trusting the\ncomment.\n\n**Method, in one line:** for each of the 12 targets, compute the AUC lift from blending in the\nE13 (fat-sensitive-crop RadImageNet) head's own OOF versus the baseline alone, then see whether\nthe two hardcoded-excluded targets actually sit at the bottom of that list. If you maintain any\nper-target include/exclude rule in your own ensemble, this is the five-line check to run before\nyou trust it."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import glob, hashlib\nimport numpy as np\nimport pandas as pd\nfrom scipy.stats import rankdata\n\n# Dataset pinned by name (L8): tonylica/rsna-knee-bend-dinov3-0917-repro-assets ships OOF\n# predictions from the E13 (fat-sensitive-crop) RadImageNet head alone. pilkwang/rsna-knee-weights\n# ships the baseline ensemble's own OOF (\"ours\"), real labels, and the expert-labelled gold subset.\n# Same join this lane already verified in `0-35-beat-0-50-offline-the-board-said-no`.\ne13_oof_hits = sorted(glob.glob(\"/kaggle/input/**/kernel-sources/rsna-knee-e13-train/**/v52_e11_oof.csv\", recursive=True))\ne13_pt_hits = sorted(glob.glob(\"/kaggle/input/**/kernel-sources/rsna-knee-e13-train/**/v52_e11_heads.pt\", recursive=True))\nmerge_hits = sorted(glob.glob(\"/kaggle/input/**/merge_gain.npz\", recursive=True))\nassert e13_oof_hits, \"tonylica/rsna-knee-bend-dinov3-0917-repro-assets not mounted, or the e13-train OOF moved\"\nassert e13_pt_hits, \"E13 checkpoint (.pt) not found under the e13-train path\"\n# The dataset also ships an e11-train run under a sibling folder with an identically-named\n# v52_e11_oof.csv (different content, different file size). An unscoped \"**/v52_e11_oof.csv\"\n# glob matches both and silently picks whichever sorts first -- scope to the e13-train path,\n# exactly like the .pt glob already does, so this can never resolve to the wrong run.\nassert merge_hits, \"pilkwang/rsna-knee-weights not mounted?\"\n\ne13 = pd.read_csv(e13_oof_hits[0])\nd = np.load(merge_hits[0], allow_pickle=True)\nids, ours, y, gold = d[\"ids\"], d[\"ours\"].astype(np.float64), d[\"y\"], d[\"gold_mask\"]\ntargets = [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\", \"Medial OA\", \"Lateral OA\",\n           \"PF OA\", \"Effusion\", \"Synovitis\", \"Baker\\u0027s\", \"Contusion\", \"Fracture\"]\nEXCLUDED = (\"Baker\\u0027s\", \"Fracture\")  # _RAD_EXCLUDE, read verbatim from rsna-tonylica-repro cell 3\n\nprint(f\"studies: {len(ids)}, targets: {len(targets)}\")\nprint(\"E13 OOF ids match merge_gain ids, same order:\", np.array_equal(ids, e13[\"StudyInstanceUID\"].to_numpy()))\nprint(\"E13 OOF is_gold matches merge_gain gold_mask:\", np.array_equal(e13[\"is_gold\"].to_numpy().astype(bool), gold))\nprint(f\"gold (expert-labelled) subset: {int(gold.sum())} studies\")\n\nrad = e13[targets].to_numpy(dtype=np.float64)"},{"cell_type":"markdown","metadata":{},"source":"## Verify the OOF array is actually the E13 head, not a guess by file path\n\nSame check as `0-35-beat-0-50-offline-the-board-said-no`: the production pipeline pins its\ncheckpoint by sha256, not by filename. If the mounted `.pt` next to the OOF csv doesn't match\nthat pinned constant, everything past this cell is testing the wrong array."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"def sha256_file(path, chunk=8 << 20):\n    digest = hashlib.sha256()\n    with open(path, \"rb\") as handle:\n        for block in iter(lambda: handle.read(chunk), b\"\"):\n            digest.update(block)\n    return digest.hexdigest()\n\nE13_EXPECTED_SHA256 = \"ad9f19af73bfdf4e49263c0e45060dc3cb239e1195039b26dc8c0a3a6bcd1a8a\"\ngot = sha256_file(e13_pt_hits[0])\nprint(f\"expected: {E13_EXPECTED_SHA256}\")\nprint(f\"got:      {got}\")\nassert got == E13_EXPECTED_SHA256, \"checkpoint does not match the pipeline's own pinned hash\"\nprint(\"MATCH -- this OOF file is confirmed E13-head predictions, not a same-shape lookalike.\")"},{"cell_type":"markdown","metadata":{},"source":"## Sanity anchor before trusting anything new\n\nReproduces `0-35-beat-0-50-offline-the-board-said-no`'s already-published all-population and\ngold-58 numbers exactly. If this cell's asserts fail, stop — the per-target table below is not\ntrustworthy either."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"def auc_binary(y_true, scores):\n    n_pos = y_true.sum()\n    n_neg = len(y_true) - n_pos\n    if n_pos == 0 or n_neg == 0:\n        return None\n    ranks = rankdata(scores)\n    return (ranks[y_true == 1].sum() - n_pos * (n_pos + 1) / 2) / (n_pos * n_neg)\n\ndef macro_auc(pred, y, mask):\n    aucs = [a for a in (auc_binary(y[mask, t], pred[mask, t]) for t in range(y.shape[1])) if a is not None]\n    return float(np.mean(aucs))\n\ndef rank_blend(a, b, alpha):\n    ra = np.apply_along_axis(rankdata, 0, a)\n    rb = np.apply_along_axis(rankdata, 0, b)\n    return alpha * ra + (1 - alpha) * rb\n\nfull_mask = np.ones(len(y), dtype=bool)\nprint(\"=== Sanity anchor: reproduce the published table ===\")\nfor alpha in (0.00, 0.35, 0.50):\n    pred = rank_blend(rad, ours, alpha)\n    print(f\"alpha={alpha:.2f}  full-AUC={macro_auc(pred, y, full_mask):.4f}  gold58-AUC={macro_auc(pred, y, gold):.4f}\")\n\na0 = macro_auc(rank_blend(rad, ours, 0.00), y, full_mask)\nassert abs(a0 - 0.8497) < 0.0005, \"sanity check failed -- join or split logic has drifted\"\nprint(\"MATCH -- alpha=0.00 full-AUC reproduces the published 0.8497 anchor.\")"},{"cell_type":"markdown","metadata":{},"source":"## The new check: per-target lift, ranked, exclude list overlaid\n\n`alpha=0.35` is the nested-CV-preferred weight from this lane's own earlier notebook (not the\n`0.50` shipped default) — using the same weight both notebooks already agree is the stronger\noffline read. For each target: AUC with the E13 blend minus AUC without it, on the full 4,407-study\npopulation and on the 58 expert-labelled studies separately, sorted by full-population lift,\ndescending. `EXCLUDED` targets are flagged inline — the question is whether they cluster at the\nbottom."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"alpha = 0.35\npred = rank_blend(rad, ours, alpha)\nbase = rank_blend(rad, ours, 0.0)  # baseline: alpha=0 is \"ours\" alone, by construction\n\nrows = []\nfor j, t in enumerate(targets):\n    a_base = auc_binary(y[:, j], base[:, j])\n    a_blend = auc_binary(y[:, j], pred[:, j])\n    g_base = auc_binary(y[gold, j], base[gold, j])\n    g_blend = auc_binary(y[gold, j], pred[gold, j])\n    rows.append({\n        \"target\": t,\n        \"excluded\": t in EXCLUDED,\n        \"n_pos_full\": int(y[:, j].sum()),\n        \"full_delta\": a_blend - a_base,\n        \"n_pos_gold\": int(y[gold, j].sum()),\n        \"gold_delta\": g_blend - g_base,\n    })\n\ntable = pd.DataFrame(rows).sort_values(\"full_delta\", ascending=False).reset_index(drop=True)\ntable.insert(0, \"rank\", table.index + 1)\nwith pd.option_context(\"display.width\", 120, \"display.float_format\", \"{:.4f}\".format):\n    print(table.to_string(index=False))\n\nprint()\nexcl = table[table.excluded]\nincl = table[~table.excluded]\nprint(f\"mean full-population delta, EXCLUDED targets (n={len(excl)}): {excl.full_delta.mean():+.4f}\")\nprint(f\"mean full-population delta, INCLUDED targets (n={len(incl)}): {incl.full_delta.mean():+.4f}\")\nprint(f\"rank (of 12) of each excluded target by full-population lift: \" +\n      \", \".join(f\"{r.target}={r['rank']}\" for _, r in excl.iterrows()))"},{"cell_type":"markdown","metadata":{},"source":"## What the rank costs in the number that actually gets submitted\n\nThe table above ranks targets, but the pipeline is scored on one macro-AUC number, not twelve\nseparate ones. This cell answers the question the rank alone can't: how much of that number is the\nexclude list actually costing, right now, on the same OOF data."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"predA = pred.copy()          # config A: shipped -- both excluded targets held at baseline (alpha=0.0)\nfor name in EXCLUDED:\n    j = targets.index(name)\n    predA[:, j] = base[:, j]\n\npredB = predA.copy()         # config B: un-exclude Fracture only, Baker's stays excluded\njf = targets.index(\"Fracture\")\npredB[:, jf] = pred[:, jf]\n\npredC = pred                 # config C: exclude list removed entirely (alpha=0.35 on all 12)\n\nauc_A, gold_A = macro_auc(predA, y, full_mask), macro_auc(predA, y, gold)\nauc_B, gold_B = macro_auc(predB, y, full_mask), macro_auc(predB, y, gold)\nauc_C, gold_C = macro_auc(predC, y, full_mask), macro_auc(predC, y, gold)\n\nprint(f\"config A (shipped, both excluded):        full-AUC={auc_A:.6f}  gold58-AUC={gold_A:.6f}\")\nprint(f\"config B (only Fracture un-excluded):      full-AUC={auc_B:.6f}  gold58-AUC={gold_B:.6f}\")\nprint(f\"config C (exclude list removed entirely):  full-AUC={auc_C:.6f}  gold58-AUC={gold_C:.6f}\")\nprint()\nprint(f\"A -> B (un-exclude Fracture alone):   {auc_B - auc_A:+.6f} macro-AUC\")\nprint(f\"A -> C (drop the exclude list):       {auc_C - auc_A:+.6f} macro-AUC\")"},{"cell_type":"markdown","metadata":{},"source":"## What this actually tells you\n\nRead the printed rank column above, not this prose — this cell states no number the code above it\ndid not just print. One excluded target sits where you'd expect (near the bottom of the lift\nranking — the E13 blend does little for it). **The other does not.** It sits above the median of\nall twelve targets on the full population, ahead of several targets the production pipeline *does*\nblend — despite being one of only two hardcoded to skip the blend entirely.\n\nUn-excluding just that one target (config B) is worth **+0.00064 macro-AUC** on this OOF, on top of\nthe already-shipped 0.920 pipeline, for a one-line diff. Dropping the exclude list entirely\n(config C) — which also puts back the small amount `Baker's` was still worth — is worth **+0.00086**.\nNeither number is a promise about the private leaderboard; it is what the same OOF data this\npipeline was tuned on already says about a config nobody has actually scored yet.\n\n**What this does not tell you.** This is a two-source simplification of the pipeline's actual\nthree-stage blend (it skips the separate `reference`-arm nested blend and the `_RAD_V48_SECOND_\nALPHA=0.15` pass2 rescore stage, both real parts of the shipped pipeline this notebook does not\nreproduce), so it cannot say the shipped exclude list is a bug — the full pipeline may have a\nreason this simplification can't see (a clinical rationale, an interaction with the pass2 stage,\nor a private-test consideration never visible in this public OOF). The gold-58 column is included\nfor a reason: it is the expert-labelled ground truth, but at 12-19 positives per target the\nper-target gold delta is noisy enough to reverse sign on its own — treat it as a caution flag, not\na tiebreaker, and prefer the full-population column when the two disagree.\n\n**The actionable part.** If you maintain a per-target include/exclude rule in your own ensemble —\nhand-picked or inherited from a fork — this is the check to run before trusting it: rank every\ntarget by its own measured lift, on your own OOF, and see whether your rule actually tracks the\nranking. A rule that groups two targets together because they're both \"hard\" or \"rare\" is a guess\nuntil you've run this cell on them separately."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11.0"}},"nbformat":4,"nbformat_minor":5}