{"cells":[{"cell_type":"markdown","metadata":{},"source":"# Should you blend in a second RSNA prediction?\n\nOur reproduced pilkwang ensemble scores **0.891** on the public leaderboard (submission\n`55613482`, rank ~932/1,938, inside a ~90-team tie block of everyone else running the same public\ncheckpoint family). This notebook isn't about that score. It's a control: **before you spend\ninference budget merging a second prediction source into your own ensemble, run this check on an\nindependent held-out split first.**\n\nThe public dataset `pilkwang/rsna-knee-weights` ships an extra file, `merge_gain.npz`, alongside\nthe well-known `oof.npz` — it contains pilkwang's own reproduced ensemble prediction (`ours`) and\na second, unlabeled prediction source (`imported`) evidently benchmarked when the file was built.\nMounted below unmodified, no scraping, no third-party claim beyond what the file itself contains."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import glob, hashlib\nimport numpy as np\nfrom scipy.stats import rankdata\n\nmerge_hits = sorted(glob.glob(\"/kaggle/input/**/merge_gain.npz\", recursive=True))\noof_hits = sorted(glob.glob(\"/kaggle/input/**/oof.npz\", recursive=True))\nassert merge_hits and oof_hits, \"pilkwang/rsna-knee-weights not mounted?\"\n\nd = np.load(merge_hits[0], allow_pickle=True)\nids, ours, imported = d[\"ids\"], d[\"ours\"].astype(np.float64), d[\"imported\"].astype(np.float64)\ny, gold = d[\"y\"], d[\"gold_mask\"]\ntargets = np.load(oof_hits[0], allow_pickle=True)[\"targets\"]\n\nprint(f\"studies: {len(ids)}, targets: {list(targets)}\")\nprint(f\"gold (expert-labelled) subset: {int(gold.sum())} studies\")"},{"cell_type":"markdown","metadata":{},"source":"## An independent split — not pilkwang's own\n\nEvery earlier check on this file (this campaign's own prior pass included) reused pilkwang's own\nreport-hash holdout, baked into the dataset the checkpoint author shipped. That's not a genuinely\nindependent test of a *second* source you're considering merging in. Here the studies are\nre-grouped by `sha256(salt + StudyInstanceUID)` parity — a fresh, fixed 2-fold split with no\ndependency on any file the checkpoint author built. Change `SALT` and every result below is a\ndifferent, equally valid independent split — the conclusion should not move if the finding is real."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"SALT = \"rsna-blend-check-2026-08-19\"\n\ndef group_of(uid):\n    h = hashlib.sha256((SALT + str(uid)).encode()).hexdigest()\n    return int(h[:8], 16) % 2\n\ngroups = np.array([group_of(u) for u in ids])\nprint(\"fold sizes:\", (groups == 0).sum(), (groups == 1).sum())\nprint(\"gold per fold:\", int(gold[groups == 0].sum()), int(gold[groups == 1].sum()))"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"def auc_binary(y_true, scores):\n    n_pos = y_true.sum()\n    n_neg = len(y_true) - n_pos\n    if n_pos == 0 or n_neg == 0:\n        return None\n    ranks = rankdata(scores)\n    return (ranks[y_true == 1].sum() - n_pos * (n_pos + 1) / 2) / (n_pos * n_neg)\n\ndef macro_auc(pred, y, mask):\n    aucs = [a for a in (auc_binary(y[mask, t], pred[mask, t]) for t in range(y.shape[1]))\n            if a is not None]\n    return float(np.mean(aucs)), len(aucs)\n\ndef rank_blend(a, b, w):\n    ra = np.apply_along_axis(rankdata, 0, a)\n    rb = np.apply_along_axis(rankdata, 0, b)\n    return w * ra + (1 - w) * rb\n\nprint(\"=== Full 4407 (weak/derived labels), by independent fold ===\")\nfor name, w in [(\"ours(w=1.0)\", 1.0), (\"blend(w=0.7)\", 0.7), (\"blend(w=0.5)\", 0.5),\n                 (\"blend(w=0.3)\", 0.3), (\"imported(w=0.0)\", 0.0)]:\n    pred = rank_blend(ours, imported, w)\n    a0, _ = macro_auc(pred, y, groups == 0)\n    a1, _ = macro_auc(pred, y, groups == 1)\n    aall, _ = macro_auc(pred, y, np.ones(len(y), dtype=bool))\n    print(f\"{name:16s} foldA={a0:.4f} foldB={a1:.4f} all={aall:.4f}\")\n\nprint()\nprint(\"=== Gold-58 only (expert labels) ===\")\nfor name, w in [(\"ours\", 1.0), (\"blend50\", 0.5), (\"imported\", 0.0)]:\n    pred = rank_blend(ours, imported, w)\n    m0, m1 = (groups == 0) & gold, (groups == 1) & gold\n    a0, _ = macro_auc(pred, y, m0)\n    a1, _ = macro_auc(pred, y, m1)\n    aall, _ = macro_auc(pred, y, gold)\n    print(f\"{name:10s} foldA(n={int(m0.sum())})={a0:.4f} foldB(n={int(m1.sum())})={a1:.4f} all(n={int(gold.sum())})={aall:.4f}\")"},{"cell_type":"markdown","metadata":{},"source":"Every blend ratio tested is worse than `ours` alone, on both folds, monotonically as more weight\nshifts to `imported`. That pattern — a clean, monotonic slide with blend weight, on every fold —\nis itself the signal worth checking for, independent of which two sources you're comparing."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"print(\"=== Per-target: does the second source (imported) beat ours on BOTH folds, for ANY target? ===\")\nany_win = False\nwins = 0\nfor t in range(len(targets)):\n    oa = auc_binary(y[groups == 0, t], ours[groups == 0, t])\n    ob = auc_binary(y[groups == 1, t], ours[groups == 1, t])\n    ia = auc_binary(y[groups == 0, t], imported[groups == 0, t])\n    ib = auc_binary(y[groups == 1, t], imported[groups == 1, t])\n    beats = ia > oa and ib > ob\n    any_win = any_win or beats\n    wins += int(beats)\n    print(f\"{targets[t]:18s} ours_A={oa:.4f} ours_B={ob:.4f} imp_A={ia:.4f} imp_B={ib:.4f}  imp_wins_both_folds={beats}\")\n\nprint()\nprint(f\"Targets where the second source wins both folds: {wins}/{len(targets)}\")"},{"cell_type":"markdown","metadata":{},"source":"## What this actually tells you\n\nZero of these 12 targets favor the second source on both folds — 24/24 fold-target comparisons go\nto `ours`, and every intermediate blend weight sits strictly between the two endpoints, worse than\nthe stronger source alone. That's a decisive negative on *this* pair, not a general claim that\nmerging prediction sources never helps.\n\n**If you're evaluating whether to merge a second prediction source into your own ensemble, run\nthis exact check first:** swap `imported` above for your own candidate array (same shape,\n`(n_studies, n_targets)`, same target order as printed in cell 1), keep your current best as\n`ours`, and look at the per-target-per-fold table. A clean monotonic slide with blend weight,\nlosing on every target on every fold — what's shown above — is a strong sign to skip the merge and\nsave the inference budget. A mixed pattern, winning on some targets or folds and losing on others,\nis the case actually worth investigating further; this notebook's own pair didn't produce one, but\nyours might."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11.0"}},"nbformat":4,"nbformat_minor":5}