{"cells":[{"cell_type":"markdown","metadata":{},"source":"# RSNA Blend-Gate Harness — 58 Golds, 12 Targets\n\n**The competition metric is macro ROC-AUC over 12 study-level labels, and only 58 of the\n4,407 training studies carry those labels.** Every blend decision in this competition —\n\"should I swap arm B in for Medial OA?\", \"does this TTA help?\", \"is my +0.002 real?\" — is\nultimately settled against those same 58 studies.\n\nThis notebook is the harness for that decision:\n\n`your two submissions → per-target AUC → paired bootstrap over studies → multiplicity\ncorrection → held-out cross-fit → SHIP / NOISE verdict`\n\n> Selecting the winner on the same 58 studies you measured it on is not a small bias. It\n> is the dominant term. This notebook measures how large it is for *your* setup, then\n> gives you a selection rule that survives it.\n\n**What you get**\n\n1. Per-target AUC for both of your prediction files, with paired bootstrap confidence\n   intervals resampled over **studies** (the correct unit — not rows, not targets).\n2. **A null calibration.** We simulate the exact per-target selection procedure that is\n   standard in this competition, under a true null where B is genuinely no better than A,\n   and report how often it manufactures a \"+0.002 macro win\" anyway. That number is your\n   false-discovery rate, and it is not small at n=58.\n3. **The fix:** a cross-fit selection that chooses on one half of the golds and scores on\n   the other, so the reported gain is one you can expect to keep on the private split.\n4. A fail-closed **SHIP / NOISE** verdict.\n\n## Use it on your own submissions\n\n1. **Copy & Edit.**\n2. Add your two submission files as inputs (or point at any two CSVs with\n   `StudyInstanceUID` plus the 12 target columns), then set `PRED_A_PATH` and\n   `PRED_B_PATH` in the config cell.\n3. Run all. With both paths left empty the notebook runs a **demo** on synthetic\n   predictors scored against the real 58 gold labels, so it executes end to end with no\n   edits and shows you exactly what the output looks like.\n\n**No GPU. No training. Runs in about a minute.** Nothing here touches the leaderboard —\nthe point is to stop spending submissions on decisions this can settle for free.\n"},{"cell_type":"markdown","metadata":{},"source":"## 0. Mission map\n\n| Step | Cell | What it settles |\n|---|---|---|\n| Load gold labels + canonical folds | 2 | the 58 labeled studies, and a study-grouped split |\n| Pre-flight assertions | 3 | that your files line up with the golds before any compute |\n| Metric | 4 | one pure AUC function used everywhere, no library drift |\n| Load / synthesize predictions | 5 | A vs B, aligned on StudyInstanceUID |\n| Paired bootstrap | 6 | per-target deltas with honest CIs over studies |\n| **Null calibration** | 7 | **how often your selection rule invents a win** |\n| Multiplicity correction | 8 | which per-target wins survive 12 simultaneous tests |\n| Cross-fit | 9 | the gain you can actually expect to keep |\n| Verdict | 10 | SHIP or NOISE, failing closed |\n\n**Deliberate choices, and why:**\n\n- **Bootstrap resamples studies, not rows.** A study is the independent unit; its 12\n  labels are correlated. Resampling anything smaller understates the interval.\n- **AUC is computed from ranks in NumPy**, not from a library, so the number cannot\n  silently change with a version bump — and every tie is handled identically in the\n  observed data and in every resample.\n- **The null is two *correlated* predictors of equal quality**, because that is what a\n  blend variant actually is. An independent-noise null would be far too easy to beat and\n  would understate the problem.\n- **Nothing uses the private test set, and nothing here needs the leaderboard.**\n"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"# ============================ CONFIG — the knobs =============================\n# Point these at YOUR two prediction files (submission format: StudyInstanceUID + the 12\n# target columns). Leave empty to run the demo on synthetic predictors.\nPRED_A_PATH = \"\"      # e.g. \"/kaggle/input/my-baseline/submission.csv\"\nPRED_B_PATH = \"\"      # e.g. \"/kaggle/input/my-blend/submission.csv\"\n\nSEED          = 2026\nN_BOOT        = 400    # paired bootstrap resamples over studies\nN_NULL        = 1000   # null-calibration replications\nSELECT_RULE   = 0.90   # per-target swap-in rule: bootstrap win fraction >= this\nCLAIMED_GAIN  = 0.002  # the macro gain size you want the null rate for\nALPHA         = 0.05   # family-wise error rate for the 12-target Holm correction\n\n# Demo-only knobs (ignored when you supply real files)\nDEMO_BASE_AUC = 0.87   # per-target AUC of the synthetic predictors\nDEMO_RHO      = 0.95   # correlation between A and B — blend variants are near-clones\nDEMO_TRUE_EDGE = 0.0   # true AUC advantage of B. 0.0 = null. Try 0.01 to see a real win.\n\nTARGETS = [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\", \"Medial OA\", \"Lateral OA\",\n           \"PF OA\", \"Effusion\", \"Synovitis\", \"Baker's\", \"Contusion\", \"Fracture\"]\nID = \"StudyInstanceUID\"\n"},{"cell_type":"markdown","metadata":{},"source":"## 1. Load the golds and the canonical folds\n\nOnly 58 of the 4,407 studies carry the 12 expert labels, and it is all-or-nothing — no study has a partial label row. Those 58 are the entire evaluation surface for any decision you make locally."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"# --------------------- load: gold labels + canonical folds ---------------------\nimport os\nimport numpy as np\nimport pandas as pd\n\ntry:              # `display` is a notebook builtin; the fallback lets this file also be\n    display       # exported to a plain .py and run as a script without edits.\nexcept NameError:\n    display = print\n\ndef first_existing(paths, what, fix, required=True):\n    for p in paths:\n        if os.path.exists(p):\n            return p\n    if required:\n        raise FileNotFoundError(f\"could not find {what}. Tried: {paths}\\nFix: {fix}\")\n    print(f\"NOTE: no {what} found. {fix}\")\n    return None\n\nTRAIN_CSV = first_existing(\n    [\"/kaggle/input/rsna-knee-abnormality-detection/train.csv\",\n     \"/kaggle/input/competitions/rsna-knee-abnormality-detection/train.csv\",\n     \"./rsna_data/train.csv\"],\n    \"the competition train.csv\",\n    \"add the competition as an input (Add Input -> Competitions -> RSNA Knee).\")\n\nFOLDS_CSV = first_existing(\n    [\"/kaggle/input/rsna-knee-2026-grouped-cv-folds/study_folds.csv\",\n     \"./rsna_data/study_folds.csv\"],\n    \"the grouped-fold pack\",\n    \"Add the dataset 'rsna-knee-2026-grouped-cv-folds' as an input for a study-grouped \"\n    \"cross-fit split; without it the harness falls back to a seeded random halving, \"\n    \"which is still unbiased but does not group series by study.\",\n    required=False)\n\ntrain = pd.read_csv(TRAIN_CSV)\ngold = train.dropna(subset=TARGETS, how=\"all\").reset_index(drop=True)\nprint(f\"train.csv        : {len(train):,} studies\")\nprint(f\"gold-labeled     : {len(gold)} studies ({100*len(gold)/len(train):.2f}%)\")\n\nfolds = None\nif FOLDS_CSV:\n    folds = pd.read_csv(FOLDS_CSV)\n    gold = gold.merge(folds[[ID, \"fold\"]], on=ID, how=\"left\")\n    print(f\"folds attached   : {gold['fold'].notna().sum()}/{len(gold)} golds carry a fold\")\n"},{"cell_type":"markdown","metadata":{},"source":"## 2. Pre-flight\n\nAssert the shape of the world before spending compute on it. The prevalence table is the thing to read twice: several targets have a minority class in the low teens, and an AUC on a dozen positives is a wide number no matter how it is computed."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"# ------------- pre-flight: assert before you compute anything expensive -------------\nassert gold[ID].is_unique, \"duplicate StudyInstanceUID among the gold studies\"\nmissing = [t for t in TARGETS if t not in gold.columns]\nassert not missing, f\"train.csv is missing target columns: {missing}\"\n\nY = gold[TARGETS].to_numpy(dtype=float)\nassert np.isfinite(Y).all(), \"gold labels contain NaN — the all-or-nothing rule was violated\"\nassert set(np.unique(Y)) <= {0.0, 1.0}, f\"labels are not binary: {set(np.unique(Y))}\"\n\nprev = pd.DataFrame({\n    \"target\": TARGETS,\n    \"positives\": Y.sum(axis=0).astype(int),\n    \"negatives\": (len(Y) - Y.sum(axis=0)).astype(int),\n})\nprev[\"prevalence\"] = (prev.positives / len(Y)).round(3)\nprev[\"min_class_n\"] = prev[[\"positives\", \"negatives\"]].min(axis=1)\ndisplay(prev)\n\nprint(f\"\\nsmallest minority class: {prev.min_class_n.min()} studies \"\n      f\"({prev.loc[prev.min_class_n.idxmin(), 'target']})\")\nprint(\"Every AUC below is computed on 58 studies. Read the intervals, not the point \"\n      \"estimates.\")\n"},{"cell_type":"markdown","metadata":{},"source":"## 3. One metric function, used everywhere\n\nThe competition scores macro ROC-AUC across the 12 labels. Computing it from ranks in NumPy keeps the observed value and every bootstrap replicate on identical tie-handling, and removes a library version from the set of things that can change your answer."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"# ------------------------- the metric, as one pure function -------------------------\n# Mann-Whitney form of ROC-AUC computed from ranks. Defining it here rather than importing\n# a scorer keeps the observed value and every bootstrap replicate on identical\n# tie-handling, and removes one library version from the set of things that can move your\n# answer between runs.\nfrom scipy.stats import rankdata\n\ndef auc(y, s):\n    \"\"\"ROC-AUC for one target. Returns nan when a class is absent (undefined, not 0.5).\"\"\"\n    n1 = y.sum()\n    n0 = len(y) - n1\n    if n1 == 0 or n0 == 0:\n        return np.nan\n    r = rankdata(s)\n    return (r[y == 1].sum() - n1 * (n1 + 1) / 2) / (n1 * n0)\n\ndef auc_rows(y_rows, s_rows):\n    \"\"\"Vectorised AUC over many resamples at once. y_rows/s_rows are (B, n).\n\n    Bootstrap resampling draws the same study more than once, so ties are real and are\n    given average ranks — the same convention `auc` uses on the observed data.\n    \"\"\"\n    ranks = rankdata(s_rows, axis=1)\n    n1 = y_rows.sum(axis=1)\n    n0 = y_rows.shape[1] - n1\n    pos_rank_sum = (ranks * y_rows).sum(axis=1)\n    out = (pos_rank_sum - n1 * (n1 + 1) / 2) / np.maximum(n1 * n0, 1)\n    out[(n1 == 0) | (n0 == 0)] = np.nan\n    return out\n\ndef macro(y, S):\n    \"\"\"Macro AUC across the 12 targets, skipping any that are undefined.\"\"\"\n    vals = [auc(y[:, t], S[:, t]) for t in range(y.shape[1])]\n    return float(np.nanmean(vals)), np.array(vals)\n\nprint(\"metric ready: macro ROC-AUC over\", len(TARGETS), \"targets\")\n"},{"cell_type":"markdown","metadata":{},"source":"## 4. Your two prediction sets\n\nA is your incumbent, B is the candidate. They are aligned on `StudyInstanceUID` and reduced to the gold studies. If you supply a raw *submission* file it will not contain the golds — score your model on the training golds and export those rows."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"# --------------------- predictions: yours, or the demo pair ---------------------\nrng = np.random.default_rng(SEED)\n\ndef load_preds(path, name):\n    df = pd.read_csv(path)\n    assert ID in df.columns, f\"{name}: no {ID} column in {path}\"\n    miss = [t for t in TARGETS if t not in df.columns]\n    assert not miss, f\"{name}: missing target columns {miss}\"\n    m = gold[[ID]].merge(df[[ID] + TARGETS], on=ID, how=\"left\")\n    cov = m[TARGETS].notna().all(axis=1).sum()\n    assert cov > 0, (f\"{name}: none of the 58 gold studies appear in {path}. A public \"\n                     f\"submission file only covers the test set — score your model on \"\n                     f\"the gold studies and export those rows instead.\")\n    print(f\"{name}: {cov}/{len(gold)} gold studies covered\")\n    return m[TARGETS].to_numpy(dtype=float)\n\ndef demo_pair():\n    \"\"\"Two correlated predictors of equal quality — what a blend variant really is.\"\"\"\n    from scipy.stats import norm\n    mu = np.sqrt(2) * norm.ppf(DEMO_BASE_AUC)\n    mu_b = np.sqrt(2) * norm.ppf(min(0.999, DEMO_BASE_AUC + DEMO_TRUE_EDGE))\n    A = np.empty_like(Y); B = np.empty_like(Y)\n    for t in range(Y.shape[1]):\n        za = rng.standard_normal(len(Y))\n        zb = DEMO_RHO * za + np.sqrt(1 - DEMO_RHO ** 2) * rng.standard_normal(len(Y))\n        A[:, t] = za + mu * Y[:, t]\n        B[:, t] = zb + mu_b * Y[:, t]\n    return A, B\n\nif PRED_A_PATH and PRED_B_PATH:\n    SA, SB = load_preds(PRED_A_PATH, \"A\"), load_preds(PRED_B_PATH, \"B\")\n    MODE = \"your files\"\nelse:\n    SA, SB = demo_pair()\n    MODE = (f\"DEMO (synthetic predictors, base AUC {DEMO_BASE_AUC}, rho {DEMO_RHO}, \"\n            f\"true edge {DEMO_TRUE_EDGE:+.3f}) scored on the REAL 58 gold labels\")\n    print(\"no prediction files supplied — running the demo.\")\n    print(\"Set PRED_A_PATH / PRED_B_PATH in the config cell to test your own.\")\n\nmacro_A, per_A = macro(Y, SA)\nmacro_B, per_B = macro(Y, SB)\nprint(f\"\\nmode      : {MODE}\")\nprint(f\"macro A   : {macro_A:.5f}\")\nprint(f\"macro B   : {macro_B:.5f}\")\nprint(f\"raw delta : {macro_B - macro_A:+.5f}   <- the number people quote\")\n"},{"cell_type":"markdown","metadata":{},"source":"## 5. Paired bootstrap over studies\n\nThe interval that matters is on the **delta**, not on either AUC alone, and both predictors must see the same resample to get it. The `SWAP` flags below apply the selection rule this competition's public pipelines use: swap B in for a target when it wins a high fraction of paired resamples."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"# ------------- paired bootstrap over STUDIES (the independent unit) -------------\n# A study's 12 labels are correlated, so the study is what gets resampled. Both\n# predictors see the identical resample: the comparison is paired, which is what makes a\n# delta interval this narrow legitimate at n=58.\nboot_idx = rng.integers(0, len(Y), size=(N_BOOT, len(Y)))\n\nrows = []\ndelta_boot = np.zeros((N_BOOT, len(TARGETS)))\nfor t, name in enumerate(TARGETS):\n    yb = Y[boot_idx, t]\n    da = auc_rows(yb, SA[boot_idx, t])\n    db = auc_rows(yb, SB[boot_idx, t])\n    d = db - da\n    delta_boot[:, t] = np.where(np.isnan(d), 0.0, d)\n    ok = ~np.isnan(d)\n    lo, hi = np.percentile(d[ok], [2.5, 97.5]) if ok.any() else (np.nan, np.nan)\n    winfrac = float((d[ok] > 0).mean()) if ok.any() else np.nan\n    rows.append(dict(target=name, auc_A=round(per_A[t], 4), auc_B=round(per_B[t], 4),\n                     delta=round(per_B[t] - per_A[t], 4), ci_lo=round(lo, 4),\n                     ci_hi=round(hi, 4), boot_win_frac=round(winfrac, 3),\n                     selected=winfrac >= SELECT_RULE))\n    flag = \"SWAP\" if winfrac >= SELECT_RULE else \"    \"\n    print(f\"  {flag}  {name:<18} A {per_A[t]:.4f}  B {per_B[t]:.4f}  \"\n          f\"delta {per_B[t]-per_A[t]:+.4f}  [{lo:+.4f},{hi:+.4f}]  win {winfrac:.2f}\")\n\nper_target = pd.DataFrame(rows)\nn_sel = int(per_target.selected.sum())\nsel_macro = float(np.where(per_target.selected, per_B, per_A).mean())\nprint(f\"\\nrule: swap in B where bootstrap win fraction >= {SELECT_RULE}\")\nprint(f\"targets swapped     : {n_sel}/12\")\nprint(f\"macro after swaps   : {sel_macro:.5f}   ({sel_macro - macro_A:+.5f} vs A)\")\nprint(\"\\nThat gain is measured on the same 58 studies the swaps were chosen on. \"\n      \"The next cell prices that.\")\n"},{"cell_type":"markdown","metadata":{},"source":"## 6. What that rule does when there is nothing to find\n\nNow the load-bearing cell. We hand the identical procedure a pair of predictors that are *equally good by construction* — same discriminative power, correlated like two members of one ensemble family — and let it select. Any macro gain it reports is manufactured, because there is nothing there to find.\n\n> Run this before believing any per-target gate, including your own."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"# ================== THE NULL CALIBRATION — the number that matters ==================\n# Question: if B were genuinely no better than A, how often would this exact procedure\n# still report a macro gain of CLAIMED_GAIN or more?\n#\n# The null pair is two predictors with the SAME discriminative power, correlated at\n# DEMO_RHO, evaluated on the real gold labels and their real prevalences. That is a blend\n# variant: a near-clone, not an independent model. Selection then runs unchanged.\nfrom scipy.stats import norm\n\ndef null_replication(rs):\n    mu = np.sqrt(2) * norm.ppf(DEMO_BASE_AUC)\n    gains = np.empty(len(TARGETS))\n    sel = np.zeros(len(TARGETS), dtype=bool)\n    a_auc = np.empty(len(TARGETS)); b_auc = np.empty(len(TARGETS))\n    bidx = rs.integers(0, len(Y), size=(N_BOOT, len(Y)))\n    for t in range(len(TARGETS)):\n        za = rs.standard_normal(len(Y))\n        zb = DEMO_RHO * za + np.sqrt(1 - DEMO_RHO ** 2) * rs.standard_normal(len(Y))\n        sa = za + mu * Y[:, t]\n        sb = zb + mu * Y[:, t]          # identical mu => TRUE NULL, no real edge\n        a_auc[t] = auc(Y[:, t], sa); b_auc[t] = auc(Y[:, t], sb)\n        yb = Y[bidx, t]\n        d = auc_rows(yb, sb[bidx]) - auc_rows(yb, sa[bidx])\n        ok = ~np.isnan(d)\n        sel[t] = (d[ok] > 0).mean() >= SELECT_RULE if ok.any() else False\n    base = np.nanmean(a_auc)\n    kept = np.nanmean(np.where(sel, b_auc, a_auc))\n    return kept - base, int(sel.sum())\n\nrs = np.random.default_rng(SEED + 1)\nnull_gain = np.empty(N_NULL); null_nsel = np.empty(N_NULL, dtype=int)\nfor i in range(N_NULL):\n    null_gain[i], null_nsel[i] = null_replication(rs)\n    if (i + 1) % max(1, N_NULL // 5) == 0:\n        print(f\"  null replications {i+1}/{N_NULL}  \"\n              f\"running mean gain {null_gain[:i+1].mean():+.5f}\")\n\np_claim = float((null_gain >= CLAIMED_GAIN).mean())\nprint(f\"\\nUnder a TRUE NULL (B is exactly as good as A), this procedure reports:\")\nprint(f\"  mean macro 'gain'          : {null_gain.mean():+.5f}\")\nprint(f\"  median                     : {np.median(null_gain):+.5f}\")\nprint(f\"  95th percentile            : {np.percentile(null_gain, 95):+.5f}\")\nprint(f\"  mean targets swapped       : {null_nsel.mean():.2f} of 12\")\nprint(f\"  P(gain >= {CLAIMED_GAIN:+.3f} | null)   : {p_claim:.1%}   <- your false-discovery rate\")\nprint(f\"\\nYour own selected gain was {sel_macro - macro_A:+.5f}; the null produces at least \"\n      f\"that {float((null_gain >= sel_macro - macro_A).mean()):.1%} of the time.\")\n"},{"cell_type":"markdown","metadata":{},"source":"## 7. Twelve tests, not one\n\nA per-target rule is twelve simultaneous decisions. Holm controls the family-wise error rate and makes no independence assumption, which matters here because every test rides on the same 58 studies."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"# ---------------- 12 simultaneous tests: Holm correction ----------------\n# Twelve targets means twelve chances to be surprised. Holm controls the family-wise error\n# rate without assuming the tests are independent (they are not — the same studies).\np_raw = np.array([2 * min((delta_boot[:, t] <= 0).mean(), (delta_boot[:, t] >= 0).mean())\n                  for t in range(len(TARGETS))])\np_raw = np.clip(p_raw, 1.0 / N_BOOT, 1.0)\norder = np.argsort(p_raw)\nholm = np.empty(len(TARGETS))\nrunning = 0.0\nfor rank, t in enumerate(order):\n    running = max(running, (len(TARGETS) - rank) * p_raw[t])\n    holm[t] = min(1.0, running)\n\nper_target[\"p_boot\"] = p_raw.round(4)\nper_target[\"p_holm\"] = holm.round(4)\nper_target[\"survives_holm\"] = per_target.p_holm < ALPHA\ndisplay(per_target[[\"target\", \"delta\", \"ci_lo\", \"ci_hi\", \"boot_win_frac\",\n                    \"selected\", \"p_holm\", \"survives_holm\"]])\nn_holm = int(per_target.survives_holm.sum())\nprint(f\"targets passing the swap rule            : {n_sel}/12\")\nprint(f\"targets surviving Holm at alpha={ALPHA}      : {n_holm}/12\")\nif n_sel > n_holm:\n    print(f\"-> {n_sel - n_holm} of your swaps do not survive being one of twelve tests.\")\n"},{"cell_type":"markdown","metadata":{},"source":"## 8. The fix, not just the warning\n\nSelection bias is not an argument for doing nothing — it is an argument for choosing on data that does not also score you. This is the cheapest honest version: split the golds, select on one half, report the other."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"# ------------- the fix: choose on one half of the golds, score on the other -------------\n# Selection and evaluation on the same 58 studies is the winner's curse, and the null cell\n# just priced it. Cross-fit removes it: pick the swaps on half the golds, then report the\n# gain on the half that had no vote. Halves are study-grouped where a fold file was found.\nif folds is not None and gold[\"fold\"].notna().all():\n    half = (gold[\"fold\"].to_numpy() % 2).astype(int)\n    split_kind = \"study-grouped folds (from the mounted fold pack)\"\nelse:\n    half = rng.integers(0, 2, size=len(Y))\n    split_kind = \"seeded random halving (no fold file mounted)\"\n\ncross_gains = []\nfor hold in (0, 1):\n    tr, te = half != hold, half == hold\n    sel_tr = np.zeros(len(TARGETS), dtype=bool)\n    bidx = rng.integers(0, tr.sum(), size=(N_BOOT, int(tr.sum())))\n    for t in range(len(TARGETS)):\n        ytr = Y[tr, t]\n        d = auc_rows(ytr[bidx], SB[tr, t][bidx]) - auc_rows(ytr[bidx], SA[tr, t][bidx])\n        ok = ~np.isnan(d)\n        sel_tr[t] = (d[ok] > 0).mean() >= SELECT_RULE if ok.any() else False\n    a_te = np.array([auc(Y[te, t], SA[te, t]) for t in range(len(TARGETS))])\n    b_te = np.array([auc(Y[te, t], SB[te, t]) for t in range(len(TARGETS))])\n    g = np.nanmean(np.where(sel_tr, b_te, a_te)) - np.nanmean(a_te)\n    cross_gains.append(g)\n    print(f\"  select on half {1-hold}, score on half {hold}: \"\n          f\"{sel_tr.sum()} swaps, held-out gain {g:+.5f}\")\n\ncross_gain = float(np.mean(cross_gains))\nprint(f\"\\nsplit               : {split_kind}\")\nprint(f\"in-sample gain      : {sel_macro - macro_A:+.5f}\")\nprint(f\"CROSS-FIT gain      : {cross_gain:+.5f}   <- the one you can expect to keep\")\nprint(f\"selection premium   : {(sel_macro - macro_A) - cross_gain:+.5f} of the in-sample \"\n      f\"number was the act of choosing\")\n"},{"cell_type":"markdown","metadata":{},"source":"## 9. Two pictures\n\nLeft: how wide the per-target intervals really are at n=58. Right: the distribution of gains this procedure reports from a true null, with your own in-sample and cross-fit numbers marked on it."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"# ------------------------------- diagnostics -------------------------------\nimport matplotlib.pyplot as plt\n\nfig, axes = plt.subplots(1, 2, figsize=(13, 4.6))\n\nax = axes[0]\no = np.argsort(per_target.delta.to_numpy())\nyy = np.arange(len(TARGETS))\nax.errorbar(per_target.delta.to_numpy()[o], yy,\n            xerr=[per_target.delta.to_numpy()[o] - per_target.ci_lo.to_numpy()[o],\n                  per_target.ci_hi.to_numpy()[o] - per_target.delta.to_numpy()[o]],\n            fmt=\"o\", capsize=3, color=\"#31708E\")\nax.axvline(0, color=\"#999\", lw=1)\nax.set_yticks(yy); ax.set_yticklabels([TARGETS[i] for i in o], fontsize=9)\nax.set_xlabel(\"AUC delta (B - A), paired bootstrap 95% CI\")\nax.set_title(\"Per-target delta at n=58\\nintervals this wide are the whole story\", fontsize=11)\n\nax = axes[1]\nax.hist(null_gain, bins=40, color=\"#C6D8E4\", edgecolor=\"#31708E\")\nax.axvline(CLAIMED_GAIN, color=\"#B03A2E\", lw=2,\n           label=f\"claimed gain {CLAIMED_GAIN:+.3f}  (null exceeds it {p_claim:.0%})\")\nax.axvline(sel_macro - macro_A, color=\"#1E8449\", lw=2, ls=\"--\",\n           label=f\"your selected gain {sel_macro - macro_A:+.4f}\")\nax.axvline(cross_gain, color=\"#7D3C98\", lw=2, ls=\":\",\n           label=f\"your cross-fit gain {cross_gain:+.4f}\")\nax.set_xlabel(\"macro gain reported under a TRUE NULL\")\nax.set_ylabel(\"replications\")\nax.set_title(\"What the selection procedure invents from nothing\", fontsize=11)\nax.legend(fontsize=8)\n\nplt.tight_layout(); plt.show()\n"},{"cell_type":"markdown","metadata":{},"source":"## 10. Verdict\n\nFails closed: a delta is only called real when it survives cross-fit, survives Holm, and beats what the null manufactures."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"# ================================ THE VERDICT ================================\n# Fails closed. A gain is only called real when it survives selection bias (cross-fit),\n# multiplicity (Holm), and the null calibration.\n# Two hard conditions, both calibrated against THIS null rather than a rule of thumb:\n#   1. the gain transfers to studies that had no vote in choosing the swaps, and\n#   2. the in-sample gain is bigger than selection alone manufactures (null 95th pct).\n# Holm is reported but NOT required: at n=58 across 12 tests it is a very high bar that a\n# genuine, useful effect can legitimately miss, and a gate that can never say SHIP is not\n# a gate — it is a mood.\nnull_p95 = float(np.percentile(null_gain, 95))\nin_gain = sel_macro - macro_A\nreasons = []\nif cross_gain <= 0:\n    reasons.append(f\"cross-fit gain is {cross_gain:+.5f} — the swaps do not transfer to \"\n                   f\"studies that had no vote in choosing them\")\nif in_gain <= null_p95:\n    reasons.append(f\"in-sample gain {in_gain:+.5f} does not clear what pure selection \"\n                   f\"manufactures at this n (null 95th pct {null_p95:+.5f})\")\n\nprint(f\"mode            : {MODE}\")\nprint(f\"macro A / B     : {macro_A:.5f} / {macro_B:.5f}\")\nprint(f\"in-sample gain  : {in_gain:+.5f}   (null 95th pct {null_p95:+.5f})\")\nprint(f\"cross-fit gain  : {cross_gain:+.5f}\")\nprint(f\"Holm survivors  : {n_holm}/12  (informational — a high bar at n=58)\")\nprint(f\"\\nPROCEDURE NOTE: at this sample size, the per-target swap rule reports a \"\n      f\"{CLAIMED_GAIN:+.3f} macro gain {p_claim:.0%} of the time when there is nothing \"\n      f\"there. That is a property of n=58 and 12 targets, not of your models.\")\nprint()\nif not reasons:\n    print(\"VERDICT: SHIP — the gain transfers to held-out golds AND exceeds what \"\n          \"selection alone invents. Spend the submission.\")\nelse:\n    print(\"VERDICT: NOISE — do not spend a submission on this delta. Because:\")\n    for r in reasons:\n        print(f\"  - {r}\")\n    print(\"\\nThis is not a claim that B is worse than A. It is that 58 studies cannot \"\n          \"tell them apart here, so the leaderboard move would be a coin flip you paid \"\n          \"for with a submission.\")\n\nout = per_target.copy()\nout[\"macro_A\"] = macro_A; out[\"macro_B\"] = macro_B\nout[\"in_sample_gain\"] = in_gain; out[\"cross_fit_gain\"] = cross_gain\nout[\"null_p95\"] = null_p95; out[\"null_fdr_at_claimed\"] = p_claim\nout.to_csv(\"blend_gate_report.csv\", index=False)\npd.DataFrame({\"null_gain\": null_gain, \"n_selected\": null_nsel}).to_csv(\n    \"null_calibration.csv\", index=False)\nprint(\"\\nwrote blend_gate_report.csv, null_calibration.csv\")\n"},{"cell_type":"markdown","metadata":{},"source":"---\n## What this does and does not say\n\n**It says:** at n=58 with 12 targets, per-target selection on the same studies you measure\non carries a selection premium large enough to manufacture the size of gain this\ncompetition trades in. The null calibration prices that premium for your own configuration,\nand the cross-fit removes it.\n\n**It does not say** that any particular public notebook's result is wrong. Selection bias is\na property of a *procedure at a sample size*, not an accusation about a person. Several\nstrong public pipelines gate their swaps carefully; the point is that careful-looking gates\n(bootstrap win-fraction thresholds) are exactly the ones the null cell shows to be\npermeable at this n. Run it on your own numbers and see where you land.\n\n**Known limits, stated:**\n\n- The null pair is Gaussian-latent with a single correlation knob. Real prediction pairs\n  have structured disagreement. `DEMO_RHO` and `DEMO_BASE_AUC` are exposed precisely so\n  you can move them to your own measured values — a higher correlation makes the null\n  *harder* to beat, so leaving rho low would flatter the procedure, not damn it.\n- Cross-fit at n=58 halves an already small sample; the held-out gain has a wide interval\n  of its own. Treating its sign as informative and its magnitude as approximate is the\n  honest reading.\n- AUC is undefined for a target with no positives in a resample; those are dropped rather\n  than imputed at 0.5, which would bias every interval toward zero.\n- The gold-58 remain the only labeled studies. Nothing here manufactures more of them —\n  and that, not architecture, is the ceiling this competition is actually up against.\n\n## Extending it\n\n- Swap `SELECT_RULE` to see how much of your gain is the threshold rather than the signal.\n- Set `DEMO_TRUE_EDGE = 0.01` to watch the harness correctly identify a *real* effect —\n  a gate that never says SHIP is useless, so check it can.\n- Point `PRED_A_PATH`/`PRED_B_PATH` at your own out-of-fold predictions on the golds\n  rather than submissions; the harness does not care where the numbers came from.\n- The fold pack it mounts (`rsna-knee-2026-grouped-cv-folds`) keeps every series of a\n  study inside one fold, so per-series models can be scored here without a scanner-session\n  leak.\n\n> **The rule worth keeping:** choose on data that does not get to score you, and price\n> your selection rule against a null before you trust the number it hands back.\n"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11"}},"nbformat":4,"nbformat_minor":5}