{"cells":[{"cell_type":"markdown","id":"cell00","metadata":{},"source":"# Your second submission is probably worth zero\n\nAs of 2026-08-26 this competition has **2,463 teams**. **228 of them display exactly 0.936**, and\nthat block spans ranks **74 through 301**. The bronze cutoff, around rank 246, sits *inside* the\nblock. Two teams whose displayed scores are identical to three decimals differ by less than 0.001\nof macro-AUC, which - as section 1 works out - is a fraction of one private standard error at\nevery test size considered there. Ordering inside that block is noise.\n\nThat is the setting for a question almost nobody prices. Kaggle scores your **best selected\nsubmission** on the private split, so a final pair pays `E[max(X, Y)]`. Most teams let the site\nauto-select their two best public scores, which are usually two siblings out of one pipeline. This\nnotebook works out what that second slot is actually worth.\n\nThree things fall out, and the third is the one I did not expect.\n\n1. A second entry more than about **two private standard errors below your best** stops paying,\n   and that threshold is derived rather than guessed. Two standard errors is where the gain\n   crosses the inert line at `rho = 0` - the best case decorrelation can offer - and every\n   `rho > 0` crosses sooner. Push the gap to four standard errors and the gain rounds to zero.\n   Decorrelation without proximity buys nothing.\n2. The correlation term has to be **measured, not assumed**. On a sibling pair out of one of our\n   own pipeline families the measured Spearman was **0.9966 to 0.9996**. Intuition had suggested\n   something near 0.5, and that substitution is the difference between a `LIVE` verdict and an\n   `INERT` one.\n3. Prediction correlation **cannot tell a complementary model from a merely weak one** - noise is\n   uncorrelated with everything. A per-label win count can. Worked example below: a model at\n   rho = 0.231 against our strongest entry, which looks beautifully decorrelated, wins **0 of 12\n   labels**, and a 3,000-shuffle permutation null says 0 wins is exactly what no signal predicts.\n\nEverything here is expressed in **standard errors**, not in points of macro-AUC. That is not a\nstylistic choice - section 1 explains why it is the only honest unit available, and it is what\nmakes the conclusions independent of a quantity nobody outside the host can measure.\n\nNothing here needs a GPU, the internet, or a mounted model. You can re-run it against your own two\nsubmissions with the helper in section 4."},{"cell_type":"markdown","id":"cell01","metadata":{},"source":"## Correction to the first version of this notebook\n\nAn earlier version quoted a single private-split standard error - **0.00305**, with **0.00466** for\nthe public split - and converted its thresholds into points of macro-AUC on that basis, for example\n\"two standard errors is 0.0061 of score\". **Those three figures are withdrawn.**\n\nThey rested on a test set of 1,322 studies. I took that number from a log line in a public notebook\nreading `sizing for 4407 train + 1322 test studies`. It is not a measurement. It is a\nmemory-planning heuristic - `int(0.3 * n_train)` applied to the 4,407 training rows - and it was\nprinted for cache sizing, not reported as a fact about the test split. The hidden test size is not\npublished. I read a printed number as a measured one, which is the same mistake as assuming a\ncorrelation instead of measuring it, and section 4 is about exactly that.\n\n**What is verified and unchanged:** the split is **30% public / 70% private**, which the\ncompetition's own metadata reports as `leaderboardPercentage: 30`.\n\n**What the correction does not touch:** every threshold in this notebook is denominated in standard\nerrors, and expected gain divided by `sigma` depends only on the mean gap *measured in sigmas* and\non `rho` - the `sigma` cancels. The next-to-last cell of section 2 checks that claim numerically\nrather than asserting it. So the break-even correlation, the inert crossings, the measured\ncorrelations, the per-label result and its permutation null are all exactly as they were. What\nchanges is that `sigma` is now shown as a **range** across plausible test sizes, and any conversion\ninto points of macro-AUC is quoted as a range with it.\n\nIf you took the 0.0061 figure from the earlier version and used it as your proximity window, the\nreplacement is in section 1: the window is **2.00 sigma**, which is somewhere between 0.003 and\n0.009 of macro-AUC depending on a test size neither of us knows.\n\nThe argument is in better shape for this. It now holds by construction rather than by assumption."},{"cell_type":"code","id":"cell02","execution_count":null,"metadata":{},"outputs":[],"source":"import math\nimport os\nimport tempfile\nfrom statistics import NormalDist\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\nND = NormalDist()\n\n# ---- board, read 2026-08-26 -------------------------------------------------\nN_TEAMS = 2463          # teams on the public leaderboard\nTIE_BLOCK = 228         # teams displaying exactly 0.936\nTIE_LO, TIE_HI = 74, 301\nBRONZE_RANK = 246       # cutoff rank, inside that block\nCEILING = 0.936         # the displayed score at the top of the block\n\n# ---- test-set geometry ------------------------------------------------------\n# VERIFIED: the competition metadata reports leaderboardPercentage = 30, so the split is\n# 30% public / 70% private.\n# NOT KNOWN: the absolute size of the hidden test set. It is not published, and it cannot be\n# recovered from outside a rerun. Rather than pick one, we carry a range of plausible sizes and\n# report every score-unit quantity across all of them.\nPUBLIC_FRAC = 0.30\nPRIVATE_FRAC = 1.0 - PUBLIC_FRAC\nTEST_SIZES = [600, 1000, 1322, 2000, 3000, 5000]\n\n# ---- the twelve study-level labels, and their positive counts on the 58\n#      gold-labelled studies we hold. Used only as a prevalence estimate.\nLABELS = [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\", \"Medial OA\",\n          \"Lateral OA\", \"PF OA\", \"Effusion\", \"Synovitis\", \"Baker's\",\n          \"Contusion\", \"Fracture\"]\nN_GOLD = 58\nPOS_GOLD = [24, 9, 26, 23, 15, 11, 21, 35, 27, 12, 19, 18]\nPREV = [p / N_GOLD for p in POS_GOLD]\n\nprint(f\"teams {N_TEAMS:,} | {TIE_BLOCK} display {CEILING:.3f}, ranks {TIE_LO}-{TIE_HI}\"\n      f\" | bronze cutoff ~{BRONZE_RANK}, inside the block\")\nprint(f\"split {PUBLIC_FRAC:.0%} public / {PRIVATE_FRAC:.0%} private  (verified)\")\nprint(f\"test-set size    NOT PUBLISHED - carried as a range: \"\n      f\"{min(TEST_SIZES):,} to {max(TEST_SIZES):,} studies\")\nprint(f\"metric           macro ROC-AUC over {len(LABELS)} study-level labels\")\nrarest = LABELS[POS_GOLD.index(min(POS_GOLD))]\ncommonest = LABELS[POS_GOLD.index(max(POS_GOLD))]\nprint(f\"prevalence on the 58 gold studies: {min(PREV):.3f} ({rarest})\"\n      f\" to {max(PREV):.3f} ({commonest})\")"},{"cell_type":"markdown","id":"cell03","metadata":{},"source":"## 1. Sigma is a range, not a constant\n\nBefore any selection maths we need a scale. The metric is a macro ROC-AUC over twelve binary\nstudy-level labels, so its sampling standard error follows from Hanley and McNeil (1982) per label,\naveraged. That formula needs the number of scored samples, and there is the problem: **the hidden\ntest size is not published.** The split is 30/70, but 30% of what is not something we can read off\nthe site.\n\nSo the cell below computes `sigma` across a range of plausible test sizes instead of choosing one.\nThree further caveats, stated before the numbers rather than after them.\n\n- Averaging the per-label variances as if the labels were **independent** gives a **lower bound**.\n  Real pathologies co-occur on the same knee. An exchangeable label correlation `rho_L` inflates a\n  `k`-label mean by `sqrt(1 + (k-1) * rho_L)`, and the cell prints that sensitivity too.\n- Hanley's standard error is **unpaired**. Comparing two models on the *same* test studies is a\n  paired comparison whose standard error is smaller. The two corrections have opposite signs.\n- Prevalence is estimated from the 58 gold-labelled studies we hold, not from the test set.\n\nGiven all that, the honest position is that `sigma` is known to within a factor of about three, and\nthe rest of this notebook is written so that the factor of three does not matter."},{"cell_type":"code","id":"cell04","execution_count":null,"metadata":{},"outputs":[],"source":"def hanley_se(auc, n_pos, n_neg):\n    \"\"\"Hanley-McNeil standard error of a single ROC-AUC.\"\"\"\n    if n_pos < 1 or n_neg < 1:\n        return float(\"nan\")\n    q1 = auc / (2 - auc)\n    q2 = 2 * auc * auc / (1 + auc)\n    var = (auc * (1 - auc)\n           + (n_pos - 1) * (q1 - auc * auc)\n           + (n_neg - 1) * (q2 - auc * auc)) / (n_pos * n_neg)\n    return math.sqrt(max(var, 0.0))\n\n\ndef macro_auc_se(auc, n, prevalences, label_rho=0.0):\n    \"\"\"SE of a macro-averaged AUC over len(prevalences) labels on n samples.\n\n    label_rho = 0 assumes independent labels and is a LOWER BOUND on the true SE.\n    \"\"\"\n    ses = []\n    for p in prevalences:\n        n_pos = max(1, round(n * p))\n        ses.append(hanley_se(auc, n_pos, n - n_pos))\n    k = len(ses)\n    indep = math.sqrt(sum(s * s for s in ses)) / k\n    return indep * math.sqrt(1 + (k - 1) * label_rho)\n\n\nrows = []\nfor n_test in TEST_SIZES:\n    n_priv = round(n_test * PRIVATE_FRAC)\n    n_pub = n_test - n_priv\n    rows.append({\n        \"test studies\": n_test,\n        \"private\": n_priv,\n        \"public\": n_pub,\n        \"sigma private\": round(macro_auc_se(CEILING, n_priv, PREV, 0.0), 5),\n        \"sigma public\": round(macro_auc_se(CEILING, n_pub, PREV, 0.0), 5),\n        \"tie of 0.001, in SE\": round(0.001 / macro_auc_se(CEILING, n_priv, PREV, 0.0), 2),\n    })\nsigma_table = pd.DataFrame(rows)\nprint(\"private-split macro-AUC SE, independent labels (a LOWER bound), by assumed test size:\")\nprint(sigma_table.to_string(index=False))\nprint()\n\nSIGMAS = [macro_auc_se(CEILING, round(n * PRIVATE_FRAC), PREV, 0.0) for n in TEST_SIZES]\nSIGMA_LO, SIGMA_HI = min(SIGMAS), max(SIGMAS)\n# A single mid-range value, used ONLY where a worked example needs one number. Nothing in the\n# argument depends on it, and every threshold below is quoted in sigmas.\nSIGMA_MID = sorted(SIGMAS)[len(SIGMAS) // 2]  # a midpoint of the sweep, not a privileged value\nprint(f\"sigma spans {SIGMA_LO:.5f} to {SIGMA_HI:.5f} - a factor of {SIGMA_HI / SIGMA_LO:.2f}.\")\nprint(f\"no single sigma is adopted anywhere below: every result is reported either in units of\")\nprint(\"  threshold is quoted in sigmas so that the choice does not matter.\")\nprint()\nprint(\"label-correlation sensitivity, on top of all of the above:\")\nfor lr in (0.0, 0.1, 0.25, 0.5):\n    lo = macro_auc_se(CEILING, round(max(TEST_SIZES) * PRIVATE_FRAC), PREV, lr)\n    hi = macro_auc_se(CEILING, round(min(TEST_SIZES) * PRIVATE_FRAC), PREV, lr)\n    print(f\"  rho_L = {lr:<5} -> sigma {lo:.5f} to {hi:.5f}\")\nprint()\ntie_lo = 0.001 / SIGMA_HI\ntie_hi = 0.001 / SIGMA_LO\nprint(f\"Two entries displaying the same 3-decimal score differ by less than 0.001, which is\")\nprint(f\"  between {tie_lo:.2f} and {tie_hi:.2f} private SE across the whole range of test sizes,\")\nprint(\"  and less again once label correlation inflates the SE.\")\nprint(f\"  Under every assumption tried, a displayed tie is a fraction of one standard error.\")\nprint(f\"  Matching {CEILING:.3f} therefore secures nothing. It buys a ticket inside a\"\n      f\" {TIE_BLOCK}-team block.\")"},{"cell_type":"markdown","id":"cell05","metadata":{},"source":"## 2. The pair pays E[max], and E[max] has a proximity condition\n\nLet `X` and `Y` be the private scores of your two selected entries. Kaggle keeps the better one, so\nthe payoff is `E[max(X, Y)]` and the value of adding `Y` alongside your best entry `X` is\n\n```\ngain = E[max(X, Y)] - E[X]\n```\n\nModel both as normal with the same standard deviation `sigma`, means `mu_X >= mu_Y`, and correlation\n`rho`. Clark (1961) gives `E[max]` exactly:\n\n```\ns      = sqrt(sigma_X^2 + sigma_Y^2 - 2*rho*sigma_X*sigma_Y)\ntheta  = (mu_X - mu_Y) / s\nE[max] = mu_X*Phi(theta) + mu_Y*Phi(-theta) + s*phi(theta)\n```\n\nTwo limits are worth holding in your head.\n\n**Equal means.** With `mu_X = mu_Y` this collapses to `gain = sigma * sqrt((1 - rho) / pi)`. That is\nthe version people quote, and it is where the whole \"decorrelate your second submission\" folklore\ncomes from. It is correct, and it applies only here.\n\n**Unequal means.** As `mu_Y` falls below `mu_X`, `theta` grows and `Phi(-theta)` - the probability\n`Y` is ever the max - collapses like a Gaussian tail. At `rho = 0`, the most favourable case there\nis, the slot goes inert at a gap of just over two `sigma`, and by four `sigma` the gain rounds to\nzero. Since `rho = 0` is the ceiling on what decorrelation can do, nothing further back than that\nis rescuable. A perfectly decorrelated model three standard errors behind is not a hedge. It is an\nunused slot.\n\n**Why the units matter.** Write `d = (mu_X - mu_Y) / sigma`. Then `s = sigma*sqrt(2(1-rho))`,\n`theta = d / sqrt(2(1-rho))`, and `gain / sigma` is a function of `d` and `rho` alone - `sigma`\ndivides out. Every threshold below is therefore a statement about *standard errors*, and it holds\nwhatever the test size turns out to be. The first cell checks this numerically at three very\ndifferent sigmas instead of taking my word for it.\n\nWorth being blunt about how this notebook came to exist: I applied the equal-means formula to a pair\nwhose means were not close, recommended a hedge on that basis, and only caught it when I computed\n`E[max]` properly. The proximity condition is the part the folklore drops."},{"cell_type":"code","id":"cell06","execution_count":null,"metadata":{},"outputs":[],"source":"def e_max(mu1, mu2, s1, s2, rho):\n    \"\"\"E[max(X, Y)] for a bivariate normal (Clark 1961). Exact, not an approximation.\"\"\"\n    sd = math.sqrt(max(s1 * s1 + s2 * s2 - 2 * rho * s1 * s2, 0.0))\n    if sd <= 1e-15:\n        return max(mu1, mu2)\n    th = (mu1 - mu2) / sd\n    return mu1 * ND.cdf(th) + mu2 * ND.cdf(-th) + sd * ND.pdf(th)\n\n\ndef gain_sigmas(gap_sigmas, rho):\n    \"\"\"Expected gain from the second slot, in units of sigma.\n\n    gap_sigmas is how far the second entry's mean sits BELOW the first, measured in sigmas.\n    Returns gain / sigma, which depends on nothing else - see the invariance check below.\n    \"\"\"\n    sigma = 1.0\n    return e_max(0.0, -gap_sigmas * sigma, sigma, sigma, rho) - 0.0\n\n\ndef verdict(gain_in_sigmas):\n    \"\"\"A blunt three-way read on whether the second slot is doing anything.\"\"\"\n    if gain_in_sigmas < 0.05:\n        return \"INERT\"\n    if gain_in_sigmas < 0.25:\n        return \"WEAK\"\n    return \"LIVE\"\n\n\n# ---- invariance check: gain/sigma must not depend on sigma ------------------\n# This is the claim the whole correction rests on, so it is tested, not asserted.\nfor probe_sigma in (SIGMA_LO, SIGMA_MID, SIGMA_HI, 1.0):\n    for d, r in [(0.0, 0.2), (0.0, 0.9996), (1.0, 0.5), (2.0, 0.0)]:\n        absolute = (e_max(CEILING, CEILING - d * probe_sigma, probe_sigma, probe_sigma, r)\n                    - CEILING) / probe_sigma\n        assert abs(absolute - gain_sigmas(d, r)) < 1e-12, (probe_sigma, d, r)\nprint(f\"invariance check: gain/sigma identical at sigma = {SIGMA_LO:.5f}, {SIGMA_MID:.5f},\"\n      f\" {SIGMA_HI:.5f} and 1.0, to 1e-12.\")\n\n# Check the exact formula against the equal-means closed form.\nfor r in (0.0, 0.5, 0.9, 0.99):\n    assert abs(gain_sigmas(0.0, r) - math.sqrt((1 - r) / math.pi)) < 1e-12\nprint(\"Clark E[max] matches sqrt((1-rho)/pi) at equal means, to 1e-12.\")\nprint()\n\n\ndef bisect(f, lo, hi, iters=200):\n    \"\"\"Smallest x in [lo, hi] where the decreasing predicate f stops holding.\"\"\"\n    for _ in range(iters):\n        mid = (lo + hi) / 2\n        if f(mid):\n            lo = mid\n        else:\n            hi = mid\n    return lo\n\n\n# Where does the INERT verdict bite at equal means? Solve gain(rho) = 0.05 sigma.\nRHO_BREAKEVEN = bisect(lambda r: gain_sigmas(0.0, r) > 0.05, 0.0, 1.0 - 1e-12)\nprint(f\"break-even rho at equal means, gain = 0.05 sigma: {RHO_BREAKEVEN:.4f}\")\nprint(f\"  closed form 1 - pi*0.05^2 = {1 - math.pi * 0.05 ** 2:.4f}\")\nprint(\"  Above that correlation, even a perfectly level second entry is INERT.\")\nprint()\n\n# And the other axis: how far back can the second entry be before the slot goes INERT?\nprint(\"gap at which the second slot goes INERT, by correlation:\")\nfor r in (0.0, 0.2, 0.5, 0.9):\n    g = bisect(lambda t: gain_sigmas(t, r) > 0.05, 0.0, 12.0)\n    print(f\"  rho = {r:<4} -> {g:.4f} private SE\")\nGAP_INERT_BEST = bisect(lambda t: gain_sigmas(t, 0.0) > 0.05, 0.0, 12.0)\nprint()\nprint(f\"The rho = 0 row is the whole proximity rule: {GAP_INERT_BEST:.2f} private SE is where even\")\nprint(\"  a PERFECTLY uncorrelated second entry stops being worth anything. No correlation\")\nprint(\"  rescues a candidate further back than that, because rho = 0 is the best case.\")\nprint(\"  Rounding it to two SEs is the rule of thumb; it is derived, not guessed.\")\nprint()\nprint(\"In points of macro-AUC that window depends on the test size we do not know:\")\nfor n_test, s in zip(TEST_SIZES, SIGMAS):\n    print(f\"  test {n_test:>5} studies -> sigma {s:.5f} -> proximity window\"\n          f\" {GAP_INERT_BEST * s:.4f} of score\")\nprint(f\"  So: between {GAP_INERT_BEST * SIGMA_LO:.4f} and {GAP_INERT_BEST * SIGMA_HI:.4f}\"\n      \" of macro-AUC. The sigma statement is the one to carry.\")"},{"cell_type":"markdown","id":"cell07","metadata":{},"source":"## 3. The gain surface, as a table and as a cliff\n\nBoth axes are unitless, which is the point. Rows are how far the second entry sits below your best,\n**in private standard errors**. Columns are the correlation between the two entries' private scores.\nCells are the expected gain, also **in standard errors**, so you can read the verdict straight off:\nbelow `0.05` is inert, below `0.25` is weak.\n\nThis table is the same whatever the hidden test size is. Read the last column first: `rho` near 1 is\nwhat you get when both selections come out of one pipeline, which is what the site picks for you by\ndefault."},{"cell_type":"code","id":"cell08","execution_count":null,"metadata":{},"outputs":[],"source":"GAPS_SE = [0.0, 0.25, 0.5, 1.0, 1.5, 2.0, 3.0, 4.0]\nRHOS = [0.0, 0.2, 0.5, 0.9, 0.99, 0.9996]\n\nrows = []\nfor d in GAPS_SE:\n    row = {\"gap (SE)\": f\"{d:.2f}\"}\n    for r in RHOS:\n        row[f\"rho={r}\"] = round(gain_sigmas(d, r), 5)\n    rows.append(row)\ntable = pd.DataFrame(rows)\nprint(\"expected gain from the second slot, in units of sigma\")\nprint(\"(INERT below 0.05, WEAK below 0.25)\")\nprint(table.to_string(index=False))\nprint()\n\nfor note, d, r in [(\"level and decorrelated\", 0.0, 0.2),\n                   (\"level but a sibling\", 0.0, 0.9996),\n                   (\"decorrelated, 2 SE back\", 2.0, 0.2),\n                   (\"decorrelated, 4 SE back\", 4.0, 0.0)]:\n    g = gain_sigmas(d, r)\n    print(f\"  {note:<25}gap {d:.1f} SE   rho {r:<8}gain {g:.5f} sigma   {verdict(g)}\")\nprint()\nprint(\"Only one of those four is doing any work, and it is the one that is level.\")\nprint()\nbest = gain_sigmas(0.0, 0.2)\nprint(f\"A level, decorrelated second entry is worth {best:.3f} sigma. In points of macro-AUC that\")\nprint(f\"  is {best * SIGMA_LO:.5f} to {best * SIGMA_HI:.5f}, depending on the test size - which is\")\nprint(\"  exactly the kind of conversion this notebook now refuses to collapse to one number.\")"},{"cell_type":"code","id":"cell09","execution_count":null,"metadata":{},"outputs":[],"source":"fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(12.5, 4.6))\nINERT_LINE = 0.05\nYTOP = 0.62\n\n# Left: gain vs mean gap, one curve per correlation. Both axes in sigmas.\ngaps = np.linspace(0.0, 3.0, 241)\nfor r, style in [(0.0, \"-\"), (0.5, \"-\"), (0.9, \"--\"), (0.99, \":\"), (0.9996, \"-.\")]:\n    ax1.plot(gaps, [gain_sigmas(g, r) for g in gaps], style, linewidth=1.9, label=f\"rho = {r}\")\nax1.axhline(INERT_LINE, color=\"0.35\", linewidth=1.0)\nax1.text(2.97, 0.072, \"INERT below (0.05 sigma)\", fontsize=8.5, color=\"0.3\",\n         ha=\"right\", va=\"bottom\")\nax1.axvline(GAP_INERT_BEST, color=\"0.75\", linewidth=1.0)\nax1.text(GAP_INERT_BEST - 0.06, 0.13, \"rho = 0 goes inert here\", fontsize=8.5, color=\"0.4\",\n         rotation=90, va=\"bottom\", ha=\"right\")\nax1.set_xlim(0, 3)\nax1.set_ylim(-0.015, YTOP)\nax1.set_xlabel(\"private SEs the second entry sits below your best\")\nax1.set_ylabel(\"expected gain, in units of sigma\")\nax1.set_title(\"Proximity is the binding constraint\")\nax1.legend(fontsize=8.5, frameon=False, loc=\"upper right\")\nax1.grid(alpha=0.25, linewidth=0.6)\n\n# Right: gain vs correlation, one curve per mean gap.\nrs = np.linspace(0.0, 0.9999, 400)\nfor g, style, name in [(0.0, \"-\", \"level with your best\"),\n                       (0.5, \"-\", \"0.5 SE back\"),\n                       (1.0, \"--\", \"1.0 SE back\"),\n                       (2.0, \":\", \"2.0 SE back\")]:\n    ax2.plot(rs, [gain_sigmas(g, r) for r in rs], style, linewidth=1.9, label=name)\nax2.axhline(INERT_LINE, color=\"0.35\", linewidth=1.0)\nax2.axvline(RHO_BREAKEVEN, color=\"0.75\", linewidth=1.0)\nax2.text(0.02, 0.072, \"INERT below (0.05 sigma)\", fontsize=8.5, color=\"0.3\",\n         ha=\"left\", va=\"bottom\")\nax2.text(RHO_BREAKEVEN - 0.015, 0.025, f\"break-even rho {RHO_BREAKEVEN:.3f}\",\n         fontsize=8.5, color=\"0.3\", ha=\"right\", va=\"bottom\")\nax2.set_xlim(0, 1.0)\nax2.set_ylim(-0.015, YTOP)\nax2.set_xlabel(\"correlation between the two entries' private scores\")\nax2.set_ylabel(\"expected gain, in units of sigma\")\nax2.set_title(\"Decorrelation only pays once you are level\")\nax2.legend(fontsize=8.5, frameon=False, loc=\"upper right\")\nax2.grid(alpha=0.25, linewidth=0.6)\n\nfig.suptitle(\"What a second selected submission is worth, in standard errors\", fontsize=11.5)\nfig.tight_layout()\nplt.show()\n\nprint(f\"Left: every curve is at or under the INERT line past {GAP_INERT_BEST:.2f} SE, including\")\nprint(\"  rho = 0, which is the best case. That is the whole finding in one line.\")\nprint(\"Right: only the level curve has meaningful height anywhere, and it still collapses\")\nprint(\"  as rho approaches 1, which is exactly where sibling submissions live.\")\nprint(\"Neither panel has a unit that depends on the test-set size.\")"},{"cell_type":"markdown","id":"cell10","metadata":{},"source":"## 4. Measure rho. Do not assume it.\n\nThe `rho` in `E[max]` is the correlation between the two entries' **private scores across\nresamples**. What you can actually measure before the deadline is the correlation between their\n**predictions**. Those move together but are not the same quantity, so prediction-rho is a proxy -\nan informative one, since near-independent predictions cannot produce near-identical scores, but a\nproxy. It licenses an ordering claim (\"this pair is more correlated than that pair\"), not a point\nforecast of the gain.\n\nIt is still far better than a guess. We rank-correlated two of our own submissions out of the same\npipeline family and measured Spearman between **0.9966 and 0.9996**. Intuition had said something\naround 0.5. The cell below prices the same pair under both, so you can see what the substitution\ncosts. Like everything else here, the answer is in sigmas and does not depend on the test size.\n\n`measure_rho` takes two submission CSVs, aligns them on the id column, pools every shared prediction\ncolumn, and returns a tie-averaged Spearman. Point it at your own pair."},{"cell_type":"code","id":"cell11","execution_count":null,"metadata":{},"outputs":[],"source":"def measure_rho(path_a, path_b, id_col=None):\n    \"\"\"Tie-averaged Spearman between two submission CSVs.\n\n    Aligns rows on the id column, pools every shared prediction column, and returns\n    (rho, n_aligned_rows, n_shared_columns). A proxy for score-correlation, not\n    score-correlation itself - see the note above.\n    \"\"\"\n    a = pd.read_csv(path_a)\n    b = pd.read_csv(path_b)\n    idc = id_col or a.columns[0]\n    if idc not in a.columns or idc not in b.columns:\n        raise ValueError(f\"id column {idc!r} missing from one of the files\")\n    shared = [c for c in a.columns if c in b.columns and c != idc]\n    if not shared:\n        raise ValueError(\"no shared prediction columns between the two files\")\n    merged = a[[idc] + shared].merge(b[[idc] + shared], on=idc, suffixes=(\"_a\", \"_b\"))\n    if len(merged) == 0:\n        raise ValueError(\"no overlapping ids between the two files\")\n    xa = merged[[c + \"_a\" for c in shared]].to_numpy(dtype=float).ravel()\n    xb = merged[[c + \"_b\" for c in shared]].to_numpy(dtype=float).ravel()\n    ra = pd.Series(xa).rank()          # pandas averages ties, as Spearman requires\n    rb = pd.Series(xb).rank()\n    return float(ra.corr(rb)), len(merged), len(shared)\n\n\n# --- three throwaway submission files, so the helper is demonstrated and not just defined.\n# The row count here is arbitrary and affects nothing; it is a demo file, not a test-size claim.\nrng = np.random.default_rng(20260826)\nn_demo = 1000\nuids = [f\"1.2.826.0.{i:07d}\" for i in range(n_demo)]\nbase = rng.random((n_demo, len(LABELS)))\ndemo_dir = tempfile.mkdtemp(prefix=\"selection-demo-\")\n\n\ndef write_sub(name, arr):\n    frame = pd.DataFrame(arr, columns=LABELS)\n    frame.insert(0, \"StudyInstanceUID\", uids)\n    path = os.path.join(demo_dir, name)\n    frame.to_csv(path, index=False)\n    return path\n\n\np_base = write_sub(\"base.csv\", base)\n# a sibling: same pipeline, one hyperparameter nudged\np_sib = write_sub(\"sibling.csv\", np.clip(base + rng.normal(0, 0.004, base.shape), 0, 1))\n# a structurally different model\np_diff = write_sub(\"different.csv\", 0.35 * base + 0.65 * rng.random(base.shape))\n\nr_self, n_self, k_self = measure_rho(p_base, p_base)\nprint(f\"self-correlation check, a file against itself: {r_self:.6f}\"\n      f\"   ({n_self:,} rows x {k_self} columns)\")\nassert abs(r_self - 1.0) < 1e-9, \"a file must correlate 1.000000 with itself\"\nr_sib, _, _ = measure_rho(p_base, p_sib)\nr_diff, _, _ = measure_rho(p_base, p_diff)\nprint(f\"sibling pair                : rho = {r_sib:.6f}\")\nprint(f\"structurally different pair : rho = {r_diff:.6f}\")\nprint()\n\n# --- what assuming rho costs. All of this is in sigmas, so it is test-size free.\nprint(\"Same pair, same formula, one input changed:\")\nprint(f\"{'second entry':<22}{'rho':<20}{'gain (sigma)':>14}   verdict\")\nprint(\"-\" * 65)\nfor gap_se, gap_name in [(0.0, \"level with your best\"), (0.5, \"0.5 SE back\")]:\n    for r in (0.5, 0.9966, 0.9996):\n        g = gain_sigmas(gap_se, r)\n        tag = f\"{'assumed' if r == 0.5 else 'measured'} {r}\"\n        # Gaussian-tail values this deep are numerically fine and practically meaningless;\n        # showing 73 decimal places would imply a precision nobody should read into it.\n        shown = f\"{g:.5f}\" if g >= 1e-5 else (\"< 1e-6\" if g < 1e-6 else f\"{g:.1e}\")\n        print(f\"{gap_name:<22}{tag:<20}{shown:>14}   {verdict(g)}\")\nprint()\nlvl_assumed = gain_sigmas(0.0, 0.5)\nlvl_measured = gain_sigmas(0.0, 0.9996)\nprint(f\"Level with your best, the guess prices this pair {lvl_assumed / lvl_measured:.0f} times\"\n      \" higher than the measurement does\")\nprint(f\"  ({lvl_assumed:.5f} against {lvl_measured:.5f} sigma), and it is the difference between\"\n      f\" {verdict(lvl_assumed)} and {verdict(lvl_measured)}.\")\nprint(\"Move the second entry half a standard error back - one or two displayed thousandths,\")\nprint(\"  depending on the test size - and the measured-rho gain is zero for practical purposes\")\nprint(\"  while the assumed-rho gain still looks worth having. An assumed rho makes the\")\nprint(\"  whole calculation decorative.\")"},{"cell_type":"markdown","id":"cell12","metadata":{},"source":"## 5. Low correlation is ambiguous, and the ambiguity is expensive\n\nSection 4 makes `rho` look like the thing to minimise. It is not, and this is the part I got wrong\nfirst.\n\nA model can be uncorrelated with your best one for two very different reasons.\n\n- **Complementary** - it is right where the other is wrong. This is what you want.\n- **Weak** - it is closer to noise, and noise is uncorrelated with everything.\n\nPrediction-rho cannot distinguish them. Both produce a small number.\n\nWe ran into this directly. Our strongest entry is a stack that builds on public community work; the\nother model is one we trained ourselves on a different data path, and its measured rho against that\nentry is **0.231** - far below break-even, comfortably `LIVE` on the pricing above, the only\nartifact we hold that has ever returned `LIVE` from this calculation. I wrote at the time that we\nhad the decorrelation and only lacked the level.\n\nThat framing was too generous, and the next two cells show why.\n\n**The diagnostic: count per-label wins.** The metric is a mean over twelve labels. If a second model\nis genuinely complementary it should beat the reference on *at least one* of them. If it loses on\nall twelve, its decorrelation is weakness, and no amount of hedging converts that into score.\n\nThis section is measured on 58 labelled studies we hold, so nothing in it depends on the test-set\nsize either. It is the one part of the notebook that was never in sigma units and never needed to\nbe."},{"cell_type":"code","id":"cell13","execution_count":null,"metadata":{},"outputs":[],"source":"# Measured on our 58 gold-labelled studies, 58/58 aligned, tie-aware Mann-Whitney AUC.\n# These twelve pairs are the output of that run. The raw per-study predictions are not in this\n# notebook, so the table is stated as measured rather than recomputed here.\nREFERENCE_AUC = [0.9804, 0.9773, 0.9868, 0.9366, 0.9868, 0.9497,\n                 0.9447, 0.9975, 0.9677, 0.9928, 0.9710, 0.9903]\nCANDIDATE_AUC = [0.6752, 0.5079, 0.6611, 0.7416, 0.7659, 0.7834,\n                 0.6692, 0.8149, 0.6464, 0.6630, 0.5992, 0.7139]\n# 2,000-rep study-bootstrap SD of the reference AUC, same run. The reference is an estimate\n# too, which limits how hard anyone should lean on any single row.\nREFERENCE_BOOT_SD = [0.016, 0.018, 0.011, 0.029, 0.011, 0.030,\n                     0.029, 0.003, 0.020, 0.009, 0.018, 0.010]\n\nperlab = pd.DataFrame({\n    \"label\": LABELS,\n    \"pos/58\": POS_GOLD,\n    \"reference\": REFERENCE_AUC,\n    \"ref bootSD\": REFERENCE_BOOT_SD,\n    \"candidate\": CANDIDATE_AUC,\n})\nperlab[\"delta\"] = (perlab[\"candidate\"] - perlab[\"reference\"]).round(4)\nperlab[\"win\"] = perlab[\"delta\"] > 0\nprint(perlab.to_string(index=False))\nprint()\n\nOBS_WINS = int(perlab[\"win\"].sum())\nprint(f\"macro: reference {np.mean(REFERENCE_AUC):.6f}   candidate {np.mean(CANDIDATE_AUC):.6f}\")\nprint(f\"per-label wins: {OBS_WINS} of {len(LABELS)}\")\nprint(f\"worst deficit {perlab['delta'].min():+.4f}, best {perlab['delta'].max():+.4f}\"\n      \" - the candidate never leads.\")\nprint()\nprint(\"Per-label AUC 0.51 to 0.81 against 0.94 to 0.99. The rho = 0.231 is the second kind of\")\nprint(\"decorrelation, not the first.\")"},{"cell_type":"code","id":"cell14","execution_count":null,"metadata":{},"outputs":[],"source":"def auc_ranked(y_true, scores):\n    \"\"\"Tie-aware ROC-AUC via the Mann-Whitney rank sum.\"\"\"\n    y = np.asarray(y_true).astype(int)\n    r = pd.Series(np.asarray(scores, dtype=float)).rank().to_numpy()\n    n_pos = int(y.sum())\n    n_neg = len(y) - n_pos\n    if n_pos == 0 or n_neg == 0:\n        return float(\"nan\")\n    return (r[y == 1].sum() - n_pos * (n_pos + 1) / 2.0) / (n_pos * n_neg)\n\n\n# Self-test in three directions before the helper is used for anything.\n_y = np.array([0, 0, 1, 1])\nassert abs(auc_ranked(_y, [0.1, 0.2, 0.8, 0.9]) - 1.0) < 1e-12\nassert abs(auc_ranked(_y, [0.9, 0.8, 0.2, 0.1]) - 0.0) < 1e-12\nassert abs(auc_ranked(_y, [0.5, 0.5, 0.5, 0.5]) - 0.5) < 1e-12\nprint(\"auc_ranked self-test passed: perfect 1.0, reversed 0.0, all-tied 0.5.\")\nprint()\n\n# ---- permutation null -------------------------------------------------------\n# Under the null \"the candidate carries no signal\", its per-label AUC distribution depends only on\n# n and the label's positive count. So the null is fully reproducible from n = 58, the twelve\n# positive counts, and the measured reference AUCs. 3,000 shuffles, seeded.\nN_SHUFFLE = 3000\nrng_null = np.random.default_rng(20260826)\nnull_wins = np.zeros(N_SHUFFLE, dtype=int)\n\nfor j, n_pos in enumerate(POS_GOLD):\n    n_neg = N_GOLD - n_pos\n    # each row is a random permutation of the ranks 1..58\n    perm = np.argsort(rng_null.random((N_SHUFFLE, N_GOLD)), axis=1) + 1\n    rank_sum = perm[:, :n_pos].sum(axis=1)\n    aucs = (rank_sum - n_pos * (n_pos + 1) / 2.0) / (n_pos * n_neg)\n    null_wins += (aucs > REFERENCE_AUC[j]).astype(int)\n\nprint(f\"permutation null, {N_SHUFFLE:,} shuffles of a no-signal candidate:\")\nprint(f\"  mean wins {null_wins.mean():.2f} | sd {null_wins.std():.2f}\"\n      f\" | 95th pct {np.percentile(null_wins, 95):.0f} | max {null_wins.max()}\")\nprint(f\"  P(at least one win | no signal) = {(null_wins >= 1).mean():.4f}\")\nprint(f\"  observed = {OBS_WINS}\")\nprint(f\"  So {OBS_WINS} wins is exactly what no signal predicts. The per-label evidence for\")\nprint(\"  complementarity is zero, not merely small.\")\nprint()\n\n# ---- and the check that the check needs: can this diagnostic ever fire? -----\npositive_control = list(CANDIDATE_AUC)\nfor j in (3, 5, 6):        # Lateral Meniscus, Lateral OA, PF OA - the reference's weakest labels\n    positive_control[j] = REFERENCE_AUC[j] + 0.02\npc_wins = sum(1 for a, b in zip(positive_control, REFERENCE_AUC) if a > b)\nprint(\"positive control, a candidate built to beat the reference on three labels:\")\nprint(f\"  wins {pc_wins} of {len(LABELS)}\"\n      f\" | P(>= {pc_wins} | no signal) = {(null_wins >= pc_wins).mean():.4f}\")\nprint(\"  A diagnostic that can never return a positive would be worthless, so that direction\")\nprint(\"  is checked too. It fires.\")"},{"cell_type":"markdown","id":"cell15","metadata":{},"source":"## 6. What to do with your two slots\n\nThe procedure this implies is short, and every step is in units you can compute for yourself.\n\n1. **Rank candidates by expected private score, not public score.** Public is 30% of the test set,\n   so its standard error is the larger of the two.\n2. **Slot one is your ceiling.** No cleverness.\n3. **For slot two, check proximity first.** If its mean is more than about `2 * sigma` below your\n   ceiling, stop. That gap is where the slot goes inert even at `rho = 0`, and `rho = 0` is the\n   best case, so nothing you do to its correlation changes the answer. Estimate your own `sigma`\n   by bootstrapping your held-out predictions rather than taking a number from a notebook - this\n   one included.\n4. **Only then measure rho**, with `measure_rho` above, on the actual CSVs. Siblings out of one\n   pipeline routinely sit past 0.99, where the gain is already gone.\n5. **Then count per-label wins.** If the candidate loses on every label, its low rho is weakness.\n   Fix the model or accept that the slot is empty, but do not select it and call it a hedge.\n\nA qualifying second entry has to clear both conditions at once, and that is a harder object to own\nthan it sounds. In our own portfolio the near-ceiling candidates are all siblings of one another,\nand the one genuinely decorrelated model sits far below the ceiling. We do not currently hold a\nqualifying second entry. That is a build problem rather than a selection problem, and saying so is\nmore useful than selecting something and calling it a hedge.\n\nThe default matters here too. Left alone, the site selects your two best public scores, which for\nmost people are two runs of one pipeline. On the table in section 3 that pair is the `rho = 0.9996`\ncolumn at a gap of zero: about `0.011` sigma, which the verdict function calls `INERT`."},{"cell_type":"markdown","id":"cell16","metadata":{},"source":"## 7. Limits, stated where you can see them\n\nI would rather you distrust the right parts of this than all of it.\n\n- **The test-set size is unknown, and that is now handled rather than hidden.** The 30/70 split is\n  published; the absolute size is not. Section 1 carries a range from 600 to 5,000 studies, over\n  which `sigma` varies by a factor of about three. Every threshold in the notebook is quoted in\n  sigmas, where that factor cancels. If you want points of macro-AUC, take the range, not a\n  midpoint.\n- **The standard error is a lower bound within each scenario.** Hanley-McNeil averaged across labels\n  assumes the twelve labels are independent. They co-occur on the same knee, so the true `sigma` is\n  larger than the table shows - the sensitivity rows give the inflation for `rho_L` up to 0.5.\n- **And it is unpaired, which pushes the other way.** Two models scored on the *same* test studies\n  form a paired comparison with a smaller SE than either marginal figure. The two corrections have\n  opposite signs and I have not netted them.\n- **Prevalence is estimated from 58 studies,** not from the test set. One study moves a label's\n  prevalence by about 0.017.\n- **Prediction-rho is a proxy for score-rho.** They move together and are not the same quantity, so\n  the pricing supports an ordering claim rather than a point forecast.\n- **Rank translations are scale, not prediction, and now doubly so.** At today's density around the\n  cutoff, 228 teams occupy one displayed thousandth. A level, decorrelated second entry is worth\n  about half a sigma, which converts to somewhere between roughly 0.0008 and 0.0023 of macro-AUC\n  across the test sizes in section 1, and thence to a few hundred displayed ranks. That arithmetic\n  holds every rival's score fixed, uses public density as a stand-in for private density, and rests\n  on a test size nobody outside the host knows. Read it as \"this is not a rounding error\", and\n  nothing more.\n- **The permutation null tests one thing.** It asks whether a no-signal candidate would produce the\n  observed win count. It does not establish that the candidate is worthless in general, only that it\n  never leads the reference on any label at this sample size. Per-label differences at n = 58\n  resolve effects of roughly 0.01 to 0.03; smaller ones would not show.\n- **Normality.** `E[max]` assumes bivariate normal private scores. A macro-average over twelve\n  labels on a test set of this order is a reasonable place to assume that, and I have not verified\n  it.\n- **The rho measurement is n = 1 in the sense that matters.** \"Siblings correlate past 0.99\" is\n  measured on our own pipeline families. Yours may differ, which is the entire reason the helper\n  exists rather than a constant.\n\nTwo references, both worth reading rather than merely citing: Hanley and McNeil (1982) for the AUC\nstandard error, and Clark (1961) for the exact moments of the maximum of correlated normals.\n\nIf you run `measure_rho` on your own pair and get something surprising, that is the interesting\noutcome, not this notebook."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11.0","mimetype":"text/x-python","file_extension":".py","pygments_lexer":"ipython3","nbconvert_exporter":"python","codemirror_mode":{"name":"ipython","version":3}}},"nbformat":4,"nbformat_minor":5}