{"cells":[{"cell_type":"markdown","metadata":{},"source":"# Both medal lines are inside one tie block\n\nOn this competition the displayed leaderboard score has stopped carrying medal information.\n\nSnapshot taken 2026-08-28, from the public leaderboard export: **2,548 teams**. The silver line\nfalls at rank 127 and the bronze line at rank 255. Both of those ranks display the same score,\n**0.936** - because 278 teams display 0.936, spanning **rank 93 to rank 370**.\n\nBoth medal lines run *through* a single block of tied displayed scores. A team showing 0.936 may\nbe silver, may be bronze, may be neither, and the number on the page cannot tell them which.\n\nThis notebook recomputes that from the snapshot rather than asserting it, prices the tie in units\nof the metric's own standard error, and is explicit about the one quantity nobody outside the\ncompetition knows.\n","id":"c00"},{"cell_type":"markdown","metadata":{},"source":"## The thing nobody outside can measure\n\nThe public/private split is stated by the competition and verified: **30% public, 70% private**.\n\nThe **absolute size of the test set is not public**. That matters, because the standard error of a\nmacro-averaged AUC depends on how many studies it is averaged over. Rather than adopt a single\nvalue, this notebook sweeps a wide range of plausible test sizes and reports every result as a\nrange.\n\nThat is not a hedge, it is the stronger form of the argument: if a conclusion holds across the\nentire sweep, it never depended on knowing the size in the first place.\n","id":"c01"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import math\n\n# --- snapshot, public leaderboard export, 2026-08-28 -------------------------------\nTEAMS = 2646\nBLOCK_SCORE = \"0.936\"\nBLOCK_LO, BLOCK_HI = 116, 446         # first and last rank displaying BLOCK_SCORE\nHISTORY = [dict(date=\"2026-08-27\", teams=2524, lo=87, hi=350),\n           dict(date=\"2026-08-28\", teams=2548, lo=93, hi=370)]\n\n# Kaggle's published medal thresholds for a competition of this size\ndef medal_lines(n_teams):\n    return {\"gold\": 10 + round(n_teams * 0.002),\n            \"silver\": round(n_teams * 0.05),\n            \"bronze\": round(n_teams * 0.10)}\n\nLINES = medal_lines(TEAMS)\nBLOCK_N = BLOCK_HI - BLOCK_LO + 1\n\nprint(f\"teams                : {TEAMS:,}\")\nprint(f\"tie block            : {BLOCK_N} teams displaying {BLOCK_SCORE}, ranks {BLOCK_LO}-{BLOCK_HI}\")\nfor k in (\"gold\", \"silver\", \"bronze\"):\n    r = LINES[k]\n    inside = BLOCK_LO <= r <= BLOCK_HI\n    print(f\"{k+' line':<21}: rank {r:>4}   {'INSIDE the tie block' if inside else 'outside the block'}\")\n\nboth_inside = all(BLOCK_LO <= LINES[k] <= BLOCK_HI for k in (\"silver\", \"bronze\"))\nprint()\nif both_inside:\n    span = LINES[\"bronze\"] - LINES[\"silver\"]\n    print(f\"Both the silver and bronze lines fall inside one block of tied displayed scores.\")\n    print(f\"{span} teams sit between them, and every one of them displays {BLOCK_SCORE}.\")\nelse:\n    print(\"On this snapshot the finding described above does not hold; the lines are not both inside.\")\n","id":"c02"},{"cell_type":"markdown","metadata":{},"source":"## How the block is changing","id":"c03"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"for h in HISTORY:\n    print(f\"{h['date']}: {h['hi'] - h['lo'] + 1:>3} teams at this score, ranks {h['lo']}-{h['hi']}\"\n          f\"   (field {h['teams']:,})\")\nprint(f\"2026-08-29: {BLOCK_N:>3} teams at this score, ranks {BLOCK_LO}-{BLOCK_HI}\"\n      f\"   (field {TEAMS:,})\")\nprint()\nfirst = HISTORY[0]\ngrew_block = BLOCK_N - (first[\"hi\"] - first[\"lo\"] + 1)\ngrew_field = TEAMS - first[\"teams\"]\nprint(f\"Over two days the block grew by {grew_block} teams while the field grew by {grew_field}.\")\nprint(f\"So {grew_block / grew_field:.0%} of all new entrants landed on this one displayed score,\")\nprint(f\"and the block now holds {BLOCK_N / TEAMS:.1%} of the entire field.\")\nprint(\"New arrivals concentrate here far above the base rate, which is what a widely shared\")\nprint(\"public approach looks like from the outside.\")\n","id":"c04"},{"cell_type":"markdown","metadata":{},"source":"## Pricing the tie in standard errors","id":"c05"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"def hanley_se(auc, n_pos, n_neg):\n    \"Standard error of a single AUC (Hanley and McNeil, 1982).\"\n    if n_pos < 1 or n_neg < 1:\n        return float(\"nan\")\n    q1 = auc / (2 - auc)\n    q2 = 2 * auc * auc / (1 + auc)\n    var = (auc * (1 - auc) + (n_pos - 1) * (q1 - auc * auc)\n           + (n_neg - 1) * (q2 - auc * auc)) / (n_pos * n_neg)\n    return math.sqrt(max(var, 0.0))\n\ndef macro_se(auc, n, prevalences, label_rho=0.0):\n    \"SE of a macro-averaged AUC. label_rho=0 assumes independent labels: a LOWER bound.\"\n    ses = [hanley_se(auc, max(1, round(n * p)), n - max(1, round(n * p))) for p in prevalences]\n    k = len(ses)\n    return math.sqrt(sum(s * s for s in ses)) / k * math.sqrt(1 + (k - 1) * label_rho)\n\n# per-label positive rates, from the competition's own labelled training studies\nPREV = [0.4138, 0.1552, 0.4483, 0.3966, 0.2586, 0.1897, 0.3621,\n        0.6034, 0.4655, 0.2069, 0.3276, 0.3103]\nAUC = 0.936\nPRIVATE_FRAC = 0.70\nTIE = 0.001                     # two adjacent displayed scores differ by at most this\n\nprint(f\"{'assumed test size':>18} {'private n':>10} {'private SE':>12} {'a tie, in SE':>14}\")\nsigmas = []\nfor t in (600, 1000, 1500, 2000, 3000, 5000):\n    n = round(t * PRIVATE_FRAC)\n    s = macro_se(AUC, n, PREV)\n    sigmas.append(s)\n    print(f\"{t:>18,} {n:>10,} {s:>12.5f} {TIE / s:>14.2f}\")\n\nlo, hi = min(sigmas), max(sigmas)\nprint()\nprint(f\"sigma spans {lo:.5f} to {hi:.5f} across the sweep, a factor of {hi / lo:.2f}.\")\nprint(f\"A displayed tie is between {TIE / hi:.2f} and {TIE / lo:.2f} standard errors.\")\nprint()\nprint(\"Across every test size tried, the gap between two tied displayed scores is a fraction\")\nprint(\"of one standard error. The conclusion does not depend on knowing the test size.\")\n","id":"c06"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import matplotlib.pyplot as plt\n\nsizes = [600, 1000, 1500, 2000, 3000, 5000]\nties = [TIE / macro_se(AUC, round(t * PRIVATE_FRAC), PREV) for t in sizes]\n\nfig, ax = plt.subplots(figsize=(8, 4.2))\nax.plot(sizes, ties, marker=\"o\", color=\"#33628f\")\nax.axhline(1.0, color=\"#b04a3a\", linestyle=\"--\", linewidth=1)\nax.text(sizes[-1], 1.02, \"one standard error\", ha=\"right\", va=\"bottom\",\n        color=\"#b04a3a\", fontsize=9)\nax.set_xlabel(\"assumed test-set size (studies)\")\nax.set_ylabel(\"a displayed tie, in standard errors\")\nax.set_title(\"A tied displayed score is a fraction of one standard error, at every assumed size\")\nax.set_ylim(0, 1.15)\nax.grid(alpha=0.25)\nfig.tight_layout()\nplt.show()\n","id":"c07"},{"cell_type":"markdown","metadata":{},"source":"## What this changes for a reader\n\n**Your displayed score is not your medal.** If your score sits inside the block, your position\nrelative to the lines is decided by digits the leaderboard does not show you.\n\n**Chasing the next displayed digit is the wrong target.** The block here is 0.001 wide and holds\n278 teams. An improvement large enough to change the printed number is far larger than the\nimprovement needed to move a long way inside the block, and far larger than the noise that\nseparates neighbours.\n\n**Treat your position as a distribution, not a fact.** The private split is a different sample of\nstudies. Everyone inside the block is re-drawn from it. A rank inside a tie block is closer to a\ntimestamp than to a measurement of skill.\n","id":"c08"},{"cell_type":"markdown","metadata":{},"source":"## Limits, stated plainly\n\n- The Hanley and McNeil standard error assumes labels are independent. They are not: these twelve\n  findings co-occur in the same knees. Positive correlation makes the true standard error *larger*\n  than the figures above, which makes a tie an even smaller fraction of it. The direction of that\n  error favours the conclusion, so it is stated rather than relied upon.\n- The per-label positive rates come from a small number of labelled training studies. They set the\n  scale of the standard error, not its order of magnitude.\n- A displayed tie is a *display* tie. The leaderboard ranks on full precision, so the teams in this\n  block are not identical - the point is that the visible number cannot separate them, and the\n  invisible difference is far smaller than the noise the private split will add.\n- Public-board density is not private-board density. Nothing here predicts anyone's final rank.\n- The absolute test-set size is deliberately not assumed anywhere in this notebook.\n","id":"c09"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"}},"nbformat":4,"nbformat_minor":5}