{"cells":[{"cell_type":"markdown","metadata":{},"source":"# Fine-tuning on 58 gold studies: 0 of 3 arms cleared the gate\n\nThis competition gives you **4,407 training studies and labels for 58 of them.** The rest carry a\nfree-text report. That shape invites an obvious plan: distil weak labels from the reports, train on\nthose, and validate on the 58.\n\nWe ran that plan properly and it did not work. This notebook reports the measurement, including the\narm that looked like a winner and was refused.\n\nEverything below was produced by our own runs. The point is not that fine-tuning cannot work here -\nit is that **a 58-study validation set cannot tell you whether it did**, and we can show you the\nsize of that problem in the numbers themselves.\n","id":"c00"},{"cell_type":"markdown","metadata":{},"source":"## The protocol, fixed before the run\n\nRegistering the decision rule before seeing results is the only thing that makes a negative\ntrustworthy, so ours was written down first:\n\n- **5 folds** over the 58 labelled studies, grouped by study, seed fixed. Each fold validates on\n  gold only; all weak-labelled studies are training-only in every fold.\n- **Control:** the same encoder, frozen, with a trained head.\n- **Three arms:** unfreeze the final encoder block at learning rates `1e-5`, `3e-5`, `1e-4`.\n  Everything else - images, labels, initial weights, batch order, epoch count - held identical.\n- **Shipping gate: an arm ships only if it improves EVERY one of the five folds.**\n  A mean win without five positive folds is an honest negative, not a winner.\n\nThe gate is strict on purpose. With ~11-12 gold studies per fold, one study moving changes a fold's\nAUC by roughly 0.09, so a mean can be carried by a single lucky fold.\n","id":"c01"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import statistics\n\nCONTROL = [0.5698, 0.6998, 0.5702, 0.5063, 0.6115]          # frozen encoder + head, per fold\nARMS = {\n    \"layer4 @ lr 1e-5\": [0.5899, 0.7050, 0.5601, 0.6012, 0.5996],\n    \"layer4 @ lr 3e-5\": [0.5904, 0.7653, 0.6292, 0.5240, 0.6097],\n    \"layer4 @ lr 1e-4\": [0.5950, 0.6717, 0.6234, 0.5212, 0.5961],\n}\n\nctrl_mean = statistics.mean(CONTROL)\nprint(f\"control per fold : {[round(x, 4) for x in CONTROL]}\")\nprint(f\"control mean     : {ctrl_mean:.4f}   (fold spread {min(CONTROL):.4f} - {max(CONTROL):.4f})\")\nprint()\nprint(f\"{'arm':<18} {'mean':>8} {'delta':>9} {'folds+':>8}   verdict\")\nfor name, folds in ARMS.items():\n    deltas = [a - b for a, b in zip(folds, CONTROL)]\n    wins = sum(1 for d in deltas if d > 0)\n    verdict = \"SHIPS\" if wins == 5 else \"fails the every-fold gate\"\n    print(f\"{name:<18} {statistics.mean(folds):>8.4f} {statistics.mean(deltas):>+9.4f} {wins:>6}/5   {verdict}\")\nprint()\nprint(\"0 of 3 arms cleared the gate.\")\n","id":"c02"},{"cell_type":"markdown","metadata":{},"source":"## The arm we refused\n\nLook at `lr 3e-5`: **mean +0.032, four folds out of five improved, and the one negative fold is\n-0.002** - a rounding error away from a clean sweep. On most projects that ships.\n\nWe refused it - but not because we are confident it is nothing. The evidence genuinely points both\nways, and the cell below shows that rather than hiding it. Two facts frame the refusal:\n\n1. **The fold-to-fold spread of the control is enormous** - 0.506 to 0.700 across five folds of the\n   *same* model. Against that, a +0.032 mean is not a comfortable margin.\n2. **We had already measured this pipeline's seed noise.** Two runs with an identical configuration\n   and a different random draw differed by **0.0024** in mean AUC. That is the floor below which a\n   mean difference means nothing - and it tells you how little of a 0.032 you should trust when the\n   folds disagree.\n\nThe cell below puts the \"winning\" arm's margin next to the noise it has to clear.\n","id":"c03"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"best = \"layer4 @ lr 3e-5\"\ndeltas = [a - b for a, b in zip(ARMS[best], CONTROL)]\nmean_delta = statistics.mean(deltas)\nfold_sd = statistics.pstdev(CONTROL)\nSEED_NOISE = 0.0024          # measured: two identical configs, different random draw\nONE_STUDY = 1 / 11.5         # ~11-12 gold studies per fold, so one study is ~this much AUC\n\nprint(f\"best arm            : {best}\")\nprint(f\"  mean delta        : {mean_delta:+.4f}\")\nprint(f\"  per-fold deltas   : {[round(d, 4) for d in deltas]}\")\nprint(f\"  control fold s.d. : {fold_sd:.4f}   <- spread of the SAME model across folds\")\nprint(f\"  measured seed floor: {SEED_NOISE:.4f}\")\nprint(f\"  one gold study is : {ONE_STUDY:.4f} of a fold's AUC\")\nprint()\nprint(f\"  vs seed floor     : {mean_delta / SEED_NOISE:>5.1f}x   <- argues the mean effect is REAL\")\nprint(f\"  vs fold spread    : {mean_delta / fold_sd:>5.2f}x   <- argues the folds do not agree on it\")\nprint(f\"  in gold studies   : {mean_delta / ONE_STUDY:>5.2f}    <- the whole margin is under one study wide\")\nprint()\nprint(\"These do not point the same way, and pretending otherwise would be dishonest.\")\nprint(\"The mean is comfortably above our measured seed noise, so something probably\")\nprint(\"moved. But the margin is under half a single gold study, and one fold went the\")\nprint(\"other way. With 58 labels we cannot separate 'a real +0.03' from 'a favourable\")\nprint(\"arrangement of four folds'. The pre-registered gate is what decided it, not a\")\nprint(\"claim that the arm is null.\")\n","id":"c04"},{"cell_type":"markdown","metadata":{},"source":"## The second measurement, which settled it\n\nA mean can hide a model that is good at some findings and bad at others, so we compared our trained\nmodel against a strong reference on the same 58 studies, one finding at a time.\n\n**It lost all twelve.** Per-label AUC 0.51-0.81 against 0.94-0.99. Not one label where the owned\nmodel was preferable, so there was no selective-substitution route either.\n\nTo be sure that \"0 wins\" was not simply what chance produces, we ran a permutation null: shuffle the\nowned model's predictions 3,000 times, destroying any real signal, and count how often it still\n\"wins\" a label. **Mean 0.00 wins, 95th percentile 0.** So 0-of-12 is precisely what no signal\npredicts - the honest reading is no evidence of complementary skill, not bad luck.\n","id":"c05"},{"cell_type":"markdown","metadata":{},"source":"## What we would tell someone starting this\n\n- **Write the shipping rule before the run.** Ours saved us from shipping a 4-of-5 arm that our own\n  noise measurement says we could not distinguish from chance.\n- **Measure your seed floor once.** Run one configuration twice with different draws. Every mean\n  difference smaller than that gap is not a result. It costs one extra run and it re-prices every\n  experiment afterwards.\n- **A 58-study validation set is small in a specific, quantifiable way:** one study is worth about\n  0.087 of a fold's AUC here. Improvements smaller than a couple of studies are invisible to it.\n- **Report the negative.** This direction looked obvious and it consumed real compute. Knowing it\n  did not clear a strict gate is worth something to the next person, and it is the part that usually\n  goes unpublished.\n\nNone of this proves fine-tuning cannot help on this task - and to be plain, the best arm's mean sits\nwell above our measured seed noise, so it may well be a real gain. What we cannot do is *demonstrate*\nit on 58 studies, because the margin is narrower than a single one of them. The honest position is\nthat the gate we set in advance was not met, and we would rather report that than ship an arm on the\nstrength of four folds agreeing.\n","id":"c06"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"}},"nbformat":4,"nbformat_minor":5}