{"cells":[{"cell_type":"markdown","metadata":{},"source":"# Your 58-study validation set is lying to you — measured, with a fix\n\nI trained four models for this competition and picked checkpoints on the 58 radiologist-annotated studies,\nbecause they are the only real labels we have. Then the leaderboard told me the model my validation ranked\n**last** was actually my **best**:\n\n| model | 58-study val | 250-study derived holdout | public LB |\n|---|---|---|---|\n| v2 — single sagittal series, ResNet-34 | 0.858 | — | 0.873 |\n| v3 — three planes + per-label plane mixture | **0.869** (best) | — | 0.874 |\n| v4 — ConvNeXt-T + slice attention, EMA | 0.850 (worst) | **0.861** | **0.883** (best) |\n\nThe 58-study set ranked v4 last by 0.019. The leaderboard ranked it first by 0.009. The larger holdout — the\none built from *derived* labels, which are individually noisier but far more numerous — got the order right.\n\nThis notebook measures why that happens, then shows the fix. It is short and CPU-only."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import glob, os\nimport numpy as np\nimport pandas as pd\nfrom sklearn.metrics import roc_auc_score\nimport matplotlib.pyplot as plt\n\nPATH = os.path.dirname(glob.glob('/kaggle/input/**/train.csv', recursive=True)[0])\ntrain = pd.read_csv(os.path.join(PATH, 'train.csv'))\nLABELS = list(train.columns[2:14])\nann = train[train[LABELS].notna().all(axis=1)].reset_index(drop=True)\nprint(f'annotated studies: {len(ann)} of {len(train)}')\nprint((ann[LABELS].sum().astype(int).rename('positives')).to_frame().assign(\n      negatives=lambda d: len(ann) - d.positives).to_string())"},{"cell_type":"markdown","metadata":{},"source":"## 1. How much can an AUC move on 58 studies, by luck alone?\n\nNine to thirty-five positives per label. To see what that means, take a **fixed** set of predictions — here,\na public LLM label table used as a stand-in for a model — and bootstrap the 58 studies. The predictions never\nchange; only which studies land in the sample does. Every bit of spread below is noise."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"src = glob.glob('/kaggle/input/**/llm_labels_v4_blend.csv', recursive=True)\ntbl = pd.read_csv(src[0]).set_index('StudyInstanceUID')\npred = tbl.loc[ann['StudyInstanceUID'], LABELS].values.astype(float)\ntruth = ann[LABELS].values.astype(int)\n\nrng = np.random.RandomState(0)\nB = 2000\nboot = np.zeros((B, len(LABELS))); boot_macro = np.zeros(B)\nfor b in range(B):\n    idx = rng.randint(0, len(ann), len(ann))\n    per = []\n    for j in range(len(LABELS)):\n        y, p = truth[idx, j], pred[idx, j]\n        per.append(roc_auc_score(y, p) if 0 < y.sum() < len(y) else np.nan)\n    boot[b] = per\n    boot_macro[b] = np.nanmean(per)\n\nlo, hi = np.percentile(boot_macro, [2.5, 97.5])\nprint(f'macro AUC point estimate : {np.nanmean([roc_auc_score(truth[:, j], pred[:, j]) for j in range(12)]):.4f}')\nprint(f'95% bootstrap interval   : {lo:.4f} - {hi:.4f}   (width {hi-lo:.4f})')"},{"cell_type":"markdown","metadata":{},"source":"**That interval width is the headline.** It is comparable to — often larger than — the entire spread between\nmy four models. Two models whose 58-study scores differ by less than that are indistinguishable, and picking\nthe higher one is a coin flip dressed up as a decision."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"fig, axes = plt.subplots(1, 2, figsize=(13, 4.5))\naxes[0].hist(boot_macro, bins=50, color='#4477AA')\naxes[0].axvline(lo, color='#CC3311', ls='--'); axes[0].axvline(hi, color='#CC3311', ls='--')\naxes[0].set_title(f'macro AUC over 2000 bootstraps of the same 58 studies\\n95% CI width = {hi-lo:.3f}')\naxes[0].set_xlabel('macro AUC')\n\nw = np.nanpercentile(boot, 97.5, axis=0) - np.nanpercentile(boot, 2.5, axis=0)\no = np.argsort(-w)\naxes[1].barh([LABELS[i] for i in o], w[o], color='#EE6677')\naxes[1].set_xlabel('95% CI width (AUC)')\naxes[1].set_title('Per-label uncertainty on 58 studies')\naxes[1].invert_yaxis()\nplt.tight_layout(); plt.show()\n\nprint('per-label 95% CI width:')\nfor i in o: print(f'   {LABELS[i]:18s} +/- {w[i]/2:.3f}   ({int(truth[:, i].sum())} positives)')"},{"cell_type":"markdown","metadata":{},"source":"Labels with the fewest positives (MCL at 9, Baker's at 12) carry CI widths around ±0.1. A model that \"improves\nMCL by 0.05\" has told you nothing."},{"cell_type":"markdown","metadata":{},"source":"## 2. The fix: a large holdout with imperfect labels beats a small one with perfect labels\n\nThe instinct is that real radiologist labels must be the better yardstick. For *selection*, they are not — the\nvariance from n=58 swamps the bias from imperfect targets. Derived labels agree with radiologists at ~0.89\n(see my [label benchmark](https://www.kaggle.com/code/starkhushi/rsna-knee-which-report-labels-should-you-train-on)),\nand you can have hundreds of studies instead of dozens.\n\nSimulate it: draw holdouts of increasing size from the derived-label pool and watch the interval collapse."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"pool_ids = train.loc[~train['StudyInstanceUID'].isin(ann['StudyInstanceUID']), 'StudyInstanceUID']\npool_ids = [u for u in pool_ids if u in tbl.index]\npool = tbl.loc[pool_ids, LABELS].apply(pd.to_numeric, errors='coerce').dropna()\nprint(f'derived-label pool available: {len(pool)} studies')\n\nsizes = [58, 100, 250, 500, 1000]\nwidths = []\nfor n in sizes:\n    m = []\n    for b in range(300):\n        idx = rng.choice(len(pool), n, replace=True)\n        sub = pool.values[idx]\n        y = (sub >= 0.5).astype(int)\n        p = sub + rng.normal(0, 0.25, sub.shape)          # a model that is good but not perfect\n        per = [roc_auc_score(y[:, j], p[:, j]) for j in range(12) if 0 < y[:, j].sum() < n]\n        m.append(np.mean(per))\n    widths.append(np.percentile(m, 97.5) - np.percentile(m, 2.5))\n\nfig, ax = plt.subplots(figsize=(7, 4))\nax.plot(sizes, widths, 'o-', color='#4477AA')\nfor x, y_ in zip(sizes, widths): ax.annotate(f'{y_:.3f}', (x, y_), textcoords='offset points', xytext=(0, 8))\nax.set_xlabel('holdout size (studies)'); ax.set_ylabel('95% CI width of macro AUC')\nax.set_title('Selection noise falls with holdout size, not with label purity')\nax.grid(alpha=0.3); plt.tight_layout(); plt.show()"},{"cell_type":"markdown","metadata":{},"source":"## 3. Building that holdout without leaking\n\nTwo traps specific to this dataset:\n\n**Duplicate reports.** 131 studies share a byte-identical report with another study (46 groups, the largest\nspanning 37 studies). Since the targets are derived *from* the report, splitting such a group across\ntrain/holdout scores the model on a target it was trained on. Keep groups whole.\n\n**The annotated 58.** Keep them out of training entirely and report on them — just do not *select* on them."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"dup_groups = train.groupby('Report').size()\nprint(f'reports shared by >1 study: {(dup_groups > 1).sum()} groups, '\n      f'{int(dup_groups[dup_groups > 1].sum())} studies, largest {dup_groups.max()}')\n\ndef make_split(train_df, label_tbl, n_holdout=250, seed=7):\n    \"\"\"Annotated studies -> report-only. Holdout drawn from studies with a UNIQUE report.\"\"\"\n    ann_ids = set(train_df[train_df[LABELS].notna().all(axis=1)]['StudyInstanceUID'])\n    ann_reports = set(train_df[train_df['StudyInstanceUID'].isin(ann_ids)]['Report'])\n    df = train_df[train_df['StudyInstanceUID'].isin(label_tbl.index)].copy()\n    df = df[~(df['Report'].isin(ann_reports) & ~df['StudyInstanceUID'].isin(ann_ids))]   # report leak\n    unique_report = ~df['Report'].duplicated(keep=False)\n    cand = df[unique_report & ~df['StudyInstanceUID'].isin(ann_ids)]\n    hold = set(cand.sample(n_holdout, random_state=seed)['StudyInstanceUID'])\n    tr = df[~df['StudyInstanceUID'].isin(hold | ann_ids)]\n    return tr['StudyInstanceUID'].tolist(), sorted(hold), sorted(ann_ids)\n\ntr_ids, ho_ids, va_ids = make_split(train, tbl)\nprint(f'train {len(tr_ids)} | selection holdout {len(ho_ids)} | annotated report-only {len(va_ids)}')\nassert not (set(tr_ids) & set(ho_ids)) and not (set(tr_ids) & set(va_ids))\nprint('no overlap between splits')"},{"cell_type":"markdown","metadata":{},"source":"## What I changed, and what happened\n\nFrom v5 onward I select checkpoints on the 250-study derived holdout and treat the 58 annotated studies as a\nreport-only sanity check. On the run where I first did this, the holdout curve rose monotonically while the\n58-study number bounced ±0.02 between epochs — the two disagreed about the best epoch three times out of eight.\n\n**Rules of thumb I would now follow in any competition with a tiny labelled set:**\n\n- Bootstrap your validation metric once, at the start. If the CI is wider than the differences you intend to\n  chase, that set cannot referee your experiments.\n- Prefer more samples over cleaner labels *for selection*. Keep the clean set for a final, honest report.\n- When a big holdout and a small clean set disagree, the big one is usually right about ranking.\n- The leaderboard is a validation set too — with hundreds of studies behind it. Sometimes the cheapest way to\n  break a tie between two models is one submission.\n\nCorrections and counter-examples welcome. If this saved you from a bad model choice, an upvote helps others\nfind it."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11"}},"nbformat":4,"nbformat_minor":5}