{"cells":[{"cell_type":"markdown","metadata":{},"source":"# Rank-mean ensembling: +0.015 AUC for free, and why probability-mean is the wrong tool\n\nThree models of mine plateaued within 0.001 of each other on the public leaderboard. Rank-averaging them\nmoved the score more than a week of architecture work had:\n\n| submission | public LB |\n|---|---|\n| v2 — single sagittal series, ResNet-34, mean-pooled slices | 0.873 |\n| v3 — three planes, per-label plane-mixture head | 0.874 |\n| v4 — ConvNeXt-T, gated slice attention, EMA weights | 0.883 |\n| **rank mean of v2 + v3 + v4** | **0.898** |\n\n**+0.015 over the best member**, from arithmetic on predictions I already had — no training, one submission.\nThis notebook explains why the gain exists, why *rank* mean rather than probability mean, when it fails, and\ngives drop-in code. Everything is CPU-only and reproducible here."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import glob, os\nimport numpy as np\nimport pandas as pd\nfrom scipy.stats import rankdata\nfrom sklearn.metrics import roc_auc_score\nimport matplotlib.pyplot as plt\n\nPATH = os.path.dirname(glob.glob('/kaggle/input/**/train.csv', recursive=True)[0])\ntrain = pd.read_csv(os.path.join(PATH, 'train.csv'))\nLABELS = list(train.columns[2:14])\nann = train[train[LABELS].notna().all(axis=1)].reset_index(drop=True)\ntruth = ann[LABELS].values.astype(int)\nprint(f'{len(ann)} annotated studies used as ground truth')"},{"cell_type":"markdown","metadata":{},"source":"## 1. The metric only reads order\n\nThis competition scores macro ROC AUC. AUC is invariant under **any strictly increasing transform** of a\nlabel's scores — multiply by 10, take a square root, apply a sigmoid twice: identical AUC. Three consequences\nfollow, and each one removes a decision:\n\n* Calibration is worth nothing. Do not spend effort on it.\n* Thresholds are worth nothing. Do not tune them.\n* **Combining models should combine orderings, not magnitudes.**\n\nThat last point is the whole argument for rank mean. Demonstration on real data below: I take one label table\nand rescale it three ways that leave every ordering intact."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"src = glob.glob('/kaggle/input/**/llm_labels_v4_blend.csv', recursive=True)[0]\ntbl = pd.read_csv(src).set_index('StudyInstanceUID')\nP = tbl.loc[ann['StudyInstanceUID'], LABELS].values.astype(float)\n\nvariants = {'raw p': P,\n            'p ** 3 (same order)': P ** 3,\n            'p * 100 (same order)': P * 100,\n            'logit-ish (same order)': np.log(np.clip(P, 1e-6, 1-1e-6) / (1 - np.clip(P, 1e-6, 1-1e-6)))}\nfor name, V in variants.items():\n    m = np.mean([roc_auc_score(truth[:, j], V[:, j]) for j in range(12) if 0 < truth[:, j].sum() < len(truth)])\n    print(f'{name:24s} macro AUC {m:.6f}')"},{"cell_type":"markdown","metadata":{},"source":"Identical to six decimals. Now watch what happens when you *average* two of those representations — the\nscale, which the metric ignores, suddenly decides the outcome."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"A = P                      # model A: well-spread probabilities\nB = P ** 6                 # model B: same ordering, but crushed toward 0 (an under-confident model)\n\ndef macro(X):\n    return np.mean([roc_auc_score(truth[:, j], X[:, j]) for j in range(12) if 0 < truth[:, j].sum() < len(truth)])\n\ndef rank_norm(X):\n    return np.stack([rankdata(X[:, j]) / len(X) for j in range(X.shape[1])], axis=1)\n\n# perturb each model differently so they are not literally identical rankings\nrng = np.random.RandomState(0)\nA_n = np.clip(A + rng.normal(0, 0.10, A.shape), 0, 1)\nB_n = np.clip(B + rng.normal(0, 0.02, B.shape), 0, 1)\n\nprint(f'model A alone           {macro(A_n):.4f}')\nprint(f'model B alone           {macro(B_n):.4f}')\nprint(f'probability mean        {macro((A_n + B_n) / 2):.4f}')\nprint(f'RANK mean               {macro((rank_norm(A_n) + rank_norm(B_n)) / 2):.4f}')"},{"cell_type":"markdown","metadata":{},"source":"Probability mean lets whichever model happens to output a wider numeric range dominate the sum — a property\nthe metric never asked about. Rank mean puts every member on the same [0, 1] scale by construction, so each\none contributes its *opinion about ordering* and nothing else.\n\n**Rule:** for any rank-based metric (ROC AUC, MAP, NDCG), average ranks. For a metric that reads magnitudes\n(log loss, RMSE, Brier), average the values — and calibrate."},{"cell_type":"markdown","metadata":{},"source":"## 2. Why blending helps at all: members must fail differently\n\nAn ensemble only gains where members disagree. My three models each *won* different findings — and the blend\nbeat all three on the labels where they split:\n\n| finding | v2 | v3 | v4 | rank mean |\n|---|---|---|---|---|\n| Effusion | 0.771 | 0.829 | **0.893** | 0.850 |\n| MCL | 0.789 | **0.846** | 0.844 | **0.865** |\n| Fracture | 0.812 | 0.843 | 0.835 | **0.869** |\n| Contusion | 0.881 | 0.893 | 0.892 | **0.909** |\n| Medial OA | 0.912 | 0.941 | 0.935 | **0.946** |\n| ACL | **0.963** | 0.926 | 0.891 | 0.945 |\n\nNote ACL: the blend is *worse* than v2 alone there, because v2 was genuinely better and the other two dragged\nit down. That is the trade — you lose a little where one member is clearly best, and gain more where none is.\nAcross twelve labels it netted +0.015.\n\nCorrelation is the diagnostic. Members correlated above ~0.95 add almost nothing; the useful range is roughly\n0.7-0.9 — different enough to disagree, good enough to be worth listening to."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"# how much diversity is enough? simulate members with a controlled correlation to the truth-signal\ndef make_member(base, noise, seed):\n    r = np.random.RandomState(seed)\n    return np.clip(base + r.normal(0, noise, base.shape), 0, 1)\n\nbase = P\nrows = []\nfor noise in [0.05, 0.15, 0.30]:\n    for n_members in [1, 2, 3, 5, 8]:\n        ms = [make_member(base, noise, s) for s in range(n_members)]\n        blended = np.mean([rank_norm(m) for m in ms], axis=0)\n        corr = np.mean([np.corrcoef(ms[0][:, j], ms[-1][:, j])[0, 1] for j in range(12)]) if n_members > 1 else 1.0\n        rows.append({'member noise': noise, 'members': n_members,\n                     'mean pairwise corr': round(corr, 3), 'macro AUC': round(macro(blended), 4)})\ndiv = pd.DataFrame(rows)\nprint(div.pivot(index='members', columns='member noise', values='macro AUC').to_string())\n\nfig, ax = plt.subplots(figsize=(7, 4))\nfor noise, g in div.groupby('member noise'):\n    ax.plot(g['members'], g['macro AUC'], 'o-', label=f'member noise {noise}')\nax.set_xlabel('number of members'); ax.set_ylabel('macro AUC of rank mean')\nax.set_title('Diminishing returns: most of the gain arrives by 3-5 members')\nax.legend(); ax.grid(alpha=0.3); plt.tight_layout(); plt.show()"},{"cell_type":"markdown","metadata":{},"source":"Most of the benefit lands by three to five members, and noisier (more diverse) members benefit more from\nbeing blended — which is why adding a *weaker but different* model often helps while adding a near-copy of\nyour best one does not."},{"cell_type":"markdown","metadata":{},"source":"## 3. Drop-in code\n\nRank-normalise per label, average, write. That is the entire technique."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"def rank_blend(prediction_frames, weights=None):\n    \"\"\"prediction_frames: list of DataFrames indexed by StudyInstanceUID, one column per label.\n\n    Ranks are computed per label per member, so members never need to share a calibration.\n    \"\"\"\n    ids = prediction_frames[0].index\n    cols = list(prediction_frames[0].columns)\n    if weights is None:\n        weights = [1.0] * len(prediction_frames)\n    weights = np.array(weights, dtype=float); weights /= weights.sum()\n\n    out = np.zeros((len(ids), len(cols)))\n    for w, df in zip(weights, prediction_frames):\n        arr = df.loc[ids, cols].values.astype(float)\n        out += w * np.stack([rankdata(arr[:, j]) / len(arr) for j in range(arr.shape[1])], axis=1)\n    return pd.DataFrame(out, index=ids, columns=cols)\n\n# toy demonstration with three perturbed members\nmembers = [pd.DataFrame(make_member(P, 0.15, s), index=ann['StudyInstanceUID'], columns=LABELS)\n           for s in (1, 2, 3)]\nblend = rank_blend(members)\nprint('members :', [round(macro(m.values), 4) for m in members])\nprint('blend   :', round(macro(blend.values), 4))"},{"cell_type":"markdown","metadata":{},"source":"## Practical notes and honest caveats\n\n* **Equal weights are a strong default.** I measured a tuned weighting at 0.8779 versus 0.8769 for equal —\n  a 0.001 difference on 58 studies, which is noise. Tuned weights on a small validation set are a good way to\n  overfit; I submitted equal.\n* **A weaker member can still help.** My fourth model scores lower alone but reads Synovitis at 0.78 where\n  the others manage 0.62-0.72. Blending is about coverage, not about a leaderboard of members.\n* **Decode once, predict many.** My three members need three different preprocessings; reading each study's\n  DICOMs once and transforming three ways kept inference inside the runtime limit.\n* **This is not free of risk.** Blending models tuned against the public leaderboard can inherit that\n  overfitting. Blend models selected on your own holdout instead.\n\nIf this is useful, an upvote helps others find it — and I would genuinely like to hear counter-results where\nprobability mean beat rank mean on an AUC metric."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11"}},"nbformat":4,"nbformat_minor":5}