{"cells":[{"cell_type":"markdown","metadata":{},"source":"# Which report-label table should you train on? A benchmark against the 58 annotated studies\n\nOnly **58 of 4,407** training studies carry the twelve radiologist labels. Everyone else trains on labels\nderived from the radiology reports — and several such tables are now public. Which one you pick silently\ncaps everything downstream, yet I could not find a published comparison. So here is one.\n\n**Method:** every public table is scored by macro AUC against the 58 studies that carry real radiologist\nlabels — the same ground truth the competition scores on. No model is trained; this measures the\n*supervision*, not the network.\n\n**Results (macro AUC vs. the 58 annotated studies):**\n\n| Label source | macro AUC | coverage |\n|---|---|---|\n| `stevenleehans/rsna-knee-llm-report-labels` — v4 blend | **0.893** | 58/58 |\n| `stevenleehans/…` — v2 | 0.887 | 58/58 |\n| `pilkwang/rsna-knee-llm-labels` — report_labels_v2 | 0.870 | 57/58 |\n| a hand-written multilingual regex extractor (mine, below) | 0.711 | 58/58 |\n\nThree things worth taking away before the code:\n\n1. **The ceiling is the labels.** The best table agrees with radiologists at 0.893. A model trained on it\n   is learning a teacher that is itself ~0.89 correct, which is why single models in this competition\n   converge around 0.87-0.89 no matter the architecture.\n2. **Per-label quality is wildly uneven** — ACL is read almost perfectly (0.99), Synovitis barely better\n   than chance-plus (0.79). Your model's weakest label is likely your labels' weakest label, not your\n   network's.\n3. **Averaging two independent tables** produces a soft target that beats either alone on several labels —\n   noise in one reader cancels noise in the other.\n\n*Not affiliated with any of these dataset authors — all credit for the tables themselves is theirs.*"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import glob, os\nimport numpy as np\nimport pandas as pd\nfrom sklearn.metrics import roc_auc_score\nimport matplotlib.pyplot as plt\n\nPATH = os.path.dirname(glob.glob('/kaggle/input/**/train.csv', recursive=True)[0])\ntrain = pd.read_csv(os.path.join(PATH, 'train.csv'))\nLABELS = list(train.columns[2:14])\n\nannotated = train[train[LABELS].notna().all(axis=1)].set_index('StudyInstanceUID')\nprint(f'{len(annotated)} of {len(train)} studies carry radiologist labels '\n      f'({len(annotated)/len(train)*100:.1f}%)')\nprint(f'positives per label:')\nprint(annotated[LABELS].sum().astype(int).to_string())"},{"cell_type":"markdown","metadata":{},"source":"## The candidate tables\n\nTwo public LLM-read tables (several versions each), plus a rule-based extractor written from scratch as a\nno-dependency baseline."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"def score_table(df, name, id_col='StudyInstanceUID'):\n    df = df.set_index(id_col)\n    cols = {}\n    for c in LABELS:                          # tolerate naming differences\n        for dc in df.columns:\n            if dc.lower().replace('_', ' ').replace(chr(39), '') == c.lower().replace(chr(39), ''):\n                cols[c] = dc\n                break\n    common = annotated.index.intersection(df.index)\n    aucs = {}\n    for c in LABELS:\n        if c not in cols:\n            continue\n        y = annotated.loc[common, c].astype(int)\n        p = pd.to_numeric(df.loc[common, cols[c]], errors='coerce')\n        m = p.notna()\n        if 0 < y[m].sum() < m.sum():\n            aucs[c] = roc_auc_score(y[m], p[m])\n    return {'name': name, 'coverage': len(common), 'macro': float(np.mean(list(aucs.values()))), **aucs}\n\ncandidates = []\nfor pattern, name in [\n    ('/kaggle/input/**/llm_labels_v4_blend.csv', 'stevenleehans v4 blend'),\n    ('/kaggle/input/**/llm_labels_v2.csv',       'stevenleehans v2'),\n    ('/kaggle/input/**/llm_labels_full.csv',     'stevenleehans full'),\n    ('/kaggle/input/**/report_labels_v2.csv',    'pilkwang report_labels_v2'),\n]:\n    hits = glob.glob(pattern, recursive=True)\n    if hits:\n        candidates.append(score_table(pd.read_csv(hits[0]), name))\n\nres = pd.DataFrame(candidates).set_index('name')\nres[['coverage', 'macro']].sort_values('macro', ascending=False)"},{"cell_type":"markdown","metadata":{},"source":"## Per-label agreement — where the supervision is strong and where it is not"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"per = res[LABELS].sort_values('macro' if 'macro' in res.columns else LABELS[0]) if False else res[LABELS]\norder = per.mean().sort_values().index\nfig, ax = plt.subplots(figsize=(10, 5))\nx = np.arange(len(order)); w = 0.8 / len(per)\nfor i, (name, row) in enumerate(per.iterrows()):\n    ax.bar(x + i*w, row[order].values, w, label=name)\nax.axhline(0.5, color='k', lw=0.8, ls='--')\nax.set_xticks(x + 0.4 - w/2); ax.set_xticklabels(order, rotation=45, ha='right')\nax.set_ylabel('AUC vs radiologist labels'); ax.set_ylim(0.4, 1.02)\nax.set_title('How well each report-label table agrees with radiologists, per finding')\nax.legend(fontsize=8); ax.grid(axis='y', alpha=0.3)\nplt.tight_layout(); plt.show()\n\nprint(per.T.round(3).to_string())"},{"cell_type":"markdown","metadata":{},"source":"**Read the low bars as a ceiling, not a failure.** Synovitis sits near 0.79 in the best table: most\nreports never mention synovitis at all (I measured ~16% of the 4,407 reports contain any synovitis term in\nany of the nine languages), so the label is largely inferred rather than read. No architecture recovers a\nsignal the targets do not contain — which is why every public model I have seen scores its worst on\nSynovitis and Lateral OA."},{"cell_type":"markdown","metadata":{},"source":"## Blending two independent readers\n\nThe two dataset authors used different pipelines, so their mistakes are not the same mistakes. Averaging\nthem produces a soft target in [0, 1] that is closer to the radiologist than either alone on several labels."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"a = glob.glob('/kaggle/input/**/llm_labels_v4_blend.csv', recursive=True)\nb = glob.glob('/kaggle/input/**/report_labels_v2.csv', recursive=True)\nif a and b:\n    A = pd.read_csv(a[0]).set_index('StudyInstanceUID')[LABELS]\n    B = pd.read_csv(b[0]).set_index('StudyInstanceUID')\n    B = B[[c for c in LABELS]].apply(pd.to_numeric, errors='coerce')\n    common = A.index.intersection(B.index)\n    blend = (A.loc[common] + B.loc[common]) / 2.0\n    rows = []\n    for nm, tbl in [('A: stevenleehans v4', A.loc[common]),\n                    ('B: pilkwang v2', B.loc[common]),\n                    ('mean(A, B)', blend)]:\n        s = score_table(tbl.reset_index(), nm)\n        rows.append(s)\n    cmp = pd.DataFrame(rows).set_index('name')[['macro'] + LABELS]\n    print(cmp.round(3).T.to_string())"},{"cell_type":"markdown","metadata":{},"source":"## A from-scratch multilingual extractor, for reference\n\nIf you would rather not depend on someone else's table, the reports can be parsed directly. The catch is\nthat they arrive in about nine languages (my rough count over the 4,407: ~37% English, then Spanish, Greek,\nCyrillic-script, German, Croatian/Serbian, Turkish, Dutch, French), so a lexicon has to carry all of them at\nonce. Below is a compact version — anatomy stems crossed with pathology terms, negation scoped to the\nclause. It reaches **0.711** macro: usable, well short of the LLM tables, and included mostly to show where\nthe difficulty lies."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import re, unicodedata\n\n_PRE = str.maketrans({chr(305): 'i', chr(304): 'i', 'I': 'i', chr(223): 'ss'})\ndef norm(t):\n    t = str(t).translate(_PRE).lower()\n    t = unicodedata.normalize('NFKD', t)\n    return re.sub(r'[ \\t]+', ' ', ''.join(c for c in t if not unicodedata.combining(c)))\n\ndef clauses(t):\n    t = re.sub(r'(?<![.:;!?])\\n', ' ', t)          # unwrap hard-wrapped lines first\n    return [c.strip() for c in re.split(r'[.;:!?\\n]', t) if c.strip()]\n\nNEG = re.compile(r'\\bno\\b|\\bnot\\b|without|\\bsin\\b|ausencia|pas de|\\bsans\\b|\\bgeen\\b|\\bkeine?\\b'\n                 r'|\\bohne\\b|nicht|\\byok\\b|saptanmadi|izlenme|\\bnema\\b|\\bbez\\b|\\bnije\\b'\n                 r'|χωρις|δεν |\\bбез\\b|нема|normal\\b|unremarkable|intact|integr|preserved')\nTEAR = re.compile(r'tear|torn|ruptur|rotura|desgarro|dechir|scheur|riss\\b|yirti|ρηξη|руптур|разкъс|puknu')\nSTEMS = {\n 'ACL': r'anterior cruciate|\\bacl\\b|cruzado anterior|\\blca\\b|croise anterieur|kruisband'\n        r'|kreuzband|on capraz|prednj\\w* krizn|χιαστ|предн\\w* кръстн',\n 'Medial Meniscus': r'menisco (interno|medial)|medial menisc|menisc\\w* medial|binnenmeniscus'\n        r'|innenmeniskus|ic menisk|medijaln\\w* menisk|εσω μηνισκ|медиалн\\w* мениск',\n 'Effusion': r'effusion|derrame|erguss|epanchement|hydrops|efuzij|eff?uzyon|sivi artisi|συλλογη|излив',\n 'Synovitis': r'synovit|sinovit|synovial (thickening|proliferation)|υμενιτιδ|синовит|sinovyal',\n}\n\ndef extract(report, label):\n    rx = re.compile(STEMS[label])\n    for c in clauses(norm(report)):\n        if rx.search(c):\n            if label in ('ACL', 'Medial Meniscus'):\n                if TEAR.search(c) and not NEG.search(c):\n                    return 1.0\n            elif not NEG.search(c):\n                return 1.0\n    return 0.0\n\ndemo = {}\nfor lab in STEMS:\n    p = annotated.index.map(lambda u: extract(train.set_index('StudyInstanceUID').loc[u, 'Report'], lab))\n    y = annotated[lab].astype(int)\n    demo[lab] = roc_auc_score(y, list(p))\nprint('rule extractor, four representative labels:')\nfor k, v in demo.items():\n    print(f'   {k:18s} {v:.3f}')"},{"cell_type":"markdown","metadata":{},"source":"## What I would do with this\n\n1. **Train on the blended soft target**, not a single table — it costs nothing and cancels reader noise.\n2. **Hold out the 58 annotated studies for reporting, but do not select checkpoints on them.** Fifty-eight\n   studies give 9-35 positives per label; swings of ±0.03 are noise. In my own runs the model that ranked\n   *last* on those 58 ranked *first* on the leaderboard. Carve a few hundred extra studies with derived\n   labels as a selection holdout instead.\n3. **Keep byte-identical reports on one side of your split.** 131 studies share a report with another study\n   (46 groups, the largest spanning 37 studies); split them apart and you are scoring a model on targets it\n   was trained on.\n\nIf this saved you an experiment, an upvote helps others find it. Corrections welcome — especially if you\nhave measured a table I missed."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11"}},"nbformat":4,"nbformat_minor":5}