{"metadata":{"kernelspec":{"display_name":"d2l","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.10.20"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"0e2a60d9","cell_type":"markdown","source":"# Weak Report Labels: A Path to a Better Score (Pending Validation)\n\n![Overview](https://gastigado.cnies.org/d/public/ds.png)\n\nOf the 4,407 training studies, only **58 carry gold labels**; the rest are images + multilingual reports.\nEveryone builds weak labels from the reports to train a vision model. This notebook lays out a\n**pending-validation** path toward a better score:\n\n> **Multi-model report extraction → disagreement arbitration → conf into training → end-to-end validation**\n\n---\n\n> **Credits & attribution:** This notebook builds entirely on\n> [Pilkwang Kim's rsna-knee-baseline-v1](https://www.kaggle.com/code/pilkwang/rsna-knee-baseline-v1)\n> and its LLM-label dataset [rsna-knee-llm-labels](https://www.kaggle.com/datasets/pilkwang/rsna-knee-llm-labels):\n> - The extractor's **prompt, schema, verdict/severity definitions** come verbatim from Pilkwang's `api_labeler.py`\n> - The extractor in Section 3 is **an adaptation of Pilkwang's `api_labeler.py`** (only the LLM endpoint changed; the logic is unchanged)\n> - `report_labels_v2.csv` in the 58-study comparison is the Claude LLM labels Pilkwang released\n> - Our addition is re-running four LLMs and proposing the disagreement-ensemble / conf path, all within Pilkwang's framework","metadata":{}},{"id":"21f091b9-39cf-4977-890d-862c2a4d9a99","cell_type":"markdown","source":"## 1. Hypothesis: will better report extraction improve the score?\n\n**Why it might help:**\n- Training relies entirely on weak labels. Weak labels are **noisy targets** — less noise means cleaner signal.\n- On the 58 gold studies, report-vs-label agreement is only ~82.5% (labels are image-derived).\n  The closer an extractor gets to that ceiling, the more trustworthy the weak labels.\n\n**But \"better\" must be validated end-to-end:**\nExtraction AUC → training → final macro AUC, with **attenuation in between**.\nA 0.005 extraction gap on 58 studies is noise (nekkon: σ=0.0125). To confirm a score gain,\nyou must **swap weak labels → train the same vision model → compare final macro AUC**,\nnot just look at extraction scores.","metadata":{}},{"id":"6ef82116-8142-4c50-8773-b23c93ff1185","cell_type":"markdown","source":"## 2. 58-study comparison: which extractor is stronger\n\nFive extractors are aligned on the same gold studies. Each model's scores come from local caches (the audit CSV).\n`report_labels_v2.csv` is Pilkwang's Claude LLM labels; `gold58_audit.csv` is our re-run of four LLMs.","metadata":{}},{"id":"85df5acf-d504-4c5a-913f-4a622abe6f94","cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom pathlib import Path\nfrom sklearn.metrics import roc_auc_score\n\n# Data mount paths (adjust to your actual mount layout)\nCOMP = Path(\"/kaggle/input/competitions/rsna-knee-abnormality-detection\")            # competition data\nGOLD_AUDIT = Path(\"/kaggle/input/datasets/cjlcjlcjl/gold58-audit\")      # audit CSV (5 models)\nPILKWANG = Path(\"/kaggle/input/datasets/pilkwang/rsna-knee-llm-labels\")  # Pilkwang LLM labels\n\ntrain = pd.read_csv(COMP / \"train.csv\")\naudit = pd.read_csv(GOLD_AUDIT / \"gold58_audit.csv\")\n\nTARGETS = [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\",\n           \"Medial OA\", \"Lateral OA\", \"PF OA\", \"Effusion\", \"Synovitis\",\n           \"Baker's\", \"Contusion\", \"Fracture\"]\ngold = train.dropna(subset=TARGETS).reset_index(drop=True)\nprint(f\"gold studies: {len(gold)}\")\n\n# Pilkwang CSV coverage + the missing study\ncsv = pd.read_csv(PILKWANG / \"report_labels_v2.csv\")\nmissing_pilkwang = set(gold.StudyInstanceUID) - set(csv.StudyInstanceUID)\nprint(f\"Pilkwang CSV missing on gold: {len(missing_pilkwang)} study(ies)\")\n\ndef macro_auc(Y, S):\n    aucs = []\n    for j in range(len(TARGETS)):\n        if Y[:, j].sum() in (0, len(Y)):\n            continue\n        try:\n            aucs.append(roc_auc_score(Y[:, j], S[:, j]))\n        except ValueError:\n            pass\n    return float(np.mean(aucs)) if aucs else float(\"nan\")\n\nY = gold[TARGETS].values.astype(int)\nprint(\"\\n=== coverage + macro AUC (each on its full subset) ===\")\nprint(f\"  {'model':20s} {'cov':>6s} {'macroAUC':>9s}\")\nfor m, disp in [(\"glm\",\"GLM\"),(\"ds_pro\",\"DS pro\"),\n                (\"ds_flash\",\"DS flash (thinking)\"),(\"ds_flash_nothink\",\"DS flash (no-think)\"),\n                (\"minimax\",\"MiniMax\")]:\n    cols = [f\"{m}_{t}\" for t in TARGETS]\n    S = audit[cols].values\n    keep = ~np.isnan(S).any(axis=1)\n    a = macro_auc(Y[keep], S[keep])\n    print(f\"  {disp:20s} {keep.sum():>4d}/58 {a:>9.4f}\")\n\ncsvv = csv.set_index(\"StudyInstanceUID\").reindex(gold.StudyInstanceUID)[TARGETS].values.astype(float)\nkeep = ~np.isnan(csvv).any(axis=1)\na = macro_auc(Y[keep], csvv[keep])\nprint(f\"  {'Pilkwang CSV (Claude)':20s} {keep.sum():>4d}/58 {a:>9.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-09T18:50:05.312825Z","iopub.execute_input":"2026-08-09T18:50:05.313136Z","iopub.status.idle":"2026-08-09T18:50:05.57818Z","shell.execute_reply.started":"2026-08-09T18:50:05.313107Z","shell.execute_reply":"2026-08-09T18:50:05.57744Z"}},"outputs":[],"execution_count":null},{"id":"93ebe290-041e-4170-9b8a-c2d91c8c062c","cell_type":"markdown","source":"## 3. LLM extractor (adapted from Pilkwang's `api_labeler.py`)\n\nThe complete extractor follows. The **prompt, 12-key schema, verdict/severity definitions, and\nbatch reading logic all come from Pilkwang's `api_labeler.py`**. We made two adaptations:\n1. Model switched to **DeepSeek v4-flash in non-thinking mode**\n2. Dropped `output_config` (Anthropic-compatible endpoints do not support JSON schema;\n   enforced via prompt + tolerant parsing instead)\n\n**This cell requires a local API key to run** (no key in the Kaggle environment; skip it there).","metadata":{}},{"id":"63191f49-f231-46b2-84ab-cbdf484abfe0","cell_type":"code","source":"\"\"\"Adapted from pilkwang/rsna-knee-llm-labels/api_labeler.py (LLM extractor).\n\nThe original used Claude opus-5 (Anthropic SDK). Here we use DeepSeek v4-flash in\nnon-thinking mode; the rest (prompt/schema/verdict→score mapping/batch) is unchanged.\n\"\"\"\nimport json\nimport os\nimport threading\nfrom pathlib import Path\n\n# ---- snake_case keys for the 12 findings (order matches TARGETS) ----\nKEYS = [\"acl\", \"mcl\", \"medial_meniscus\", \"lateral_meniscus\",\n        \"medial_oa\", \"lateral_oa\", \"pf_oa\", \"effusion\",\n        \"synovitis\", \"bakers\", \"contusion\", \"fracture\"]\n\nTARGETS = [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\",\n           \"Medial OA\", \"Lateral OA\", \"PF OA\", \"Effusion\", \"Synovitis\",\n           \"Baker's\", \"Contusion\", \"Fracture\"]\n\n# ---- SYSTEM prompt (verbatim from Pilkwang's api_labeler.py) ----\nFINDING_SPEC = \"\"\"For each finding below, decide what THE REPORT asserts:\n- ACL: tear or rupture of the anterior cruciate ligament\n- MCL: tear or sprain of the medial collateral ligament\n- medial_meniscus: tear of the medial meniscus\n- lateral_meniscus: tear of the lateral meniscus\n- medial_oa / lateral_oa / pf_oa: osteoarthritis (cartilage loss) in that compartment\n- effusion: joint effusion\n- synovitis: inflammation or thickening of the synovial lining\n- bakers: Baker (popliteal) cyst\n- contusion: bone contusion / bone marrow oedema from impact\n- fracture: acute fracture or cortical break\"\"\"\n\nSYSTEM = \"\"\"You are an expert musculoskeletal radiologist reading a knee MRI report.\nReports may be written in any language, including Greek, Bulgarian, Turkish, Croatian,\nDutch, German, Spanish, French or English. Read each one in its original language.\n\nFor each of the twelve findings below, state what THE REPORT IN FRONT OF YOU says.\n{spec}\nverdict: YES the report asserts it; NO the report denies it; UNK it is not mentioned.\nseverity (0 unless YES): 1 minimal/mild, 2 moderate, 3 marked/severe/full-thickness.\nRespond with JSON only: one object with all twelve keys, each {{\"verdict\",\"severity\"}}.\n\"\"\".format(spec=FINDING_SPEC)\n\n# ---- tolerant parsing (model output keys are not consistently cased / structured) ----\nimport re as _re\n\ndef parse_obj(text: str) -> dict:\n    \"\"\"Parse a 12-key object from model output. Handles mixed-case keys / ```json /\n    array-wrapped / collapsed keys.\n\n    Anthropic-compatible endpoints do not enforce a JSON schema, so the model may output:\n      - mixed-case keys (ACL / medial_meniscus)\n      - the 12 keys wrapped in an array [{...}]\n      - keys collapsed to a single letter (e.g. all \"l\")\n    Falls back layer by layer, finally scraping \"KEY: verdict\" pairs with regex.\n    \"\"\"\n    t = text.strip()\n    if t.startswith(\"```\"):\n        m = _re.match(r\"```(?:json)?\\s*(.*?)\\s*```\", t, _re.S)\n        if m:\n            t = m.group(1).strip()\n    try:\n        obj = json.loads(t)\n        if isinstance(obj, list) and len(obj) == 1 and isinstance(obj[0], dict):\n            obj = obj[0]\n        if isinstance(obj, dict):\n            alias = {k.lower(): k for k in KEYS}\n            norm = {}\n            for k, v in obj.items():\n                key = k.lower().replace(\"'\", \"\").replace(\" \", \"_\").replace(\"-\", \"_\")\n                if key in alias:\n                    norm[alias[key]] = v\n            if all(k in norm for k in KEYS):\n                return norm\n            vals = list(obj.values())\n            if len(vals) == 12 and all(isinstance(v, dict) and \"verdict\" in v for v in vals):\n                return {k: {\"verdict\": v[\"verdict\"], \"severity\": v.get(\"severity\", 0)}\n                        for k, v in zip(KEYS, vals)}\n    except json.JSONDecodeError:\n        pass\n    # whole-text regex fallback\n    out = {}\n    for k in KEYS:\n        pat = rf'\"?{_re.escape(k)}\"?(?!\\w)\\s*:\\s*\\{{[^}}]*?\"verdict\"\\s*:\\s*\"(YES|NO|UNK)\"[^}}]*?\"severity\"\\s*:\\s*([0-3])[^}}]*\\}}'\n        m = _re.search(pat, text, _re.S | _re.I)\n        if not m:\n            pat2 = rf'\"?{_re.escape(k)}\"?(?!\\w)\\s*:\\s*\\{{[^}}]*?\"severity\"\\s*:\\s*([0-3])[^}}]*?\"verdict\"\\s*:\\s*\"(YES|NO|UNK)\"[^}}]*\\}}'\n            m = _re.search(pat2, text, _re.S | _re.I)\n        if m:\n            if m.group(1) in (\"YES\", \"NO\", \"UNK\"):\n                out[k] = {\"verdict\": m.group(1), \"severity\": int(m.group(2))}\n            else:\n                out[k] = {\"verdict\": m.group(2), \"severity\": int(m.group(1))}\n    if all(k in out for k in KEYS):\n        return out\n    raise ValueError(f\"could not parse a 12-key object from output:\\\\n{text[:400]}\")\n\n\n# ---- verdict/severity -> score (5 tiers reverse-engineered from Pilkwang's CSV) ----\n_YES_SCORE = {1: 0.68, 2: 0.82, 3: 0.94}\ndef to_scores(verdicts):\n    \"\"\"[(verdict, severity), ...] -> (score, conf). UNK=0.28, NO=0.08.\"\"\"\n    sc, cf = [], []\n    for v, s in verdicts:\n        if v == \"YES\":\n            sc.append(_YES_SCORE.get(s, 0.82)); cf.append(0.95)\n        elif v == \"NO\":\n            sc.append(0.08); cf.append(0.85)\n        else:\n            sc.append(0.28); cf.append(0.05)\n    return sc, cf\n\n\nclass ReportLabeler:\n    \"\"\"Anthropic-compatible endpoint extractor (MiniMax / DeepSeek).\"\"\"\n\n    def __init__(self, model=\"deepseek-v4-flash\", base_url=None,\n                 cache_path=None, max_tokens=2000):\n        import anthropic\n        self.client = anthropic.Anthropic(\n            api_key=os.environ[\"ANTHROPIC_API_KEY\"],\n            base_url=base_url,                 # e.g. https://api.deepseek.com/anthropic\n        )\n        self.model = model\n        self.max_tokens = max_tokens\n        self.cache_path = Path(cache_path) if cache_path else None\n        self.cache = {}\n        self._lock = threading.Lock()\n\n    def _params(self, report):\n        return dict(\n            model=self.model,\n            max_tokens=self.max_tokens,\n            thinking={\"type\": \"disabled\"},     # non-thinking: ~25x faster, reads bilateral reports too\n            system=[{\"type\": \"text\", \"text\": SYSTEM,\n                     \"cache_control\": {\"type\": \"ephemeral\"}}],\n            messages=[{\"role\": \"user\", \"content\": report.strip()[:12000]}],\n        )\n\n    def label(self, uid, report):\n        if uid in self.cache:\n            return self.cache[uid]\n        resp = self.client.messages.create(**self._params(report))\n        text = next(b.text for b in resp.content if b.type == \"text\")\n        obj = parse_obj(text)                  # tolerant parse: {key: {\"verdict\",\"severity\"}}\n        verdicts = [(obj[k][\"verdict\"], obj[k][\"severity\"] if obj[k][\"verdict\"] == \"YES\" else 0)\n                    for k in KEYS]\n        sc, cf = to_scores(verdicts)\n        rec = {\"uid\": uid, \"score\": sc, \"conf\": cf,\n               \"verdict\": [v for v, _ in verdicts],\n               \"severity\": [s for _, s in verdicts]}\n        self.cache[uid] = rec\n        if self.cache_path:\n            with open(self.cache_path, \"a\", encoding=\"utf-8\") as fh:\n                fh.write(json.dumps(rec, ensure_ascii=False) + \"\\n\")\n        return rec\n\n    def label_many(self, items, workers=8, progress=None):\n        from concurrent.futures import ThreadPoolExecutor\n        done = [0]\n        def one(it):\n            r = self.label(*it)\n            done[0] += 1\n            if progress and done[0] % progress == 0:\n                print(f\"  {done[0]}/{len(items)}\", flush=True)\n            return r\n        with ThreadPoolExecutor(workers) as ex:\n            return list(ex.map(one, items))\n\n\n# ---- Usage: full 4,407 studies, DeepSeek flash non-thinking ----\n# labeler = ReportLabeler(model=\"deepseek-v4-flash\",\n#                         base_url=\"https://api.deepseek.com/anthropic\",\n#                         cache_path=\"labels.jsonl\")\n# items = [(uid, report_text) for uid, report_text in train[[\"StudyInstanceUID\",\"Report\"]]]\n# recs = labeler.label_many(items, workers=32)","metadata":{},"outputs":[],"execution_count":null},{"id":"5f06d290-cd11-4c61-9466-294059a69a61","cell_type":"markdown","source":"## 4. Next steps: disagreement ensemble + conf into training\n\n**No single extractor wins on every label** — each excels somewhere (e.g. one wins ACL/Fracture,\nanother Effusion/Contusion, another Medial OA, another MCL/Synovitis). **Disagreement carries information.**\n\n### 4.1 Disagreement arbitration\nHave multiple models read the same report independently:\n- **Consensus** (verdicts agree) → high confidence, take the median score\n- **Disagreement** (verdicts differ) → low confidence, lower the conf, or send to an arbiter\n\nThis is more likely to improve the score than picking a single \"strongest\" model, because it\ncombines each model's strengths instead of betting on one.\n\n### 4.2 conf into training\nThe extractor already outputs `__conf` (Pilkwang's UNK rows have only 0.05).\nFeed it into the loss: **high weight on consensus, low weight on disagreement** — telling the\nvision model \"these labels are trustworthy, those are not.\" This is how extraction quality is\nactually *transferred* into training.\n\n### 4.3 Validation (the only credible way)\nNot just extraction AUC on 58 studies, but **end-to-end**:\n\nTrain the same vision model on two weak-label sets and compare final macro AUC:\n- **Baseline weak labels**: what we already use — scores from a single model (e.g. DeepSeek flash)\n  reading all reports. Think of it as \"the status quo, no improvements.\"\n- **Disagreement-ensemble weak labels**: scores after §4.1's multi-model arbitration —\n  multiple models read the same report; consensus is trusted, disagreement is down-weighted.\n  Think of it as \"the improved version.\"\n\n> Controlled variable: the vision model, data, and training settings are identical;\n> **the only difference is the weak-label source**.\n> If the ensemble version scores higher → the improvement is real; if it is about the same,\n> the ensemble gain was attenuated away during training.\n\n> **Status: this end-to-end validation is not finished yet — experiments are running,\n> results expected tomorrow (2026-08-11).**\n> Until then, all of the above are **hypotheses** — but hypotheses with direction and a clear\n> validation plan are exactly what make them worth proposing.","metadata":{}},{"id":"d424e89c-164d-4925-bc3f-a651258b9109","cell_type":"markdown","source":"## 5. Summary\n\n1. **Whether extraction improves the score must be validated end-to-end**, not by the 58-study extraction AUC alone.\n2. The 58-study comparison shows: four LLMs cover 58/58, AUC 0.870-0.878.\n3. **The real opportunity is disagreement ensemble + conf into training**, not picking one \"strongest model\".\n4. **Next step: run the full report set, and build conf.**\n   - **Run the full report set**: apply the extractor (e.g. DeepSeek flash non-thinking) to all 4,407 reports,\n     producing a score for every label of every study — not just the 58 gold ones.\n   - **Build conf**: each label of each report carries a confidence (`__conf`) saying \"how much to trust\n     this judgement\" — low where models disagree, high where they agree. When training the vision model,\n     weight labels by conf so the model trusts high-confidence labels and is not led astray by low-confidence ones.","metadata":{}},{"id":"ccb9a1ed-5954-4a10-9c29-168d0ac3ae0d","cell_type":"code","source":"print('Done.')","metadata":{},"outputs":[],"execution_count":null}]}