{"cells":[{"cell_type":"markdown","metadata":{},"source":"# RSNA Knee — Prvsiyan V9 guarded reproduction\n\nThis zero-output release is adapted from **Prvsiyan**, [*RSNA Knee: read the report,\nthen the knee*, V9](https://www.kaggle.com/code/prvsiyan/rsna-knee-read-the-report-then-the-knee?scriptVersionId=340808178),\nreleased under the Apache License 2.0. The linked upstream submission is `55332748`\n(Public AUC `0.847`). The frozen upstream notebook SHA-256 is\n`37d6c51268dcab2b530701331db2f140971c841c66e2fb012b221949471ed460`.\n\nConservative co-attribution is also given to **Pilkwang Kim**, [*RSNA Knee baseline v1*,\nV14](https://www.kaggle.com/code/pilkwang/rsna-knee-baseline-v1?scriptVersionId=340738955),\nbecause independent review found substantial contiguous source overlap. This attribution\ndoes not assert an undocumented fork relationship. Pilkwang V14 is Apache-2.0; its frozen\nnotebook SHA-256 is\n`67e874fa121b2f163a090bf598815f00441789e1699772c78f96eaa1f5ba60be`.\n\nThe learned method, five-arm configuration, seeds, epochs, preprocessing, report-label\nextractor, DINOv2 loading, holdout selection, and rank ensemble are preserved. Added\nsafety checks require the requested T4 environment, complete stages and arms, finite\npredictions, exact identifiers and schema, and atomic publication of `submission.csv`.\nThe notebook schema is nbformat 4.4 because the upstream cells have no ids.\n\nReport-like examples in markdown, comments, and docstrings are synthetic or abstract\nparaphrases. No competition report row is reproduced. The executable multilingual regex\nlexicon and extractor computations are unchanged. DINOv2 Small/Base remain pinned Kaggle\nmodel dependencies and are not redistributed.\n"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"# Candidate-only environment and wall-clock preflight. This runs before data I/O.\nimport os as _audit_os\nimport signal as _audit_signal\nimport time as _audit_time\nfrom pathlib import Path as _AuditPath\n\nimport torch as _audit_torch\n\n_AUDIT_STARTED = _audit_time.monotonic()\n_AUDIT_LIMIT_SECONDS = int(8.5 * 3600)\n\n\ndef _audit_timeout(_signum, _frame):\n    for _path in (_AuditPath(\"submission.csv\"), _AuditPath(\"_submission_candidate.csv\")):\n        _path.unlink(missing_ok=True)\n    raise TimeoutError(\"candidate global 8.5-hour deadline reached\")\n\n\n_audit_signal.signal(_audit_signal.SIGALRM, _audit_timeout)\n_audit_signal.alarm(_AUDIT_LIMIT_SECONDS)\nassert _audit_os.environ.get(\"SMOKE\", \"0\") in (\"\", \"0\"), \"SMOKE must be disabled\"\nassert int(_audit_os.environ.get(\"EPOCHS\", \"16\")) == 16, \"EPOCHS override rejected\"\nassert abs(float(_audit_os.environ.get(\"TIME_BUDGET_H\", \"7.6\")) - 7.6) < 1e-9\nassert _audit_torch.cuda.is_available(), \"CUDA is unavailable\"\n_audit_name = _audit_torch.cuda.get_device_name(0)\n_audit_capability = _audit_torch.cuda.get_device_capability(0)\nassert \"T4\" in _audit_name.upper(), f\"Tesla T4 required, got {_audit_name}\"\nassert _audit_capability == (7, 5), f\"sm_75 required, got {_audit_capability}\"\nassert \"sm_75\" in _audit_torch.cuda.get_arch_list(), _audit_torch.cuda.get_arch_list()\n\n_audit_layer = _audit_torch.nn.Linear(8, 2, device=\"cuda\")\n_audit_opt = _audit_torch.optim.AdamW(_audit_layer.parameters(), lr=1e-3)\n_audit_x = _audit_torch.randn(4, 8, device=\"cuda\")\n_audit_loss = _audit_layer(_audit_x).square().mean()\n_audit_opt.zero_grad(set_to_none=True)\n_audit_loss.backward()\n_audit_opt.step()\n_audit_torch.cuda.synchronize()\ndel _audit_layer, _audit_opt, _audit_x, _audit_loss\n_audit_torch.cuda.empty_cache()\nprint(f\"candidate preflight PASS: {_audit_name} sm_{_audit_capability[0]}{_audit_capability[1]}; \"\n      \"forward/backward/optimizer; Internet disabled by pinned kernel metadata\")\n"},{"cell_type":"markdown","metadata":{},"source":"# Twelve findings, fifty-eight labels, four thousand reports\n\nA knee MRI study here has to be given twelve probabilities — anterior cruciate and medial\ncollateral ligament injury, medial and lateral meniscal tear, osteoarthritis in each of\nthe three compartments, effusion, synovitis, Baker's cyst, bone contusion, fracture — and\nthe score is the unweighted mean of the twelve ROC AUCs.\n\nThe decisive fact about this dataset is not in the images. It is in `train.csv`:\n\n| | studies | carry the twelve labels | carry a radiology report |\n|---|---|---|---|\n| train | 4 407 | **58** | 4 407 |\n| test | — | — | **none — there is no `Report` column** |\n\nFifty-eight labelled studies cannot train an imaging model, and the reports that could\nsupply the missing four thousand are unavailable at prediction time. So the pipeline is\nforced into one shape: **read the reports into targets, fit a pure imaging model against\nthem, and throw the text away.** Everything downstream — which slices to decode, how large\nto make them, which encoder to adapt — is bounded by how well that first step is done.\n\nThis notebook is organised in the order the decisions constrain each other.\n\n1. What the score rewards, and what that removes from the design.\n2. Where the targets come from — a nine-language report reader, and the two ways to tell\n   whether it works when only 58 studies can check it.\n3. What the scanner recorded — recovering the acquisition protocol from DICOM headers.\n4. Geometry — slice order, physical scale, and which knee was scanned.\n5. Reading the pixels once.\n6. Turning six views into twelve decisions.\n7. Validating without fooling yourself.\n8. Three arms, combined the only way the metric permits.\n\nEvery number quoted below is computed by the cell above it. Nothing is asserted that the\nnotebook does not measure."},{"cell_type":"markdown","metadata":{},"source":"## 1. What the score rewards\n\n$$\\text{Score} \\;=\\; \\frac{1}{12}\\sum_{i=0}^{11} \\mathrm{AUC}_i$$\n\nThree consequences follow directly, and each one removes a design choice rather than\nadding one.\n\n**Only order matters.** $\\mathrm{AUC}_i$ is invariant under any strictly increasing map of\nthe scores for label $i$. Calibration is worth nothing here, and so is any fixed\nthreshold. It also settles how to combine models: averaging raw probabilities lets\nwhichever arm happens to be most confident dominate, while averaging *ranks* combines the\nonly information the metric reads. Every combination in §8 is a rank mean.\n\n**Every label costs the same.** Write $M$ for the mean AUC a good model reaches. A label\nleft at chance contributes $0.5$ instead of roughly $M$, forfeiting\n\n$$\\frac{M-0.5}{12}$$\n\nno matter how well the other eleven do. At $M=0.85$ that is $0.029$ — larger than the gap\nbetween neighbouring places on a mature leaderboard. So a rare finding deserves *more*\nattention than a common one, because a rare finding is where a model most easily ends up\nat chance. That is why §2 spends its effort on the findings the reports mention least,\nnot on the ones they mention most.\n\n**Prevalence drift is survivable; thresholds are not.** AUC is, in expectation, invariant\nto the positive rate, and the data description warns that prevalence need not match across\nthe training, public and private sets. That would be fatal for an accuracy-like metric. It\nis not fatal here — but it does mean one cut has to be named and watched. The report-derived\ntargets below are graded, not binary; to compute an AUC against them at all they are\nbinarised at their midpoint for validation only. That cut chooses epochs and arms. It never\nreaches a submitted score."},{"cell_type":"markdown","metadata":{},"source":"## 2. Where the targets come from\n\nEvery training study carries the radiology report written when it was read, and the data\ndescription invites deriving labels from it. The structural fact that decides how is in the\nschemas rather than in the prose: `train.csv` has a `Report` column and `test.csv` does\nnot. Text is available when fitting and absent when predicting. That rules out a fusion\nmodel with a text branch — at inference it would have nothing to read — and leaves the\nreports doing two jobs:\n\n1. supplying the training targets, and\n2. saying how confidently each one could be read, which becomes a per-finding sample\n   weight.\n\nBoth matter. A study whose report never mentions synovitis should pull on the synovitis\nhead far more weakly than one that names it, and a loss that cannot express that trains\nevery silence as a confident negative.\n\n### Nine languages, one lexicon\n\nThe corpus is written in nine languages across three scripts. The extractor is built here\nin-line rather than attached as a file, for two reasons: it runs over the whole corpus in\nseconds, so there is nothing to save by precomputing; and a notebook that ships its own\nlabels cannot quietly train against a stale copy of them.\n\n**No language is identified.** Every cue list below carries all nine at once and each\nclause is tested against the union. Routing first would mean committing to a guess before\nany evidence is read, and the cheap guess fails badly here — `la` is as common in Spanish\nas in French, and whichever substring test runs first swallows both. Pooling costs little\nin exchange: Greek and Cyrillic cues cannot collide with Latin-script ones at all, and\namong the Latin-script languages the vocabulary of interest is close enough that a shared\ncue is usually right. The price is paid in coverage instead — a phrasing no listed language\ncontributes stays unmatched — and coverage is the thing §2.2 measures.\n\n**Normalise, then unwrap, then segment, then scope.** Case, diacritics and separators are\nfolded first, which also repairs a real codepoint problem: many Greek reports spell mu with\nthe MICRO SIGN `U+00B5` rather than `U+03BC`, and NFKD maps one onto the other. Then the\nlines are unwrapped — see below — then split into clauses, with a heading line attached to\nthe value beneath it, because a synthetic example with a finding heading on one line and a negative\nresult on the next still states one thing across two lines. All narrative\nreport examples in this derivative are synthetic paraphrases, not corpus rows.\n\n**Assertion, negation, hedge.** Negation is not an edge case here. For several findings\n*most* mentions are negative, because a report lists what was checked and found intact.\nExplicit normality counts as absence: a synthetic sentence saying that the cruciate and\ncollateral ligaments are within normal limits is evidence of absence, not absence of\nevidence, except where a tear or high grade is named in the same breath.\n\n**Grade, do not threshold.** §1 established that only order is read. So a mention is scored\nby its emphasis — trace, unqualified, marked — and never binarised."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"\"\"\"Report -> twelve graded targets, in nine languages.\n\nv2. Differences from the public lexicon are all coverage: the compartment scoping for\nosteoarthritis, the pathology vocabulary for cartilage, an asserted-negative path for\nfindings a report explicitly clears, and a backoff for synovitis, which most reports\nnever name.\n\"\"\"\nfrom __future__ import annotations\n\nimport re\nimport unicodedata\n\nTARGETS = [\n    \"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\",\n    \"Medial OA\", \"Lateral OA\", \"PF OA\", \"Effusion\",\n    \"Synovitis\", \"Baker's\", \"Contusion\", \"Fracture\",\n]\n\n_PRE = str.maketrans({\n    \"ı\": \"i\", \"İ\": \"i\", \"I\": \"i\", \"ß\": \"ss\", \"đ\": \"d\", \"Đ\": \"d\",\n    \"ø\": \"o\", \"Ø\": \"o\", \"æ\": \"ae\", \"Æ\": \"ae\",\n})\n\n\ndef normalize(text: str) -> str:\n    if not isinstance(text, str):\n        return \"\"\n    text = text.translate(_PRE).lower()\n    text = unicodedata.normalize(\"NFKD\", text)\n    text = \"\".join(ch for ch in text if not unicodedata.combining(ch))\n    text = text.replace(\"­\", \"\")\n    text = re.sub(r\"[_\\-/\\\\]+\", \" \", text)\n    text = re.sub(r\"[ \\t]+\", \" \", text)\n    return text\n\n\n_SENT_SPLIT = re.compile(r\"(?<=[.;!?])\\s+|\\n+\")\n\n\ndef unwrap(text: str) -> str:\n    \"\"\"Rejoin lines that a fixed-width layout broke mid-sentence.\n\n    A synthetic abstract example places a pathology cue on one hard-wrapped line and its\n    anatomy on the continuation. Splitting those lines would separate the two cues and\n    make both clauses uninformative. No corpus wording is reproduced here; a line without\n    terminal sentence punctuation is treated as a continuation.\n    \"\"\"\n    if not isinstance(text, str):\n        return \"\"\n    out = []\n    for line in text.split(\"\\n\"):\n        s = line.strip()\n        if out and out[-1] and not re.search(r\"[.;:!?>*•]$\", out[-1]) \\\n                and len(out[-1].split()) >= 4 and s and not s[:1].isupper():\n            out[-1] = out[-1] + \" \" + s\n        else:\n            out.append(s)\n    return \"\\n\".join(out)\n\n\ndef clauses(text: str):\n    \"\"\"Split into clauses; attach `header:` lines to the value beneath them.\"\"\"\n    norm = normalize(unwrap(text) if FEATURES[\"unwrap\"] else text)\n    raw = [c.strip() for c in _SENT_SPLIT.split(norm) if c and c.strip()]\n    merged = []\n    for i, c in enumerate(raw):\n        if c.endswith(\":\") and len(c.split()) <= 14 and i + 1 < len(raw):\n            merged.append(c + \" \" + raw[i + 1])\n        merged.append(c)\n    out = []\n    for c in merged:\n        out.append(c)\n        if len(c.split()) > 25:\n            out.extend(p.strip() for p in c.split(\",\") if len(p.split()) > 2)\n    return out\n\n\n# Each rule below that is not obviously right is behind a flag, so that the notebook can\n# turn it off and re-measure rather than assert that it helps.\nFEATURES = {\n    \"unwrap\": True,               # rejoin hard-wrapped lines before splitting\n    \"directional_negation\": True,  # scope negation by direction instead of by clause\n    \"oa_inherit\": True,            # unlocalised cartilage statements reach all three\n    \"graded_pathology\": True,      # read numeric grades on a per-structure scale\n    \"synovitis_backoff\": True,     # order the silent majority by inflammatory context\n}\n\n\ndef _rx(*alts: str) -> re.Pattern:\n    return re.compile(\"|\".join(alts))\n\n\n# --------------------------------------------------------------------------- #\n# polarity\n# --------------------------------------------------------------------------- #\n# Directional negation matters when a positive finding precedes a later negated\n# complication. A synthetic abstract example therefore anchors scope on the pathology\n# span: the later negator governs only what follows it. Exact corpus wording and site\n# attribution are intentionally omitted.\nPRE_NEG = _rx(\n    r\"\\bno\\b\", r\"\\bnot\\b\", r\"\\bwithout\\b\", r\"\\bnegative for\\b\", r\"\\babsence\\b\",\n    r\"\\bno evidence\\b\", r\"\\bfree of\\b\", r\"\\bnone\\b\", r\"\\bneither\\b\", r\"\\bnor\\b\",\n    r\"\\bsin\\b\", r\"\\bno hay\\b\", r\"\\bausencia\\b\", r\"\\bausentes?\\b\", r\"\\bno se\\b\",\n    r\"\\bpas de\\b\", r\"\\bsans\\b\", r\"\\baucune?\\b\",\n    r\"\\bgeen\\b\", r\"\\bzonder\\b\", r\"\\bniet\\b\",\n    r\"\\bkeine?[nmrs]?\\b\", r\"\\bohne\\b\", r\"\\bnicht\\b\", r\"\\bkein\\b\",\n    r\"\\bnema\\b\", r\"\\bbez\\b\", r\"\\bnisu\\b\", r\"\\bnije\\b\",\n    r\"\\bδεν\\b\", r\"\\bχωρις\\b\", r\"ουδεν\", r\"\\bουτε\\b\",\n    r\"\\bбез\\b\", r\"\\bне\\b\", r\"липсва\", r\"\\bняма\\b\",\n)\n\n# Turkish and a few Croatian forms put the negator at the end of the sentence, so these\n# govern what precedes them instead.\nPOST_NEG = _rx(\n    r\"\\byok\\b\", r\"\\byoktur\\b\", r\"izlenmemekte\", r\"saptanmadi\", r\"\\bdegil\\b\",\n    r\"gozlenmemekte\", r\"mevcut degil\", r\"eslik etmiyor\", r\"\\bizlenmedi\\b\",\n    r\"izlenmemistir\", r\"saptanmamistir\", r\"gorulmemistir\", r\"\\bnema znakova\\b\",\n    r\"bez znakova\",\n)\n\nNEGATION = _rx(PRE_NEG.pattern, POST_NEG.pattern, r\"\\bunremarkable\\b\")\n\nNEG_WINDOW = 90\n\n\ndef _negated(clause: str, start: int, end: int) -> bool:\n    \"\"\"True when a negation trigger governs the span [start, end) of this clause.\"\"\"\n    for m in PRE_NEG.finditer(clause):\n        if m.end() <= start and start - m.end() <= NEG_WINDOW:\n            # A contrast conjunction closes negative scope before a later positive cue.\n            if not re.search(r\"\\b(but|however|ancak|fakat|pero|maar|aber|no i|ali|\"\n                             r\"ωστοσο|αλλα|но)\\b\", clause[m.end():start]):\n                return True\n    for m in POST_NEG.finditer(clause):\n        if m.start() >= end and m.start() - end <= NEG_WINDOW:\n            return True\n    return False\n\nNORMALITY = _rx(\n    r\"\\bnormal\", r\"\\bintact\\b\", r\"\\bpreserved\\b\", r\"\\bwithin normal limits\\b\",\n    r\"limites normales\", r\"\\bconservad\", r\"\\bintegr\", r\"\\bnormales\\b\",\n    r\"\\bdoga(l|ll)\\b\", r\"korunmus\", r\"\\bnormaldir\\b\", r\"olagan\",\n    r\"\\buredn\", r\"\\bocuvan\", r\"\\bodrzan\", r\"\\bintakt\", r\"\\bprimjeren\",\n    r\"\\bodrzanog kontinuiteta\", r\"\\bodržan\",\n    r\"φυσιολογικ\", r\"ακεραι\", r\"δεν παρατηρουνται\", r\"δεν σημειωνονται\",\n    r\"unauffallig\", r\"regelrecht\", r\"\\bo\\.?b\\.?\\b\",\n    r\"нормал\", r\"запазен\", r\"съхранен\", r\"\\bбез особености\\b\", r\"интактн\",\n    r\"\\bgaaf\\b\", r\"\\bnormaal\\b\",\n)\n\n# Across the represented languages, a negator plus an abnormality noun can assert\n# normality. Exact corpus phrases are omitted; the unchanged regex below is the\n# executable lexicon. The old normality-and-no-negation guard discarded these\n# constructions, making an explicitly clear structure look unmentioned. Those states\n# must remain ordered differently for a rank metric.\nNORMAL_PHRASE = _rx(\n    r\"\\bsin alteracion\", r\"\\bsin cambios\\b\", r\"\\bsin particularidad\",\n    r\"\\bsin hallazgos\\b\", r\"\\bsin lesion\", r\"\\bsin signos de (rotura|lesion)\",\n    r\"\\bcontinu[oa]s?\\b\", r\"\\bcontinuidad conservada\\b\",\n    r\"\\bno abnormalit\", r\"\\bno significant abnormalit\", r\"\\bunremarkable\\b\",\n    r\"\\bno evidence of (tear|injury|abnormalit)\",\n    r\"\\bohne auffalligkeit\", r\"\\bkein nachweis\\b\", r\"\\bohne befund\\b\",\n    r\"\\bgeen afwijking\", r\"\\bzonder afwijking\",\n    r\"\\bsans anomalie\", r\"\\bpas d[e']anomalie\",\n    r\"\\bbez osobitosti\\b\", r\"\\bbez znakova (rupture|lezije)\\b\",\n    r\"\\bbez patoloskih\\b\",\n    r\"χωρις αλλοιωσ\", r\"χωρις παθολογ\", r\"δεν παρατηρουνται (αξιολογα|παθολογ)\",\n    r\"\\bбез особености\\b\", r\"\\bбез патологич\", r\"\\bбез данни за\\b\",\n    r\"\\bozel bir ozellik yok\", r\"\\bpatolojik bulgu (yok|izlenmemis)\",\n)\n\nUNCERTAIN = _rx(\n    r\"\\bpossible\\b\", r\"\\bprobable\\b\", r\"\\bsuspicious\\b\", r\"\\bsuspected?\\b\",\n    r\"cannot (be )?exclude\", r\"\\bmay\\b\", r\"\\bquestionable\\b\", r\"\\bequivocal\\b\",\n    r\"\\br/o\\b\", r\"\\bdd\\b\", r\"\\blikely\\b\", r\"\\bsuggest\", r\"\\bcompatible with\\b\",\n    r\"\\bposible\\b\", r\"sin criterios categoricos\", r\"\\bdudos\", r\"\\bsugier\",\n    r\"\\bmuhtemel\\b\", r\"\\bolasi\\b\", r\"\\bsupheli\\b\", r\"\\bizlenim\", r\"\\bdusundur\",\n    r\"\\bmoguce\\b\", r\"\\bvjerojatno\\b\", r\"\\bsumnja\\b\", r\"\\bmoze odgovarati\\b\",\n    r\"πιθαν\", r\"υποπτ\",\n    r\"\\bmoglich\", r\"\\bverdachtig\", r\"\\bfraglich\", r\"\\bv\\.?a\\.?\\b\", r\"\\bwohl\\b\",\n    r\"\\bвъзможно\\b\", r\"\\bвероятно\\b\", r\"суспект\",\n    r\"\\bmogelijk\\b\", r\"\\bverdacht\\b\",\n)"},{"cell_type":"markdown","metadata":{},"source":"### 2.1 Four rules that are not obvious\n\nMost of a lexicon is vocabulary, and vocabulary is dull work that pays. Four rules are not\nvocabulary, and each of them fixes an error that a term list cannot reach. All four are\nbehind flags in `FEATURES`, and §2.2 turns them off one at a time and re-measures.\n\n**Unwrap before splitting.** A large share of these reports arrive hard-wrapped at a fixed\ncolumn, so a sentence is broken across two lines with no punctuation at the break.\nSplitting on newlines then severs a finding from its anatomy:\n\n> *Synthetic paraphrase:* `Complex tearing involves the anterior horn, body, posterior`\n> `horn and root of the lateral meniscus, with extrusion.`\n\nThe first line says *tear* and names no structure; the second names the lateral meniscus\nand says nothing about a tear. Neither clause carries a finding on its own, and the study\ncomes out silent on a meniscus that the report describes as complexly torn. A line that\ndoes not end in sentence punctuation is a continuation, not a statement.\n\n**Scope negation by direction, and anchor it on the pathology word.** Read at clause scope,\n\n> *Synthetic paraphrase:* `A medial subchondral fracture is present without collapse of the joint surface.`\n\nis a denial: it contains *without*. It is in fact an assertion — *without* governs what\nfollows it, and the fracture precedes it. This is the house style of one of the larger\nreporting sites in this corpus, so the error is systematic rather than occasional, and it\nfalls on `Fracture`, one of the rarest targets. Direction alone is not enough either. In\n\n> *Synthetic paraphrase:* `The medial meniscus is intact and not torn.`\n\nthe negator stands *after* the noun and before the verb, so scoping the test from the\nanatomy match reads the sentence as an assertion. The negator governs the finding, so the\ntest is anchored on the pathology word — *torn* — not on *meniscus*.\n\n**Let an unlocalised cartilage statement reach all three compartments.** Osteoarthritis is\nrarely written as \"osteoarthritis\". It is written as cartilage loss, chondrosis, a\nchondromalacia grade, joint space narrowing, or osteophytes — and often without naming a\ncompartment at all. Synthetic paraphrases range from an explicit statement of arthritis across all three\ncompartments to a compartment-specific description of cartilage degeneration. A rule that requires a compartment phrase before it will fire leaves three\nquarters of the corpus silent on the three OA targets. So a cartilage statement is\nattributed by what else its clause names — *medial* next to a tibiofemoral structure sends\nit to the medial compartment, *patella* or *trochlea* to the patellofemoral one — and a\nstatement that localises to nothing counts for all three at a discount, unless that\ncompartment was separately and explicitly cleared.\n\n**Read numeric grades on the right scale.** A grade is the most precise thing a knee report\nsays, and it means opposite things in different places. Grade 3 of a *meniscus* is signal\nreaching the articular surface, which is a tear by definition; grades 1 and 2 are\nintrasubstance change that is not. A *ligament* runs the other way — grade 1 is a stretch,\ngrade 2 a partial tear. Folding both into one pathology vocabulary scores a degenerate\nmeniscus like a torn one and throws away the ordering the metric is built on."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"# --------------------------------------------------------------------------- #\n# pathology vocabulary\n# --------------------------------------------------------------------------- #\nTEAR = _rx(\n    r\"\\btear\", r\"\\btorn\\b\", r\"\\brupture\", r\"\\bdisruption\\b\", r\"discontinuit\",\n    r\"\\bavuls\", r\"\\bmacerat\", r\"\\bbuckethandle\\b\", r\"bucket handle\",\n    r\"\\brotura\\b\", r\"\\broturas\\b\", r\"\\bruptura\", r\"\\bdesgarro\", r\"\\broto\\b\",\n    r\"\\bdechirure\", r\"\\bdechire\",\n    r\"\\bscheur\", r\"\\bruptuur\", r\"gescheurd\",\n    r\"\\briss\\b\", r\"einriss\", r\"\\bruptur\", r\"zerreiss\", r\"\\blasion\", r\"\\bausriss\",\n    r\"\\byirtik\", r\"\\byirtig\", r\"\\bkopma\\b\", r\"butunluk kaybi\", r\"\\brupturu\\b\",\n    r\"devamsizlik\", r\"\\brupture\\b\", r\"\\bdevamliligi secilememis\",\n    r\"\\bpuknuce\", r\"\\bprekid\\b\", r\"\\bpukotin\", r\"\\bruptur\",\n    r\"ρηξη\", r\"ρηξις\", r\"ρηγμα\", r\"ασυνεχεια\",\n    r\"руптура\", r\"разкъсв\", r\"разрив\", r\"скъсв\", r\"\\bлезия\\b\",\n)\n\nDEGEN = _rx(\n    r\"degenerat\", r\"\\bmucoid\\b\", r\"\\bmyxoid\\b\", r\"\\bfray\", r\"\\bfissur\",\n    r\"dejeneratif\", r\"\\bmukoid\\b\", r\"degenerativn\", r\"εκφυλ\", r\"дегенерат\",\n    r\"\\bμυξοειδ\", r\"\\bμυξωδ\", r\"\\bmeniskopat\", r\"\\bmeniscopath\",\n    r\"\\bmuco ?ide\\b\", r\"aufgefasert\", r\"\\bdejenerasyon\\b\",\n)\n\nINJURY = _rx(\n    r\"\\binjur\", r\"\\bsprain\", r\"\\blesion\", r\"\\blasion\", r\"\\bedema\\b\", r\"\\boedema\\b\",\n    r\"\\bodem\\b\", r\"\\bedem\\b\", r\"\\bοιδημα\", r\"\\bодем\", r\"\\bедем\", r\"\\bstrain\\b\",\n    r\"\\bhigh signal\\b\", r\"\\bsignal alteration\\b\", r\"\\bhiperintens\", r\"\\bhyperintens\",\n    r\"aumento de senal\", r\"alteracion de senal\", r\"cambio de senal\",\n    r\"\\bsignalanhebung\", r\"\\bsignalalteration\", r\"verhoogd signaal\", r\"sinyal artis\",\n    r\"αυξημενο σημα\", r\"повишен сигнал\", r\"\\besguince\\b\",\n    r\"\\bthicken\", r\"\\bzadebljanje\\b\", r\"\\bverdikking\\b\", r\"\\bdistenzij\",\n    r\"\\blaksite\\b\", r\"\\blaxity\\b\", r\"\\bpartial\\b\", r\"\\bparcijaln\", r\"\\bparcial\",\n    r\"\\bpartiel\", r\"\\bpartiell\",\n)\n\n# A numeric grade is the most precise thing a knee report says, and it means different\n# things in different places: grade 3 of a meniscus is a tear by definition, grade 1 or 2\n# is intrasubstance signal that never reaches the surface and is not one. A ligament runs\n# the other way round - grade 1 is a stretch, grade 2 a partial tear. So the grade is\n# read as a number and interpreted per structure rather than folded into one pathology\n# vocabulary.\n_GRADE_RX = re.compile(\n    r\"(?:grade|grad|grado|grau|derece|stupnja|stupanj|βαθμ|степен|icrs|outerbridge)\"\n    r\"[\\s:]*(?:grade\\s*)?([1-4]|iv|iii|ii|i)\\b\"\n)\n_ROMAN = {\"i\": 1, \"ii\": 2, \"iii\": 3, \"iv\": 4}\n\n\ndef _grade_of(clause: str):\n    \"\"\"Highest numeric grade stated in a clause, or None.\"\"\"\n    best = None\n    for m in _GRADE_RX.finditer(clause):\n        v = m.group(1)\n        n = _ROMAN.get(v, None) if not v.isdigit() else int(v)\n        if n is not None and (best is None or n > best):\n            best = n\n    return best\n\n# --------------------------------------------------------------------------- #\n# anatomy\n# --------------------------------------------------------------------------- #\nANAT = {\n    \"ACL\": _rx(\n        r\"anterior cruciate\", r\"\\bacl\\b\",\n        r\"cruzado anterior\", r\"\\blca\\b\",\n        r\"croise anterieur\",\n        r\"voorste kruisband\", r\"\\bvkb\\b\",\n        r\"vorderes kreuzband\", r\"vorderen kreuzband\", r\"vordere kreuzband\",\n        r\"on capraz\", r\"\\bocb\\b\", r\"anterior capraz\",\n        r\"prednji krizni\", r\"prednjeg krizn\",\n        r\"προσθι[οα][^ ]* χιαστ\", r\"προσθιου χιαστου\", r\"χιαστο[^ ]* συνδεσμ\",\n        r\"\\bχιαστ\\w*\",\n        r\"предна кръстна\", r\"предната кръстна\", r\"предна кръста\",\n        r\"cruciate ligaments\", r\"ligamentos cruzados\", r\"ligaments croises\",\n        r\"kruisbanden\", r\"kreuzbander\", r\"capraz baglar\", r\"krizn[a-z]* ligament[a-z]*\",\n        r\"χιαστοι συνδεσμ\", r\"χιαστων συνδεσμ\", r\"кръстните връзки\", r\"кръстни връзки\",\n    ),\n    \"MCL\": _rx(\n        r\"medial collateral\", r\"\\bmcl\\b\", r\"tibial collateral\",\n        r\"colateral medial\", r\"colateral interno\", r\"\\blcm\\b\",\n        r\"collateral medial\", r\"collateral interne\",\n        r\"mediale collaterale\", r\"binnenband\", r\"\\b(mediale|laterale) banden\\b\",\n        r\"\\bcollaterale banden\\b\",\n        r\"innenband\", r\"mediales? kollateral\",\n        r\"\\bic yan bag\", r\"medial kollateral\", r\"\\biyb\\b\", r\"medyal kollateral\",\n        r\"medijalni kolateraln\", r\"medijalnog kolateraln\",\n        r\"εσω πλαγι\", r\"εσωτερικο πλαγι\", r\"\\bπλαγι\\w* συνδεσμ\", r\"\\bπλαγιοι\\b\",\n        r\"медиален колатерал\", r\"вътрешна странична\", r\"\\bколатерал\\w*\",\n        r\"\\bcolaterales\\b\", r\"\\bcollateraux\\b\", r\"\\bcollateralen\\b\", r\"\\bkolateralni\\b\",\n        r\"collateral ligaments\", r\"ligamentos colaterales\", r\"ligaments collateraux\",\n        r\"collaterale banden\", r\"kollateralbander\", r\"seitenbander\", r\"yan baglar\",\n        r\"kolateraln[a-z]* ligament[a-z]*\", r\"πλαγιοι συνδεσμ\", r\"πλαγιων συνδεσμ\",\n        r\"колатерални връзки\", r\"страничните връзки\",\n    ),\n    \"Medial Meniscus\": _rx(\n        r\"medial meniscus\", r\"\\bmm\\b(?= tear)\", r\"medial menisc\",\n        r\"menisco medial\", r\"menisco interno\",\n        r\"menisque medial\", r\"menisque interne\",\n        r\"mediale meniscus\", r\"binnenmeniscus\",\n        r\"innenmeniskus\", r\"medialen? meniskus\", r\"innenmeniskushinterhorn\",\n        r\"medyal menisk\", r\"\\bic menisk\",\n        r\"medijalni meniskus\", r\"medijalnog meniskusa\", r\"medijalnom meniskusu\",\n        r\"medijaln\\w* menisk\\w*\", r\"\\bmedijalnog meniska\\b\", r\"medijalni menisk\",\n        r\"εσω μηνισκ\", r\"μηνισκ[^ ]* του εσω\", r\"εσω διαμερισμα[^.]{0,40}μηνισκ\",\n        r\"медиалния менискус\", r\"медиален менискус\", r\"вътрешния менискус\",\n        r\"oba meniska\", r\"both menisci\", r\"ambos meniscos\", r\"beide menisci\",\n        r\"her iki menisku\", r\"amfoteroi\\w* mhnisk\", r\"αμφοτερ\\w* μηνισκ\",\n        r\"двата менискуса\", r\"medial (and|&) lateral menisc\",\n    ),\n    \"Lateral Meniscus\": _rx(\n        r\"lateral meniscus\", r\"lateral menisc\",\n        r\"menisco lateral\", r\"menisco externo\",\n        r\"menisque lateral\", r\"menisque externe\",\n        r\"laterale meniscus\", r\"buitenmeniscus\",\n        r\"aussenmeniskus\", r\"lateralen? meniskus\", r\"aussenmeniskushinterhorn\",\n        r\"lateral menisk\", r\"\\bdis menisk\",\n        r\"lateralni meniskus\", r\"lateralnog meniskusa\", r\"lateralnom meniskusu\",\n        r\"lateraln\\w* menisk\\w*\", r\"\\blateralnog meniska\\b\",\n        r\"εξω μηνισκ\", r\"μηνισκ[^ ]* του εξω\", r\"εξω διαμερισμα[^.]{0,40}μηνισκ\",\n        r\"латералния менискус\", r\"латерален менискус\", r\"външния менискус\",\n        r\"oba meniska\", r\"both menisci\", r\"ambos meniscos\", r\"beide menisci\",\n        r\"her iki menisku\", r\"αμφοτερ\\w* μηνισκ\",\n        r\"двата менискуса\", r\"medial (and|&) lateral menisc\",\n    ),\n}\n\n# Osteoarthritis is written as cartilage damage far more often than as a diagnosis.\nOA_EVIDENCE = _rx(\n    r\"osteoarthrit\", r\"\\barthros\", r\"\\bgonarthros\", r\"\\bosteoarthros\",\n    r\"chondropath\", r\"chondromalac\", r\"condropat\", r\"condromalac\", r\"\\bchondros\",\n    r\"\\bchondrosis\\b\", r\"chondral (loss|defect|ulcer|thinning|injury|fissur|wear)\",\n    r\"cartilage (loss|thinning|defect|fissur|wear|damage|heterogeneity|irregularit)\",\n    r\"(loss|thinning|fissur|defect|ulcer|erosion|denudation) of[^.]{0,20}cartilage\",\n    r\"articular cartilage[^.]{0,30}(loss|thin|fissur|defect|erosion|wear|irregular)\",\n    r\"osteophyt\", r\"osteofit\", r\"osteofyt\", r\"osteofito\", r\"osteophyten\", r\"spurring\",\n    r\"joint space narrowing\", r\"pinzamiento articular\", r\"reduced joint space\",\n    r\"kikirdak kayb\", r\"kikirdak incelme\", r\"kondropati\", r\"kondral\", r\"kikirdak dejener\",\n    r\"eklem aralig\\w* daral\", r\"eklem mesafesi daral\", r\"kikirdak kalinlig\\w* azal\",\n    r\"kraakbeen\", r\"gonartrose\", r\"artrose\", r\"\\bknorpel\", r\"arthrose\", r\"gonarthrose\",\n    r\"hrskavic\", r\"hondromalac\", r\"artroz\", r\"osteoartrit\", r\"artrotsk\", r\"artrotick\",\n    r\"\\boa promjen\", r\"\\boa\\b\", r\"degenerativne promjene hrskav\",\n    r\"χονδρ[^ ]*παθ\", r\"αρθριτ\", r\"αρθρωσ\", r\"οστεοφυτ\", r\"χονδρομαλακ\",\n    r\"αρθρικου χονδρου\", r\"εξαλειψη του αρθρικου χονδρου\", r\"διαβρωση του αρθρικου χονδρ\",\n    r\"λεπτυνση[^.]{0,30}χονδρ\", r\"φθορα[^.]{0,20}χονδρ\",\n    r\"артроз\", r\"хондропат\", r\"остеофит\", r\"хрущял[^.]{0,40}(изтън|увред|дефект|липс)\",\n    r\"изтъняване[^.]{0,30}хрущял\", r\"хондромалац\",\n    r\"ulcera[s]? condral\", r\"cartilago[^.]{0,25}(perdida|adelgaz)\",\n    r\"icrs grade\", r\"icrs\\b\", r\"outerbridge\", r\"\\bdenudation\\b\", r\"denudacij\",\n    r\"erozivne promjene\", r\"\\berosion of[^.]{0,20}cartilage\",\n    r\"kraakbeenlijden\", r\"kraakbeenverlies\",\n)\n\n# --------------------------------------------------------------------------- #\n# where in the joint a cartilage statement sits\n# --------------------------------------------------------------------------- #\n# Tibiofemoral structures. Used only inside a clause that already carries OA evidence,\n# so bare \"condyle\" is safe here and would not be elsewhere.\nTF_SITE = _rx(\n    r\"compartment\", r\"compartimento\", r\"compartiment\", r\"kompartman\", r\"kompartiment\",\n    r\"kompartment\", r\"odjelj\", r\"διαμερισμα\", r\"компартм\", r\"\\bотдел\",\n    r\"femorotibial\", r\"tibiofemoral\", r\"femoro tibial\", r\"femorotibiaal\",\n    r\"femorotibijaln\", r\"феморотибиал\", r\"\\bft zglob\", r\"tibiofemoraln\",\n    r\"condyle\", r\"condilo\", r\"kondyl\", r\"kondil\", r\"condyl\", r\"κονδυλ\",\n    r\"кондил\", r\"\\bplateau\", r\"\\bplato\\b\", r\"platillo\", r\"meseta\", r\"плато\",\n    r\"tibiaplateau\", r\"tibijaln\\w* plato\", r\"tibyal plato\", r\"tibia plato\",\n    r\"κνημιαι\", r\"μηριαι\", r\"weightbearing\", r\"weightbaring\", r\"zona de carga\",\n    r\"dragende deel\", r\"agirlik tasiyan\", r\"\\bfemur\\b\", r\"\\btibia\\b\", r\"\\bfemoral\\b\",\n    r\"\\btibial\\b\", r\"\\bfemura\\b\", r\"\\btibije\\b\", r\"\\bmesarthrio\\b\", r\"μεσαρθριο\",\n)\n\n# Patellofemoral structures.\nPF_SITE = _rx(\n    r\"patellofemoral\", r\"femoropatellar\", r\"femoropatelar\", r\"patelofemoral\",\n    r\"retropatellar\", r\"retrorotulian\", r\"trochlea\", r\"troclea\", r\"troklea\",\n    r\"trochlear\", r\"trohlej\", r\"τροχιλ\", r\"\\bpatella\", r\"\\bpatellar\", r\"rotulian\",\n    r\"\\brotula\\b\", r\"\\bpatele\\b\", r\"patellofemoraal\", r\"femoropatellair\",\n    r\"επιγονατιδ\", r\"μηροεπιγονατιδ\", r\"пател\", r\"феморопател\",\n    r\"anterior compartment\", r\"compartimento anterior\", r\"prednj\\w* odjeljk\",\n    r\"\\bfp zglob\", r\"\\bpf zglob\", r\"\\bfaset\", r\"\\bfacet\", r\"patellofemoraln\",\n)\n\nSIDE_MEDIAL = _rx(\n    r\"\\bmedial\\w*\", r\"\\bmedyal\\w*\", r\"\\bmedijaln\\w*\", r\"\\bmediaal\\w*\",\n    r\"\\bmediale\\w*\", r\"\\binterno\\b\", r\"\\binterna\\b\", r\"\\binternos\\b\", r\"\\binterne\\b\",\n    r\"\\binnen\\w*\", r\"\\bic\\b\", r\"\\bunutarnj\\w*\", r\"\\bεσω\\w*\", r\"\\bεσωτερικ\\w*\",\n    r\"\\bмедиал\\w*\", r\"\\bвътреш\\w*\", r\"\\bbinnen\\w*\", r\"\\bmediaal\\b\", r\"\\bmediales?\\b\",\n)\nSIDE_LATERAL = _rx(\n    r\"\\blateral\\w*\", r\"\\bexterno\\b\", r\"\\bexterna\\b\", r\"\\bexternos\\b\", r\"\\bexterne\\b\",\n    r\"\\bdis\\b\", r\"\\blateraln\\w*\", r\"\\baussen\\w*\", r\"\\bbuiten\\w*\", r\"\\bεξω\\w*\",\n    r\"\\bεξωτερικ\\w*\", r\"\\bлатерал\\w*\", r\"\\bвъншн\\w*\", r\"\\bvanjsk\\w*\",\n)\nSIDE_ANTERIOR = _rx(\n    r\"\\banterior\\w*\", r\"\\bant\\b\", r\"\\bon\\b\", r\"\\bprednj\\w*\", r\"\\bvorder\\w*\",\n    r\"\\bvoorste\\b\", r\"\\bπροσθι\\w*\", r\"\\bпредн\\w*\", r\"\\banteriyor\\w*\", r\"\\bavant\\b\",\n    r\"\\banterieur\\w*\",\n)\n\nGLOBAL_OA = _rx(\n    r\"tri ?compartment\", r\"all three compartment\", r\"global(ised)? (oa|osteoarthrit)\",\n    r\"\\bgonarthros\", r\"\\bgonartros\", r\"\\bgonarthrose\", r\"\\bgonartrose\", r\"gonartro\",\n    r\"goanrtrot\", r\"gonartrot\",\n    r\"osteoarthritis of the knee\", r\"artrosis (de |)(la )?rodilla\", r\"knee osteoarthrit\",\n    r\"\\bdiz osteoartrit\", r\"\\bgonartroz\", r\"artroza koljena\",\n    r\"οστεοαρθριτιδα\", r\"αρθριτιδα του γονατος\", r\"εκφυλιστικη οστεοαρθριτ\",\n    r\"артроза на колянната\", r\"гонартроз\",\n    r\"degenerative joint disease\", r\"\\bdjd\\b\", r\"three compartments\",\n    r\"compartmens\", r\"compartments\",\n)\n\n# --------------------------------------------------------------------------- #\n# self-declaring findings\n# --------------------------------------------------------------------------- #\nDIRECT = {\n    \"Effusion\": _rx(\n        r\"\\beffusion\", r\"joint fluid\", r\"intra ?articular fluid\", r\"\\bhydrops\\b\",\n        r\"\\bhemarthros\", r\"\\bhaemarthros\",\n        r\"derrame articular\", r\"\\bderrame\\b\", r\"liquido articular\", r\"hemartrosis\",\n        r\"epanchement\",\n        r\"gewrichtsvocht\", r\"\\bvocht\\b\", r\"gewrichtseffusie\", r\"opzetting van suprapatell\",\n        r\"gelenkerguss\", r\"\\berguss\\b\", r\"gelenksergu\", r\"gelenksflussigkeit\",\n        r\"eklem\\w* ic\\w* sivi\", r\"efuzyon\", r\"eklem sivisi\", r\"eklem mesafesinde sivi\",\n        r\"sivi (miktari|artisi|birikimi)\", r\"sivi artis\", r\"\\bsivi\\b[^.]{0,25}artmis\",\n        r\"\\bizljev\", r\"\\bizliv\", r\"zglobn[^ ]* tekucin\", r\"\\bhidrops\\b\",\n        r\"αρθρικ[^ ]* υγρ\", r\"υγρου ενδαρθρικα\", r\"ενδαρθρικ[^ ]* υγρ\", r\"ποσοτητα υγρου\",\n        r\"ενδαρθρικ\", r\"αρθρικη συλλογη\", r\"υγρο στην αρθρωση\", r\"υγρου στην αρθρωση\",\n        r\"συλλογη υγρου\", r\"ενθαρθρικ\",\n        r\"ставен излив\", r\"излив\", r\"ставна течност\", r\"синовиална течност\",\n    ),\n    \"Synovitis\": _rx(\n        r\"synovit\", r\"sinovit\", r\"synovial (thickening|proliferation|hypertroph)\",\n        r\"thicken\\w* synovial\", r\"hypertroph\\w* of the synovium\",\n        r\"synoviale? (verdikking|proliferatie)\", r\"verdikkingen van (het )?synovium\",\n        r\"synovialitis\", r\"synovialis(verdickung|proliferation)\", r\"reizsynovial\",\n        r\"sinovijalitis\", r\"sinovitis\", r\"zadebljanje sinovij\", r\"proliferacij\\w* sinovij\",\n        r\"sinovijaln\\w* proliferacij\",\n        r\"υμενιτιδα\", r\"συνοβιτιδα\", r\"υμενικ[^ ]* υπερτροφ\", r\"αρθρικου υμεν\",\n        r\"παχυνση[^.]{0,20}υμεν\", r\"υμενα\",\n        r\"синовит\", r\"синовиал[^ ]* (задебел|пролифер)\",\n        r\"\\bpannus\\b\", r\"\\bhoffit\", r\"sinovyal\\w* (kalinlas|proliferas)\",\n        r\"sinovyal hipertrof\", r\"\\bartrit\\b\", r\"\\barthritis\\b\",\n    ),\n    \"Baker's\": _rx(\n        r\"baker\", r\"popliteal cyst\", r\"quiste popliteo\", r\"quistes popliteos\",\n        r\"kyste poplite\", r\"popliteale? cyst\", r\"poplitealzyste\", r\"bakerzyste\",\n        r\"popliteal kist\", r\"\\bbakerova\\b\", r\"poplitealn[^ ]* cist\", r\"popliteal\\w* cist\",\n        r\"κυστη baker\", r\"πολυχωρη συνοβιακη κυστη\", r\"κυστη του baker\",\n        r\"συνοβιακη κυστη\", r\"κυστη τυπου baker\",\n        r\"киста на бейкър\", r\"бейкърова киста\", r\"поплитеална киста\", r\"бекеров\",\n        r\"gastrocnemio ?semimembranos\", r\"gastrocnemius semimembranosus burs\",\n    ),\n    \"Contusion\": _rx(\n        r\"\\bcontusion\", r\"bone bruise\", r\"bone marrow (o?edema|contusion)\",\n        r\"marrow o?edema\", r\"\\bkontuz\", r\"medular bone o?edema\", r\"osseous contusion\",\n        r\"contusion osea\", r\"edema oseo\", r\"edema de medula osea\", r\"contusiones oseas\",\n        r\"oedeme osseux\", r\"contusion osseuse\",\n        r\"botcontusie\", r\"botoedeem\", r\"beenmergoedeem\", r\"botmergoedeem\",\n        r\"knochenmarkodem\", r\"knochenodem\", r\"knochenmarksodem\", r\"kontusion\",\n        r\"kemik kontuzyonu\", r\"kemik iligi odemi\", r\"kemik odemi\", r\"kemik iliginde odem\",\n        r\"kontuzyonel kemik\", r\"kemik iligi odemleri\",\n        r\"kostani edem\", r\"edem kosti\", r\"kontuzij\", r\"kostane srzi[^.]{0,20}edem\",\n        r\"οστεομυελικ[^ ]* οιδημα\", r\"οστικο οιδημα\", r\"μυελικο οιδημα\", r\"οστικο μωλωπ\",\n        r\"костномозъчен едем\", r\"костен едем\", r\"контузионен\", r\"костно мозъчен едем\",\n    ),\n    \"Fracture\": _rx(\n        r\"\\bfractur\", r\"\\bfract\\b\",\n        r\"\\bfractura\", r\"\\bfracturas\\b\",\n        r\"\\bfractuur\", r\"\\bbreuk\\b\",\n        r\"\\bfraktur\", r\"\\bbruch\\b\",\n        r\"\\bkirik\\b\", r\"\\bkirigi\\b\", r\"\\bkiri[kg]\\w*\",\n        r\"\\bprijelom\", r\"impresijsk[^ ]* fraktur\", r\"impaktcij\",\n        r\"καταγμα\", r\"καταγματ\",\n        r\"фрактур\", r\"счупван\", r\"фисур\",\n        r\"insufficiency fracture\", r\"stress fracture\", r\"avulsion fracture\",\n        r\"subchondral fracture\", r\"subkondral kiri\", r\"impaction (fracture|injury)\",\n        r\"osteochondral (fracture|impaction)\", r\"\\bsegond\\b\", r\"impactiefractuur\",\n        r\"subchondrale impression\", r\"subchondraler? impress\",\n    ),\n}\n\nDECOY = {\n    \"Fracture\": _rx(r\"microfractur\", r\"\\bfracture (risk|prophyla)\"),\n    \"Baker's\": _rx(r\"meniscal cyst\", r\"quiste meniscal\", r\"parameniscal\"),\n}\n\nPAIRED = {\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\"}\nOA_TARGETS = [\"Medial OA\", \"Lateral OA\", \"PF OA\"]\n\n# --------------------------------------------------------------------------- #\n# stems, for morphology the phrase lexicons cannot reach\n# --------------------------------------------------------------------------- #\n# A synthetic abstract example may clear or injure both menisci with an unqualified\n# plural, naming neither side; a side-qualified lexicon would then leave both targets\n# silent. Exact corpus wording is omitted. The plural rule is consulted only when the\n# clause names no side, preventing a mixed side-specific clause from firing the other\n# meniscus target.\nPLURAL_MENISCI = _rx(\n    r\"\\bmenisci\\b\", r\"\\bmeniscos\\b\", r\"\\bmenisques\\b\", r\"\\bmenisken\\b\",\n    r\"\\bmeniskusi\\b\", r\"\\bmenisk\\w*ler\\b\", r\"\\bμηνισκοι\\b\", r\"\\bμηνισκων\\b\",\n    r\"\\bменискуси\\b\", r\"\\bменискусите\\b\", r\"\\bmenisci\\w*\\b\",\n)\nANY_SIDE = _rx(SIDE_MEDIAL.pattern, SIDE_LATERAL.pattern)\n\nSTEM_MENISCUS = _rx(r\"menisc\\w*\", r\"menisk\\w*\", r\"μηνισκ\\w*\", r\"мениск\\w*\")\nSTEM_CRUCIATE = _rx(r\"cruciate\", r\"cruzado\", r\"croise\", r\"kruisband\", r\"kreuzband\",\n                    r\"capraz bag\\w*\", r\"krizn\\w*\", r\"χιαστ\\w*\", r\"кръстн\\w*\",\n                    r\"\\bacl\\b\", r\"\\blca\\b\", r\"\\bvkb\\b\", r\"\\bocb\\b\", r\"\\bacb\\b\")\nSTEM_COLLATERAL = _rx(r\"collateral\\w*\", r\"colateral\\w*\", r\"kollateral\\w*\",\n                      r\"collaterale\\w*\", r\"kolateraln\\w*\", r\"yan bag\\w*\",\n                      r\"πλαγι\\w*\", r\"колатерал\\w*\", r\"странич\\w*\",\n                      r\"innenband\\w*\", r\"binnenband\\w*\", r\"\\bmcl\\b\", r\"\\blcm\\b\",\n                      r\"\\biyb\\b\")\n\nSTEM_FRACTURE = _rx(r\"fractur\\w*\", r\"fraktur\\w*\", r\"fractuur\\w*\", r\"\\bfract\\b\",\n                    r\"kiri[kgğ]\\w*\", r\"prijelom\\w*\", r\"lom kosti\", r\"\\bbreuk\\w*\",\n                    r\"\\bbruch\\w*\", r\"καταγμα\\w*\", r\"καταγματ\\w*\", r\"фрактур\\w*\",\n                    r\"счупван\\w*\", r\"fisur\\w* (osea|oseas|kost)\", r\"fissur\\w* kost\")\n\n# Decoys that steal a cruciate / collateral stem match for the wrong ligament.\nPOSTERIOR_ONLY = _rx(r\"\\bpcl\\b\", r\"\\blcp\\b\", r\"\\bhkb\\b\", r\"\\bacb\\b\",\n                     r\"posterior cruciate\", r\"cruzado posterior\", r\"croise posterieur\",\n                     r\"achterste kruisband\", r\"hinteres kreuzband\", r\"arka capraz\",\n                     r\"straznji krizn\", r\"οπισθι[οα]\\w* χιαστ\", r\"задна кръстн\",\n                     r\"задната кръстн\")\nLATERAL_COLL_ONLY = _rx(r\"\\blcl\\b\", r\"\\bfcl\\b\", r\"lateral collateral\",\n                        r\"fibular collateral\", r\"colateral lateral\", r\"colateral externo\",\n                        r\"buitenband\", r\"aussenband\", r\"dis yan bag\",\n                        r\"lateralni kolateraln\", r\"εξω πλαγι\", r\"латерален колатерал\")"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"def _near(clause: str, stem_rx: re.Pattern, qual_rx: re.Pattern, window: int = 55):\n    for m in stem_rx.finditer(clause):\n        lo = max(0, m.start() - window)\n        hi = min(len(clause), m.end() + window)\n        if qual_rx.search(clause[lo:hi]):\n            return True\n    return False\n\n\nSTEM_RULES = {\n    \"ACL\": (STEM_CRUCIATE, SIDE_ANTERIOR),\n    \"MCL\": (STEM_COLLATERAL, SIDE_MEDIAL),\n    \"Medial Meniscus\": (STEM_MENISCUS, SIDE_MEDIAL),\n    \"Lateral Meniscus\": (STEM_MENISCUS, SIDE_LATERAL),\n}\n\n\nclass _Matcher:\n    def __init__(self, phrase_rx, stem=None, side=None, window=55):\n        self.phrase_rx = phrase_rx\n        self.stem = stem\n        self.side = side\n        self.window = window\n\n    def search(self, clause):\n        m = self.phrase_rx.search(clause)\n        if m is not None:\n            return m\n        if self.stem is not None and _near(clause, self.stem, self.side, self.window):\n            return self.stem.search(clause)\n        return None\n\n\nANAT_MATCH = {t: _Matcher(ANAT[t], *STEM_RULES[t]) for t in PAIRED}\nDIRECT_MATCH = {\n    t: _Matcher(_rx(rx.pattern, STEM_FRACTURE.pattern) if t == \"Fracture\" else rx)\n    for t, rx in DIRECT.items()\n}\n\n# --------------------------------------------------------------------------- #\n# severity\n# --------------------------------------------------------------------------- #\nSEV_LOW = _rx(\n    r\"\\bsmall\\b\", r\"\\bminimal\\b\", r\"\\btrace\\b\", r\"\\bmild\\b\", r\"\\bslight\\b\",\n    r\"\\btiny\\b\", r\"\\bscant\\b\", r\"\\bdiscrete\\b\", r\"\\blow ?grade\\b\", r\"\\bincipient\\b\",\n    r\"\\bleve\\b\", r\"\\bminim\", r\"\\bpeque\", r\"\\bfina\\b\", r\"\\bfino\\b\", r\"\\bligero\\b\", r\"\\bescaso\\b\", r\"\\bdiscreto\\b\",\n    r\"\\bhafif\\b\", r\"\\baz miktarda\\b\", r\"\\bsilik\\b\",\n    r\"\\bmanj\\w*\", r\"\\bblago\\b\", r\"\\bdiskretn\", r\"\\bmalo\\b\", r\"\\bpocetn\",\n    r\"\\bgering\", r\"\\bdiskret\", r\"\\bkleine?r?\\b\", r\"\\bwenig\\b\", r\"\\bzarte?\\b\",\n    r\"\\bbeperkte?\\b\", r\"\\bgeringe\\b\", r\"\\bweinig\\b\", r\"\\blichte?\\b\", r\"\\blicht\\b\",\n    r\"\\bηπι\", r\"\\bμικρ\", r\"\\bελαχιστ\", r\"\\bαρχομεν\",\n    r\"\\bминимал\", r\"\\bлек\", r\"\\bмалк\", r\"\\bнеголям\",\n)\n\nSEV_HIGH = _rx(\n    r\"\\blarge\\b\", r\"\\bmarked\\b\", r\"\\bmassive\\b\", r\"\\bsevere\\b\", r\"\\bextensive\\b\",\n    r\"\\bmoderate\\b\", r\"\\bgross\\b\", r\"\\bsignificant\\b\", r\"\\babundant\\b\", r\"\\btense\\b\",\n    r\"\\bcomplete\\b\", r\"\\bfull ?thickness\\b\", r\"\\bhigh ?grade\\b\", r\"\\badvanced\\b\",\n    r\"\\bmoderad\", r\"\\bimportante\\b\", r\"\\bsevera?\\b\", r\"\\bmarcad\", r\"\\bcuantios\",\n    r\"\\bespesor total\\b\", r\"\\bcompleta?\\b\",\n    r\"\\bbelirgin\\b\", r\"\\byaygin\\b\", r\"\\bileri\\b\", r\"\\bciddi\\b\", r\"\\bbol\\b\", r\"\\bkomplet\",\n    r\"\\bopsezan\\b\", r\"\\bveliki\\b\", r\"\\bizrazit\", r\"\\bznacajn\", r\"\\bumjeren\",\n    r\"\\buznapredoval\", r\"\\bpotpun\", r\"\\bkompleksn\",\n    r\"\\bausgepragt\", r\"\\bdeutlich\", r\"\\bmassiv\", r\"\\bmassig\", r\"\\bgross\",\n    r\"\\buitgebreid\", r\"\\bgevorderd\", r\"\\bveel\\b\", r\"\\bmatige?\\b\", r\"\\bvolledig\",\n    r\"\\bμετρι\", r\"\\bμεγαλ\", r\"\\bεκτεταμεν\", r\"\\bευμεγεθ\", r\"\\bσοβαρ\", r\"\\bπληρη\",\n    r\"\\bголям\", r\"\\bизразен\", r\"\\bзначим\", r\"\\bумерен\", r\"\\bобилен\", r\"\\bпълн\",\n)\n\nGRADE_HIGH = re.compile(r\"grade?[ao]?\\s*(3|4|iii|iv)\\b|icrs grade (iii|iv|3|4)|\"\n                        r\"stupnja iv|stupnja iii|\\bgrado (3|4)\\b|\\bgrad (3|4)\\b|\"\n                        r\"\\bgrade (3|4)\\b\")\n\nDEGENERATIVE_MARROW = _rx(\n    r\"subchondral\", r\"subcondral\", r\"subkondral\", r\"supkondraln\", r\"subchondraln\",\n    r\"υποχονδρι\", r\"υπαρθρικ\", r\"субхондрал\", r\"subchondrale?\", r\"subartikuler\",\n    r\"\\bcyst\", r\"\\bquist\", r\"\\bzyste\\b\", r\"\\bcistic\", r\"reactive\", r\"reactivo\",\n    r\"degenerative\", r\"degenerativ\", r\"reaktiv\", r\"\\bcisti\\b\",\n)\n\nTRAUMA = _rx(\n    r\"\\bbruise\\b\", r\"\\bcontusion\", r\"\\bkontuz\", r\"\\btrauma\", r\"\\bimpaction\\b\",\n    r\"\\bpivot shift\\b\", r\"\\bkissing\\b\", r\"\\bacute\\b\", r\"\\bagudo\\b\", r\"\\bakut\",\n    r\"\\bpivot kaymasi\\b\", r\"\\bcontusion osseuse\\b\", r\"\\bbone bruise\\b\",\n    r\"\\bbotcontusie\\b\", r\"\\bконтузион\", r\"\\bμωλωπ\", r\"\\bkontuzij\", r\"\\bimpaktcij\",\n    r\"\\bimpakcij\", r\"\\bfall\\b\", r\"\\binjury\\b\", r\"\\bimpression\\b\",\n)\n\n# Effusion-adjacent inflammatory signs, used only when a report never names synovitis.\nSYNOVIAL_PROXY = _rx(\n    r\"bursit\", r\"burzit\", r\"\\bbursa\\b[^.]{0,30}(fluid|distend|sivi|tekucin|opzetting)\",\n    r\"suprapatellar (bursitis|effusion|recess)\", r\"suprapatellar bursa\",\n    r\"suprapatellar bursada\", r\"suprapatelarno\", r\"suprapatellaire recessus\",\n    r\"hoffa\", r\"hoffit\", r\"plica\", r\"plika\", r\"πλικα\", r\"fat pad[^.]{0,20}(edema|oedema)\",\n    r\"kapsul\", r\"capsul\", r\"καψ\", r\"капсул\", r\"\\bpannus\\b\", r\"\\bsinov\", r\"\\bsynov\",\n)\n\n\ndef _polarity(clause: str, span=None) -> str:\n    \"\"\"positive / negative / uncertain for a term matched at `span` in this clause.\"\"\"\n    if UNCERTAIN.search(clause):\n        return \"uncertain\"\n    if span is None or not FEATURES[\"directional_negation\"]:\n        if NEGATION.search(clause):\n            return \"negative\"\n    elif _negated(clause, span[0], span[1]):\n        return \"negative\"\n    if NORMALITY.search(clause):\n        # Synthetic abstract contrast: explicit target normality denies, while unrelated\n        # normal anatomy does not override a positive pathology cue.\n        if TEAR.search(clause) or GRADE_HIGH.search(clause):\n            return \"positive\"\n        return \"negative\"\n    return \"positive\"\n\n\ndef _severity(clause: str) -> float:\n    \"\"\"How emphatic a clause is about the finding it asserts.\n\n    A numeric grade is deliberately not read here because it belongs to the structure\n    being graded. In a synthetic abstract mixed-finding clause, a high cartilage grade\n    must not be transferred to a separately mentioned mild effusion. No corpus wording\n    is reproduced here.\n    \"\"\"\n    high = SEV_HIGH.search(clause) is not None\n    low = SEV_LOW.search(clause) is not None\n    if high and not low:\n        return 1.0\n    if low and not high:\n        return 0.45\n    if high and low:\n        return 0.8\n    return 0.75\n\n\ndef _grade(n_pos, n_neg, n_unc, best):\n    \"\"\"Map counted evidence onto a score in (0, 1) and a confidence.\"\"\"\n    if n_pos or n_unc:\n        score = min(0.97, 0.50 + 0.45 * best + 0.015 * min(n_pos, 3))\n        conf = min(1.0, 0.55 + 0.15 * n_pos)\n    elif n_neg:\n        score = max(0.04, 0.20 - 0.04 * n_neg)\n        conf = min(0.9, 0.45 + 0.12 * n_neg)\n    else:\n        score, conf = 0.28, 0.05\n    return score, conf\n\n\ndef _paired_weight(clause: str, meniscus: bool) -> float:\n    \"\"\"How strongly one clause asserts damage to a meniscus or a cruciate/collateral.\n\n    Ordered, not calibrated. What has to hold is that a tear outranks a graded lesion,\n    that the grade is read on the right scale for the structure, and that intrasubstance\n    degeneration lands below both - the annotator marks a torn meniscus and leaves a\n    degenerate one, and a lexicon that scores the two alike throws that ordering away.\n    \"\"\"\n    g = _grade_of(clause) if FEATURES[\"graded_pathology\"] else None\n    tear = TEAR.search(clause) is not None\n    if meniscus:\n        if tear:\n            base = 1.0\n        elif g is not None:\n            base = 0.95 if g >= 3 else 0.30\n        elif DEGEN.search(clause):\n            base = 0.35\n        else:\n            base = 0.45\n    else:\n        if tear:\n            base = 1.0\n        elif g is not None:\n            base = 0.85 if g >= 2 else 0.30\n        elif DEGEN.search(clause):\n            base = 0.40\n        else:\n            base = 0.55\n    if SEV_HIGH.search(clause) and not SEV_LOW.search(clause):\n        base = min(1.0, base * 1.2)\n    elif SEV_LOW.search(clause) and not SEV_HIGH.search(clause):\n        base *= 0.7\n    return base\n\n\ndef _score_paired(cls, tgt):\n    \"\"\"Evidence for one of the four side-specific ligament / meniscus targets.\"\"\"\n    anat_rx = ANAT_MATCH[tgt]\n    path_rx = _rx(TEAR.pattern, DEGEN.pattern, INJURY.pattern)\n    meniscus = \"Meniscus\" in tgt\n    n_pos = n_neg = n_unc = 0\n    best = 0.0\n    for c in cls:\n        hit = anat_rx.search(c)\n        if hit is None and meniscus and PLURAL_MENISCI.search(c) \\\n                and not ANY_SIDE.search(c):\n            hit = PLURAL_MENISCI.search(c)\n        if hit is None:\n            continue\n        # Anchor negation on the pathology span, not the anatomy span. In a synthetic\n        # abstract post-nominal-negation case, scoping from the earlier anatomy noun\n        # would incorrectly turn a denial into an assertion; no corpus sentence is shown.\n        pm = path_rx.search(c)\n        if pm is None and _grade_of(c) is None:\n            if NORMAL_PHRASE.search(c) or (NORMALITY.search(c)\n                                           and not NEGATION.search(c)):\n                n_neg += 1\n            continue\n        span = (pm.start(), pm.end()) if pm is not None else None\n        pol = _polarity(c, span)\n        if pol == \"positive\":\n            n_pos += 1\n            best = max(best, _paired_weight(c, meniscus))\n        elif pol == \"negative\":\n            n_neg += 1\n        else:\n            n_unc += 1\n            best = max(best, 0.45 * _paired_weight(c, meniscus))\n    s, cf = _grade(n_pos, n_neg, n_unc, best)\n    return s, cf, n_pos, n_neg\n\n\ndef _score_clauses(cls, anat_rx, path_rx=None, decoy_rx=None, context_penalty=None,\n                   context_bonus=None):\n    n_pos = n_neg = n_unc = 0\n    best = 0.0\n    for c in cls:\n        m = anat_rx.search(c)\n        if not m:\n            continue\n        if decoy_rx is not None and decoy_rx.search(c):\n            continue\n        if path_rx is not None and not path_rx.search(c):\n            if NORMAL_PHRASE.search(c) or (NORMALITY.search(c)\n                                           and not NEGATION.search(c)):\n                n_neg += 1\n            continue\n        pol = _polarity(c, (m.start(), m.end()))\n        if pol == \"positive\":\n            n_pos += 1\n            w = _severity(c)\n            if context_penalty is not None and context_penalty.search(c):\n                w *= 0.45\n            if context_bonus is not None and context_bonus.search(c):\n                w = min(1.0, w * 1.35)\n            best = max(best, w)\n        elif pol == \"negative\":\n            n_neg += 1\n        else:\n            n_unc += 1\n            best = max(best, 0.30)\n    s, c = _grade(n_pos, n_neg, n_unc, best)\n    return s, c, n_pos, n_neg\n\n\ndef _score_oa(cls):\n    \"\"\"Osteoarthritis, scoped to the three compartments.\n\n    A cartilage statement is attributed using the anatomical context in its clause. A\n    side-qualified tibiofemoral context maps to that side, patellofemoral anatomy maps\n    to that compartment, and an unlocalised whole-joint assertion reaches all three at\n    a discount. Exact corpus phrasings are omitted; this is a synthetic abstract\n    description of the unchanged scoping rule.\n    \"\"\"\n    acc = {t: {\"pos\": 0, \"neg\": 0, \"unc\": 0, \"best\": 0.0} for t in OA_TARGETS}\n    g_pos, g_neg, g_best = 0, 0, 0.0\n\n    for c in cls:\n        m = OA_EVIDENCE.search(c)\n        if not m:\n            continue\n        pol = _polarity(c, (m.start(), m.end()))\n        sev = _severity(c)\n        tf_med = _near(c, TF_SITE, SIDE_MEDIAL, 45)\n        tf_lat = _near(c, TF_SITE, SIDE_LATERAL, 45)\n        pf = PF_SITE.search(c) is not None\n        hits = []\n        if tf_med:\n            hits.append(\"Medial OA\")\n        if tf_lat:\n            hits.append(\"Lateral OA\")\n        if pf:\n            hits.append(\"PF OA\")\n\n        if not hits:\n            # No compartment named: a synthetic abstract whole-joint statement speaks\n            # for all three. An unlocalised cartilage remark is weaker evidence but is\n            # carried at a discount rather than dropped; no corpus wording is shown.\n            if pol == \"positive\":\n                g_pos += 1\n                g_best = max(g_best, sev if GLOBAL_OA.search(c) else sev * 0.7)\n            elif pol == \"negative\":\n                g_neg += 1\n            continue\n\n        for t in hits:\n            if pol == \"positive\":\n                acc[t][\"pos\"] += 1\n                acc[t][\"best\"] = max(acc[t][\"best\"], sev)\n            elif pol == \"negative\":\n                acc[t][\"neg\"] += 1\n            else:\n                acc[t][\"unc\"] += 1\n                acc[t][\"best\"] = max(acc[t][\"best\"], 0.30)\n\n    out = {}\n    for t in OA_TARGETS:\n        a = acc[t]\n        pos, neg, unc, best = a[\"pos\"], a[\"neg\"], a[\"unc\"], a[\"best\"]\n        if not (pos or unc) and g_pos and FEATURES[\"oa_inherit\"]:\n            # Inherit the whole-joint statement, unless this compartment was separately\n            # and explicitly cleared.\n            if neg:\n                score, conf = _grade(0, neg, 0, 0.0)\n                score = max(score, 0.35)\n                conf *= 0.7\n            else:\n                score, conf = _grade(g_pos, 0, 0, g_best * 0.92)\n                conf *= 0.75\n        else:\n            score, conf = _grade(pos, neg + g_neg, unc, best)\n        out[t] = (score, conf, pos, neg)\n    return out\n\n\ndef extract(report: str) -> dict:\n    \"\"\"Twelve (score, confidence) pairs, plus the counts the coverage gauge reads.\"\"\"\n    cls = clauses(report)\n    out = {}\n\n    for tgt in PAIRED:\n        s, c, npos, nneg = _score_paired(cls, tgt)\n        out[tgt] = s\n        out[tgt + \"__conf\"] = c\n        out[tgt + \"__npos\"] = npos\n        out[tgt + \"__nneg\"] = nneg\n\n    for tgt, (s, c, npos, nneg) in _score_oa(cls).items():\n        out[tgt] = s\n        out[tgt + \"__conf\"] = c\n        out[tgt + \"__npos\"] = npos\n        out[tgt + \"__nneg\"] = nneg\n\n    for tgt in (\"Effusion\", \"Synovitis\", \"Baker's\", \"Contusion\", \"Fracture\"):\n        if tgt == \"Contusion\":\n            s, c, npos, nneg = _score_clauses(\n                cls, DIRECT_MATCH[tgt], None, DECOY.get(tgt),\n                context_penalty=DEGENERATIVE_MARROW, context_bonus=TRAUMA)\n        else:\n            s, c, npos, nneg = _score_clauses(cls, DIRECT_MATCH[tgt], None,\n                                              DECOY.get(tgt))\n        out[tgt] = s\n        out[tgt + \"__conf\"] = c\n        out[tgt + \"__npos\"] = npos\n        out[tgt + \"__nneg\"] = nneg\n\n    # --- synovitis backoff --------------------------------------------------- #\n    # Eighty-eight per cent of reports never write the word, and the annotator marks it\n    # on nearly half the studies: the label cannot be read off the term alone. What the\n    # report does say is whether the joint is wet and irritated - an effusion, a\n    # distended bursa, an inflamed fat pad - and that ordering is the only signal\n    # available on the silent majority. It enters at low confidence, so it shapes the\n    # ranking without asserting a finding.\n    if (FEATURES[\"synovitis_backoff\"] and out[\"Synovitis__npos\"] == 0\n            and out[\"Synovitis__nneg\"] == 0):\n        proxy = sum(1 for c in cls if SYNOVIAL_PROXY.search(c)\n                    and _polarity(c) == \"positive\")\n        eff = out[\"Effusion\"]\n        prior = 0.30 + 0.30 * max(0.0, (eff - 0.5) / 0.45) + 0.06 * min(proxy, 3)\n        out[\"Synovitis\"] = min(0.72, prior)\n        out[\"Synovitis__conf\"] = 0.18\n\n    return out"},{"cell_type":"markdown","metadata":{},"source":"### 2.2 How an extractor like this is validated\n\nThis is the part that decides whether any of the above is worth trusting, and it is harder\nthan writing the rules, because the obvious measurement cannot carry the weight.\n\n**The dangerous failure is silent.** A rule that never fires does not raise an error — it\nemits a negative, indistinguishable from a confident one. A lexicon complete in English and\nthin in Greek does not look broken; it looks like a corpus in which Greek patients have\nfewer findings. And the error is not random: language tracks the reporting institution,\nwhich tracks the scanner and the population, so a gap in one language is a bias aligned with\na site, not noise that averages out.\n\n**Gauge one — agreement, on the 58 annotated studies.** For each target, compare the\nextracted score against the annotation and read the AUC. This measures exactly the right\nthing and is nearly useless for tuning, because the subset is tiny. The Hanley–McNeil\napproximation for the standard error of an AUC $A$ with $n_p$ positives and $n_n$ negatives,\n\n$$\\mathrm{SE}(A)=\\sqrt{\\frac{A(1-A)+(n_p-1)(Q_1-A^{2})+(n_n-1)(Q_2-A^{2})}{n_p n_n}},\\qquad\nQ_1=\\frac{A}{2-A},\\quad Q_2=\\frac{2A^{2}}{1+A},$$\n\nat $A\\approx0.85$ with nine positives among 58 studies lands near $0.083$, so a 95% interval\nspans roughly $\\pm0.16$. Competitions are decided by differences an order of magnitude\nsmaller. Choosing between two lexicons on a single one of these numbers is choosing by coin\nflip, and it will feel like signal every time. The intervals are drawn below so that this\nis visible rather than stated.\n\n**Gauge two — coverage, on all 4 407 reports.** For each (study, target) pair, record\nwhether *any* rule fired — assertion, negation or hedge. The **silence rate** is the fraction\nwhere none did. It needs no labels, so it runs on the whole corpus rather than on the\nannotated handful, and it points straight at missing vocabulary.\n\n| | measures | sample | can decide |\n|---|---|---|---|\n| agreement | is a fired rule *right* | 58 | whether a target's labels are usable at all |\n| silence | does a rule *fire* | 4 407 | which finding to work on next |\n\nNeither substitutes for the other, and the second is the one to steer by. A common finding\nthat is silent in one language and not another is a lexicon gap. A rare finding that is\nsilent nearly everywhere is simply rare, and silence there is correct.\n\n**One limit worth naming.** The 58 annotations were read from the images; the reports were\nwritten by a different radiologist on a different day. They can disagree about findings such as fluid or a cyst even when both readings are\nreasonable; this documentation intentionally does not reproduce the underlying report text. That disagreement is not extractor error and no\nlexicon can remove it. It puts a ceiling on gauge one somewhere well below 1.0, which is\nanother reason to read the intervals rather than the point estimates."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import os\nimport time\nimport warnings\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\n\nwarnings.filterwarnings(\"ignore\")\nT0 = time.time()\n\n\ndef log(msg):\n    print(f\"[{time.time() - T0:7.1f}s] {msg}\", flush=True)\n\n\ndef find_root():\n    \"\"\"Locate the competition mount, wherever it was attached.\"\"\"\n    for c in [Path(\"/kaggle/input/rsna-knee-abnormality-detection\"),\n              Path(\"/kaggle/input/competitions/rsna-knee-abnormality-detection\"),\n              Path(\"data\"), Path(\".\")]:\n        if (c / \"test.csv\").is_file() and (c / \"test_series\").is_dir():\n            return c\n    base = Path(\"/kaggle/input\")\n    if base.is_dir():\n        for d1 in sorted(p for p in base.iterdir() if p.is_dir()):\n            for cand in [d1] + sorted(p for p in d1.iterdir() if p.is_dir()):\n                if (cand / \"test.csv\").is_file():\n                    return cand\n    raise FileNotFoundError(\"competition mount not found\")\n\n\nROOT = find_root()\nlog(f\"input root: {ROOT}\")\n\n# Safety contract: a diagnostic fallback may exist, but the competition filename is\n# materialised only after every stage, arm and final matrix passes the audit below.\n_test_df = pd.read_csv(ROOT / \"test.csv\")\n_bench = _test_df[[\"StudyInstanceUID\"]].copy()\nfor _c in TARGETS:\n    _bench[_c] = 0.5\nfor _p in (Path(\"submission.csv\"), Path(\"_submission_candidate.csv\")):\n    _p.unlink(missing_ok=True)\n_bench.to_csv(\"submission_fallback.csv\", index=False)\nlog(f\"diagnostic fallback written ({len(_bench)} rows); no submission.csv yet\")\n\nSTAGE_OK = {}\n\n\ndef stage(name):\n    \"\"\"Run one required stage and re-raise any failure.\"\"\"\n    def deco(fn):\n        def run(*a, **k):\n            t = time.time()\n            try:\n                out = fn(*a, **k)\n                STAGE_OK[name] = True\n                log(f\"stage '{name}' ok in {time.time() - t:.1f}s\")\n                return out\n            except Exception:\n                import traceback\n                traceback.print_exc()\n                STAGE_OK[name] = False\n                Path(\"submission.csv\").unlink(missing_ok=True)\n                Path(\"_submission_candidate.csv\").unlink(missing_ok=True)\n                log(f\"stage '{name}' FAILED after {time.time() - t:.1f}s; aborting\")\n                raise\n        return run\n    return deco\n\n\ntrain_df = pd.read_csv(ROOT / \"train.csv\")\nlog(f\"train {train_df.shape}  test {_test_df.shape}\")\n\nt = time.time()\nLAB = pd.DataFrame([extract(r) for r in train_df[\"Report\"].fillna(\"\")])\nLAB[\"StudyInstanceUID\"] = train_df[\"StudyInstanceUID\"].values\nLAB = LAB.set_index(\"StudyInstanceUID\")\nlog(f\"read {len(LAB)} reports in {time.time() - t:.1f}s\")\n\nGOLD = train_df.dropna(subset=TARGETS).set_index(\"StudyInstanceUID\")[TARGETS]\nlog(f\"{len(GOLD)} studies carry the twelve annotations\")\n\npos = (LAB[TARGETS] > 0.5).mean()\nsil = pd.Series({t_: float(((LAB[t_ + \"__npos\"] == 0) & (LAB[t_ + \"__nneg\"] == 0)).mean())\n                 for t_ in TARGETS})\nprint(pd.DataFrame({\"derived positive rate\": pos.round(3),\n                    \"silence rate\": sil.round(3),\n                    \"annotated positive rate\": GOLD.mean().round(3)}).to_string())"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import matplotlib.pyplot as plt\nfrom sklearn.metrics import roc_auc_score\n\nplt.rcParams.update({\"figure.dpi\": 120, \"font.size\": 8, \"axes.grid\": True,\n                     \"grid.alpha\": 0.25, \"axes.spines.top\": False,\n                     \"axes.spines.right\": False})\nINK, ACC, WARN = \"#22303f\", \"#2b7a9b\", \"#c25a3d\"\n\n\ndef agreement(lab, n_boot=2000, seed=0):\n    \"\"\"Per-target AUC against the annotated studies, with a bootstrap interval.\"\"\"\n    rng = np.random.default_rng(seed)\n    g = lab.loc[GOLD.index]\n    rows = []\n    for t_ in TARGETS:\n        y = GOLD[t_].values.astype(int)\n        p = g[t_].values\n        if len(set(y)) < 2:\n            rows.append((t_, np.nan, np.nan, np.nan, int(y.sum()), int((1 - y).sum())))\n            continue\n        a = roc_auc_score(y, p)\n        bs = []\n        for _ in range(n_boot):\n            i = rng.integers(0, len(y), len(y))\n            if len(set(y[i])) > 1:\n                bs.append(roc_auc_score(y[i], p[i]))\n        rows.append((t_, a, np.percentile(bs, 2.5), np.percentile(bs, 97.5),\n                     int(y.sum()), int((1 - y).sum())))\n    return pd.DataFrame(rows, columns=[\"target\", \"auc\", \"lo\", \"hi\", \"npos\", \"nneg\"])\n\n\nAGREE = agreement(LAB).dropna(subset=[\"auc\"])\nprint(AGREE.round(3).to_string(index=False))\nprint(f\"\\nmacro agreement AUC: {AGREE.auc.mean():.4f}   \"\n      f\"mean silence rate: {sil.mean() * 100:.1f}%\")\n\nfig, ax = plt.subplots(1, 2, figsize=(10.5, 3.6),\n                       gridspec_kw={\"width_ratios\": [1.25, 1]})\no = AGREE.sort_values(\"auc\")\ny = np.arange(len(o))\nax[0].hlines(y, o.lo, o.hi, color=ACC, lw=3, alpha=.35)\nax[0].plot(o.auc, y, \"o\", color=ACC, ms=5)\nax[0].axvline(0.5, color=INK, lw=.8, ls=\":\")\nax[0].axvline(o.auc.mean(), color=WARN, lw=1, ls=\"--\")\nax[0].text(o.auc.mean(), -0.9, f\" macro {o.auc.mean():.3f}\", color=WARN, fontsize=7)\nax[0].set_ylim(-1.4, len(o) - 0.4)\nax[0].set_yticks(y)\nax[0].set_yticklabels([f\"{t_}  ({p}+/{n}-)\" for t_, p, n in zip(o.target, o.npos, o.nneg)])\nax[0].set_xlim(0.35, 1.02)\nax[0].set_xlabel(\"AUC of the derived score against the annotation\")\nax[0].set_title(\"gauge one: agreement, n = 58\\n\"\n                \"bars are 95% bootstrap intervals — they are this wide on purpose\",\n                loc=\"left\", fontsize=8)\n\no2 = sil.sort_values()\nax[1].barh(np.arange(len(o2)), o2.values * 100, color=INK, alpha=.8, height=.65)\nax[1].set_yticks(np.arange(len(o2)))\nax[1].set_yticklabels(o2.index)\nax[1].set_xlabel(\"% of studies where no rule fired at all\")\nax[1].set_title(\"gauge two: coverage, n = 4 407\\n\"\n                \"no labels needed, so it runs on the whole corpus\",\n                loc=\"left\", fontsize=8)\nfor i, v in enumerate(o2.values * 100):\n    ax[1].text(v + 1, i, f\"{v:.0f}\", va=\"center\", fontsize=6.5, color=INK)\nfig.tight_layout()\nplt.show()"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"def macro_and_silence():\n    lab = pd.DataFrame([extract(r) for r in train_df[\"Report\"].fillna(\"\")])\n    lab[\"StudyInstanceUID\"] = train_df[\"StudyInstanceUID\"].values\n    lab = lab.set_index(\"StudyInstanceUID\")\n    g = lab.loc[GOLD.index]\n    a = float(np.nanmean([\n        roc_auc_score(GOLD[t_].values.astype(int), g[t_].values)\n        if GOLD[t_].nunique() > 1 else np.nan for t_ in TARGETS]))\n    s = float(np.mean([((lab[t_ + \"__npos\"] == 0) & (lab[t_ + \"__nneg\"] == 0)).mean()\n                       for t_ in TARGETS]))\n    return a, s\n\n\nbase_a, base_s = macro_and_silence()\nrows = [(\"all rules on\", base_a, base_s, 0.0, 0.0)]\nfor k in list(FEATURES):\n    FEATURES[k] = False\n    a, s = macro_and_silence()\n    rows.append((f\"without {k}\", a, s, a - base_a, s - base_s))\n    FEATURES[k] = True\n\nABL = pd.DataFrame(rows, columns=[\"configuration\", \"macro AUC\", \"mean silence\",\n                                  \"d(AUC)\", \"d(silence)\"])\nprint(ABL.round(4).to_string(index=False))"},{"cell_type":"markdown","metadata":{},"source":"Read that table against §2.2 rather than as a ranking. Three of the five deltas are smaller\nthan the sampling error of a 58-study AUC, so the honest reading is:\n\n* **`oa_inherit` is the one change large enough to see on gauge one.** It is also the one\n  whose effect is visible on gauge two, and the two agree in direction. That is the only\n  combination that justifies confidence at this sample size.\n* **`unwrap` barely moves agreement and clearly improves coverage.** Kept for the second\n  reason. Wrapped lines are a property of how a site exports its reports, not of what those\n  reports say, and a rule that repairs the export cannot be wrong about the medicine.\n* The rest are kept because the sentence-level argument for them is sound and their\n  measured effect is not negative — not because 58 studies said so.\n\nMost of the total quality of these labels is not in this table at all: it is in the\nvocabulary, and vocabulary shows up on the silence rate rather than on the ablation. The\nfindings the reports mention least are the ones §1 says are most expensive to leave at\nchance, and they are exactly where the silence bars are tallest — `Fracture`,\n`Lateral OA`, `Synovitis`. That is the standing to-do list for this notebook, and it is\nreadable from a gauge that needs no labels at all.\n\nTwo further points about the targets before moving to the pixels.\n\n**Confidence becomes a sample weight.** Every target carries a confidence alongside its\nscore: high when several clauses agree, low when a finding was named once in passing, and\nvery low when nothing fired. That number multiplies the per-element loss in §7, so silence\npulls weakly rather than asserting a negative.\n\n**Synovitis is a special case, handled explicitly.** Nearly nine studies in ten never use\nthe word, and the annotator marks it on close to half. No term list can close that. What a\nreport does say is whether the joint is wet and irritated — an effusion, a distended bursa,\nan inflamed fat pad — and on the silent majority that ordering is the only signal there is.\nIt enters at low confidence, so it shapes the ranking without asserting a finding. This is\nthe weakest of the twelve target derivations and the notebook says so rather than hiding it\nin an average."},{"cell_type":"markdown","metadata":{},"source":"### 2.3 Reading the coverage gauge, which is the point of having it\n\nThe silence rate is only useful if it is broken down. Aggregated over the corpus it says\n\"this target is thin\", which anyone could have guessed. Broken down by language it says\n*where the vocabulary is missing*, and that is a to-do list.\n\nThe two figures below are the ones that actually drove the last several revisions of the\nlexicon, so it is worth being precise about how to read them — including the way the left\none can mislead.\n\n**Silence has two causes, and only one of them is a bug.** A target is silent either\nbecause the report never discusses that structure, or because it does and no rule fired.\nThe first is correct behaviour and no amount of vocabulary fixes it; the second is a gap.\nThey look identical in the left figure and are separated in the right one, by asking a\nquestion that needs no labels: does the report contain *any* word for this structure at\nall?\n\nFor the Spanish cruciate cell that split came out 62% never-mentioned against 38%\nmentioned-and-missed. The 62% are short reports — a median of 336 characters against 1257\nfor the ones that do mention it — stating findings and staying silent on everything\nnormal. Nothing there to fix. The 38% were a genuine gap, and reading a handful of them\nshowed what it was:\n\n> *Synthetic paraphrase:* `No significant abnormality is identified in the cruciate or collateral ligaments.`\n\nThis synthetic paraphrase represents a negated-abnormality construction that asserts\nnormal cruciates. The extractor initially missed that construction because a negator plus\nan abnormality noun is literally a negation, not an explicit normality phrase, and the\nnormality rule refused to fire when a negator was present. So a ligament a radiologist had looked at and called intact was\nscored the same as one the report never mentioned. To a rank metric those are not the\nsame: explicitly clear must rank *below* unmentioned.\n\nComparable negated-abnormality constructions occur across the represented languages;\nthe exact corpus wording is omitted here. Adding those constructions dropped the mean silence over all 4 407 studies and twelve\ntargets from 49.97% to 48.09%, concentrated exactly where predicted: Greek −15 points on\nthe ligament and meniscus targets, Spanish −6, Bulgarian/Russian −4, Turkish and Croatian\nunchanged because they were already covered by a different construction.\n\nThat cell is now essentially closed. Spanish studies silent on the cruciate fell from 215\nto 141, and of the 141 that remain, **5%** name a cruciate at all — a median report length\nof 230 characters says the rest simply do not discuss it. The work moved from the\nfixable half of the bar to the half that is not a bug.\n\nThe same figure then pointed at a Dutch plural construction meaning normal menisci but\nnaming no side; its exact corpus wording is omitted here, and it left both meniscal\ntargets silent. The cruciates and the collaterals had\ntheir plural forms from the start; the menisci had simply been missed.\n\n**What gauge one said about all of this: nothing.** Agreement on the 58 annotated studies\nmoved by −0.0007 macro AUC, 95% CI [−0.0035, +0.0009] — a flat zero. That is not evidence\nthe change was worthless. It is evidence that 58 studies, most of them English, cannot see\na Greek and Spanish coverage fix at all. Steering by that number would have rejected the\nwork; steering by the gauge that runs on all 4 407 found it."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"_SCRIPT = {\"el\": re.compile(r\"[Ͱ-Ͽ]\"), \"bg/ru\": re.compile(r\"[Ѐ-ӿ]\")}\n_STOP = {\n    \"en\": r\"\\b(the|and|is|with|there is|normal)\\b\",\n    \"es\": (r\"\\b(del|los|las|con|sin|senal|rodilla|hallazgos|tecnica|resultados\"\n           r\"|impresion|menisco|rotura)\\b\"),\n    \"fr\": r\"\\b(des|les|avec|sans|genou|aucune)\\b\",\n    \"nl\": r\"\\b(van|het|een|geen|met|voorste|knie)\\b\",\n    \"de\": r\"\\b(der|die|und|mit|ohne|kein|keine|nachweis)\\b\",\n    \"tr\": r\"\\b(ve|ile|izlenmistir|mevcut|normaldir|diz|bulgular)\\b\",\n    \"hr\": r\"\\b(se|te|uz|bez|prikaz|uredan|koljena|meniska)\\b\",\n}\n_STOP = {k: re.compile(v) for k, v in _STOP.items()}\n\n\ndef guess_language(report):\n    \"\"\"A crude language tag, used only to *audit* the reading - never to do it.\n\n    §2 argued against routing a report to a per-language rule set, because that commits\n    to a guess before any evidence is read. None of that applies here: this classifier\n    never touches extraction. It exists so the silence rate can be broken down, and a\n    tag that is wrong now and then blurs the breakdown rather than corrupting a label.\n    Script settles Greek and Cyrillic outright; the Latin-script languages are separated\n    by counting function words, which is ugly and sufficient for a histogram.\n    \"\"\"\n    n = normalize(report)\n    for tag, rx in _SCRIPT.items():\n        if rx.search(n):\n            return tag\n    score = {k: len(rx.findall(n)) for k, rx in _STOP.items()}\n    best = max(score, key=score.get)\n    return best if score[best] >= 2 else \"?\"\n\n\nLANG = pd.Series([guess_language(r) for r in train_df[\"Report\"].fillna(\"\")],\n                 index=train_df[\"StudyInstanceUID\"])\nprint(LANG.value_counts().to_string())\n\nSIL = pd.DataFrame({t: ((LAB[t + \"__npos\"] == 0) & (LAB[t + \"__nneg\"] == 0)).values\n                    for t in TARGETS}, index=LAB.index)\nby_lang = SIL.groupby(LANG.reindex(SIL.index).values).mean() * 100\nby_lang = by_lang.loc[LANG.value_counts().index.intersection(by_lang.index)]\n\nfig, ax = plt.subplots(1, 2, figsize=(12, 3.6),\n                       gridspec_kw={\"width_ratios\": [1.7, 1]})\nim = ax[0].imshow(by_lang.values, cmap=\"RdYlBu_r\", vmin=0, vmax=100, aspect=\"auto\")\nax[0].set_xticks(range(len(TARGETS)))\nax[0].set_xticklabels(TARGETS, rotation=55, ha=\"right\", fontsize=6.5)\nax[0].set_yticks(range(len(by_lang)))\nax[0].set_yticklabels([f\"{l}  (n={int((LANG == l).sum())})\" for l in by_lang.index],\n                      fontsize=7)\nax[0].grid(False)\nfor i in range(by_lang.shape[0]):\n    for j in range(by_lang.shape[1]):\n        v = by_lang.values[i, j]\n        ax[0].text(j, i, f\"{v:.0f}\", ha=\"center\", va=\"center\", fontsize=5.5,\n                   color=\"white\" if v > 62 or v < 12 else INK)\nax[0].set_title(\"gauge two, broken down: % of studies where no rule fired\\n\"\n                \"a common finding silent in one language and not another is a \"\n                \"lexicon gap — or a reporting style\", loc=\"left\", fontsize=8)\nfig.colorbar(im, ax=ax[0], fraction=0.02, pad=0.01)\n\n# Silence has two causes needing different work, so split them: of the studies where\n# nothing fired, in how many does the report name *this* structure anyway?\n#\n# The test has to be the target's own anatomy matcher, not a bare stem for the organ. A\n# report discussing only the medial meniscus contains the word \"meniscus\", and counting\n# that as a missed *lateral* meniscus would score correct silence as a bug - the wrong\n# direction entirely for a gauge whose job is to find real gaps.\nCLAUSES = {u: clauses(r) for u, r in\n           zip(train_df[\"StudyInstanceUID\"], train_df[\"Report\"].fillna(\"\"))}\n\n\ndef names_it(uid, target):\n    cs = CLAUSES[uid]\n    if any(ANAT_MATCH[target].search(c) for c in cs):\n        return True\n    if \"Meniscus\" in target:\n        return any(PLURAL_MENISCI.search(c) and not ANY_SIDE.search(c) for c in cs)\n    return False\n\n\nrows = []\nfor t_ in [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\"]:\n    sel = SIL[t_].values\n    if not sel.sum():\n        continue\n    named = np.array([names_it(u, t_) for u in SIL.index[sel]])\n    rows.append((t_, int(sel.sum()), 100 * named.mean()))\nD = pd.DataFrame(rows, columns=[\"target\", \"silent\", \"names it anyway\"])\ny = np.arange(len(D))\nax[1].barh(y, 100 - D[\"names it anyway\"], color=\"#8a97a3\", label=\"never mentioned\")\nax[1].barh(y, D[\"names it anyway\"], left=100 - D[\"names it anyway\"], color=WARN,\n           label=\"mentioned, missed\")\nax[1].set_yticks(y)\nax[1].set_yticklabels([f\"{t_}\\n({n} silent)\" for t_, n in zip(D.target, D.silent)],\n                      fontsize=6.5)\nax[1].set_xlabel(\"% of the silent studies\")\nax[1].legend(fontsize=6.5, frameon=False, loc=\"lower right\")\nax[1].set_title(\"why it was silent\\nonly the orange half is a lexicon gap\",\n                loc=\"left\", fontsize=8)\nfig.tight_layout()\nplt.show()\nprint(D.round(1).to_string(index=False))"},{"cell_type":"markdown","metadata":{},"source":"## 3. What the scanner recorded\n\n`train_series.csv` describes each series with three flags — `Anatomical_Plane`,\n`Fluid_Sensitive`, `Fat_Suppression`. A study holds several series and the twelve findings\nare not read on the same ones: cruciates sagittally, the collateral ligaments and the\nmeniscal body coronally, patellar cartilage axially, and anything involving fluid on a\nfat-suppressed fluid-sensitive sequence, where marrow oedema and effusion light up and fat\ndoes not.\n\nSo a study is not a stack of images. It is a **bag of up to six slots**, one per\n(plane × contrast) combination, with a mask saying which of them this study actually has.\nThe mask is not a formality: the fat-suppressed fluid-sensitive series exist for nearly\nevery study, while T1 and non-suppressed series are far scarcer, and a model that averaged\nover missing slots would be averaging over zeros.\n\nThe flags in the CSV are recovered again from the DICOM headers rather than taken as given,\nfor one specific reason: two axes are collapsed into one flag. `Fluid_Sensitive = 0`\ncovers both T1 and non-fat-suppressed proton-density series, which carry very different\ntissue contrast. Reading `SeriesDescription`, `ScanningSequence`, `RepetitionTime` and\n`EchoTime` separates them, and the header pass is needed anyway for pixel spacing and\nlaterality. Both slot schemes are kept below as a switch, so the choice can be varied with\neverything else held fixed.\n\nOne header trap is worth naming because it is silent: GE writes `SAT_GEMS` into\n`ScanOptions` for *spatial* saturation. A substring test for `SAT` marks those series\nfat-suppressed when they are not, so `ScanOptions` is matched as exact tokens."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import gc\nimport re as _re\nfrom concurrent.futures import ThreadPoolExecutor\n\nimport pydicom\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\n\nfor _v in (\"OMP_NUM_THREADS\", \"OPENBLAS_NUM_THREADS\", \"MKL_NUM_THREADS\"):\n    os.environ.setdefault(_v, \"4\")\n\nSEED = 2026\nnp.random.seed(SEED)\ntorch.manual_seed(SEED)\n\n# --- geometry -------------------------------------------------------------- #\n# The centre crop has to be smaller than the field of view of nearly every series, or it\n# silently does nothing: where the requested crop is larger than the image, the crop is\n# skipped and that series keeps its physical scale unnormalised - with no error and no\n# log line. The figure two cells below plots the acquired field of view\n# (Rows x PixelSpacing) over the whole training corpus and marks this value against it,\n# together with the exact percentage of series it would fail to crop, so the choice can\n# be checked rather than taken on trust. 130 mm still contains the joint.\nCROP_MM = 130.0\n\n# --- cache ----------------------------------------------------------------- #\n# bytes = n_study x n_slot x slices x P^2. The exponent on P is what makes this a real\n# constraint: the cache grows with the SQUARE of resolution and only linearly with\n# slices, so coverage is the cheap axis and resolution the expensive one.\n#\n# Measured, not assumed. Resolution was tested twice at fixed coverage - 224 px against\n# 266 px - and moved the holdout by a thousandth of macro AUC both times, inside the\n# noise. Coverage was then tested at fixed resolution and moved it by 0.033, and the\n# curve was still climbing at the last point measured: 3 -> 6 slices gained 0.019, and\n# 6 -> 12 gained a further 0.014.\n#\n# So the trade is pushed one step further in the direction the evidence points: a coarser\n# grid buying twice the slices again. That is an extrapolation, and it has a way of being\n# wrong - resolution not mattering between 224 and 266 does not prove it stops mattering\n# below 224, and somewhere there is a floor. So the arms below include a 12-slice arm at\n# this grid, directly comparable with the 12-slice arm of the previous version at 224 px.\n# If that comparison drops, the floor has been found and the notebook says so.\n# 168 = 12 x 14 keeps the DINOv2 patch grid exact.\nCACHE_IMG = 168\nGROUP = 3                  # slices per encoder input, stacked as the three channels\nN_GROUP_MAX = 8            # -> 24 slices per slot\nCACHE_BUDGET_GB = float(os.environ.get(\"CACHE_BUDGET_GB\", 17.0))\nHDR_THREADS = 16\nPIX_THREADS = 12\nORDER_THREADS = 32         # slice ordering is latency-bound on the mount, not CPU-bound\nORDER_BUDGET_S = 3000     # the real corpus used 2318s of a 2400s cap: too close\n\n# --- training -------------------------------------------------------------- #\nEPOCHS = int(os.environ.get(\"EPOCHS\", 16))\nBATCH_STUDIES = 8\nLR_HEAD = 1e-3\nLR_BACKBONE = 8e-6         # the encoder is adapted, not retrained\nUNFREEZE_LAST = 6\nWEIGHT_DECAY = 0.02\nEVAL_BATCH = 8\nTIME_BUDGET = float(os.environ.get(\"TIME_BUDGET_H\", 7.6)) * 3600\nSMOKE = int(os.environ.get(\"SMOKE\", 0))\n\n# Arms, cheapest first. Each one checks the remaining budget before it starts. Completed\n# arms refresh a diagnostic partial ensemble only; an incomplete run never exposes the\n# ordinary submission filename and fails the final safety gate.\n#\n# The first three arms are a controlled coverage experiment: identical encoder, identical\n# grid, identical cache, identical cost per epoch - they differ only in how many of the\n# cached slice groups each may draw from and average over. Three points rather than two,\n# because two points cannot tell \"still climbing\" from \"saturated\", and that distinction\n# is the whole guidance for whoever runs this next.\n#\n# `large` was tried and is not here. It scored below `base` on the holdout while costing\n# two and a half hours, so capacity is not monotonic on this corpus and the budget is\n# better spent elsewhere. The second `base` arm differs only by seed, which buys ensemble\n# diversity at a known cost rather than a speculative one.\nARMS = [\n    # the resolution safety check: same 12 slices as the previous version, coarser grid\n    {\"name\": \"s168g4\", \"variant\": \"small\", \"img\": 168, \"lr_bb\": 8e-6, \"groups\": 4},\n    # the coverage step\n    {\"name\": \"s168g8\", \"variant\": \"small\", \"img\": 168, \"lr_bb\": 8e-6, \"groups\": 8},\n    {\"name\": \"b168g8\", \"variant\": \"base\", \"img\": 168, \"lr_bb\": 5e-6, \"groups\": 8},\n    {\"name\": \"b168g8b\", \"variant\": \"base\", \"img\": 168, \"lr_bb\": 5e-6, \"groups\": 8,\n     \"seed\": 7},\n    {\"name\": \"b168g8c\", \"variant\": \"base\", \"img\": 168, \"lr_bb\": 5e-6, \"groups\": 8,\n     \"seed\": 13},\n]\nif SMOKE:\n    # A smoke run keeps the whole control path and shrinks only the work: one epoch and\n    # one arm by default, but both overridable, so the arm bookkeeping - deduplication,\n    # selection, per-arm groups - can be exercised without a GPU.\n    EPOCHS = int(os.environ.get(\"EPOCHS\", 1))\n    ARMS = ARMS[:int(os.environ.get(\"SMOKE_ARMS\", 1))]\n\nSLOTS_RECOVERED = [\n    (\"SAG_FLUID_FS\", \"Sagittal\", True, True),\n    (\"COR_FLUID_FS\", \"Coronal\", True, True),\n    (\"AX_FLUID_FS\", \"Axial\", True, True),\n    (\"SAG_FLUID_NOFS\", \"Sagittal\", True, False),\n    (\"COR_T1\", \"Coronal\", False, False),\n    (\"SAG_T1\", \"Sagittal\", False, False),\n]\n# The alternative: plane x the single delivered flag, ignoring the recovered weighting.\n# Under this scheme one slot mixes T1 with non-fat-suppressed PD/T2.\nSLOTS_PUBLIC = [\n    (\"SAG_FLUID\", \"Sagittal\", None, True),\n    (\"COR_FLUID\", \"Coronal\", None, True),\n    (\"AX_FLUID\", \"Axial\", None, True),\n    (\"SAG_STRUCT\", \"Sagittal\", None, False),\n    (\"COR_STRUCT\", \"Coronal\", None, False),\n    (\"AX_STRUCT\", \"Axial\", None, False),\n]\nSLOTS = SLOTS_PUBLIC if os.environ.get(\"SLOT_SCHEME\") == \"public\" else SLOTS_RECOVERED\nN_SLOT = len(SLOTS)\n\nFATSAT_OPTS = {\"FS\", \"FATSAT\", \"FAT_SAT\", \"FSAT\"}\n_SEP = _re.compile(r\"[_\\-.]\")\n_FATSAT_RX = _re.compile(r\"\\bfs\\b|fatsat|fat sat|\\bstir\\b|\\bspair\\b|\\bspir\\b|\\bwe\\b|\"\n                         r\"water excit|\\btirm\\b|\\bsting\\b|\\bfatsup\\b\")\n_T1_RX = _re.compile(r\"\\bt1\\b|\\bt1w\\b\")\n_T2_RX = _re.compile(r\"\\bt2\\b|\\bt2w\\b\")\n_PD_RX = _re.compile(r\"\\bpd\\b|\\bpdw\\b|proton|\\bdp\\b|dens\")\n\ndef gpu_probe():\n    \"\"\"Check that the GPU can run a kernel - before an hour of I/O is spent assuming it.\n\n    This is not paranoia, it is a bug this notebook hit. Kaggle's generic `GPU`\n    accelerator may allocate a P100, and current PyTorch builds no longer ship kernels\n    for Pascal (sm_60). Every symptom you would test for looks healthy:\n    `cuda.is_available()` is True, the device name and capability query fine, tensors\n    allocate, and the model moves onto the device without complaint. The *first kernel\n    launch* then fails with `no kernel image is available for execution on the device`.\n\n    Discovering that after the header pass and the decode costs 56 minutes and the whole\n    run. Discovering it with a 64x64 matmul costs a millisecond, so it is done here,\n    before anything expensive, and it launches a real kernel rather than asking a\n    property.\n    \"\"\"\n    if not torch.cuda.is_available():\n        return False, \"no CUDA device visible\"\n    try:\n        name = torch.cuda.get_device_name(0)\n        major, minor = torch.cuda.get_device_capability(0)\n    except Exception as exc:\n        return False, f\"device query failed: {exc}\"\n    try:\n        a = torch.randn(64, 64, device=\"cuda\")\n        float((a @ a).sum().item())\n        float(torch.rand(1, device=\"cuda\").item())\n        return True, f\"{name} (sm_{major}{minor})\"\n    except Exception as exc:\n        return False, (f\"{name} (sm_{major}{minor}) allocates but cannot launch a \"\n                       f\"kernel: {type(exc).__name__}. This torch build has no code \"\n                       f\"for sm_{major}{minor} - select a T4 accelerator.\")\n\n\nGPU_OK, GPU_MSG = gpu_probe()\nDEV = torch.device(\"cuda\" if GPU_OK else \"cpu\")\nlog(f\"torch {torch.__version__}  |  gpu: {GPU_MSG}\")\nlog(f\"device {DEV}  |  arms {[a['name'] for a in ARMS]}  |  epochs {EPOCHS}\"\n    f\"  |  budget {TIME_BUDGET / 3600:.1f} h\")\n\n# Fine-tuning three vision transformers on CPU is not a slower plan, it is not a plan.\n# The candidate preflight normally aborts before this branch. If the later probe becomes\n# unusable, heavy stages remain incomplete and the final gate raises; only the explicitly\n# named diagnostic fallback can remain.\nRUN_HEAVY = bool(GPU_OK or SMOKE)\nif not RUN_HEAVY:\n    log(\"=\" * 72)\n    log(\"NO USABLE GPU - skipping the imaging pipeline.\")\n    log(f\"reason: {GPU_MSG}\")\n    log(\"submission_fallback.csv remains diagnostic; no submission.csv is produced.\")\n    log(\"the report reader needs no accelerator and its measurements stand.\")\n    log(\"=\" * 72)"},{"cell_type":"markdown","metadata":{},"source":"> ### An aside that will cost someone else a day\n>\n> If a GPU notebook in this competition trains fine locally and then dies with\n>\n> ```\n> CUDA error: no kernel image is available for execution on the device\n> ```\n>\n> the accelerator is the problem, not the code. Kaggle's generic **GPU** setting can\n> allocate either a T4 or a **P100**, and the P100 is Pascal — compute capability 6.0. The\n> current container ships PyTorch 2.10, and that build's kernels are compiled for:\n>\n> ```\n> ['sm_70', 'sm_75', 'sm_80', 'sm_86', 'sm_90', 'sm_100', 'sm_120']\n> ```\n>\n> There is no `sm_60` in that list, so nothing Pascal can run at all.\n>\n> What makes it expensive is that every check you would think to make passes.\n> `torch.cuda.is_available()` is `True`, the device name and capability query fine, tensors\n> allocate, and a model moves onto the device without a murmur. Only the **first kernel\n> launch** fails — which, in a pipeline like this one, is after the header pass and the\n> full decode. That cost this notebook 56 minutes and an entire run before anything\n> trained.\n>\n> Two lines of defence, both cheap:\n>\n> 1. **Pin the accelerator to T4 x2** in the notebook settings rather than leaving it on\n>    the generic GPU option.\n> 2. **Launch a real kernel during setup**, before anything expensive — the `gpu_probe()`\n>    in the configuration cell above does a 64×64 matmul and reports what it found. A\n>    property query is not enough; the failure only appears when a kernel actually runs."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"HDR_TAGS = [\"SeriesDescription\", \"SequenceName\", \"ScanOptions\", \"ScanningSequence\",\n            \"RepetitionTime\", \"EchoTime\", \"Laterality\", \"PixelSpacing\", \"Rows\",\n            \"Columns\"]\n\n\ndef probe(item):\n    \"\"\"One header read per series - the middle slice - for protocol and geometry.\"\"\"\n    split, study, series, path = item\n    row = {\"split\": split, \"StudyInstanceUID\": study, \"SeriesInstanceUID\": series,\n           \"dir\": path}\n    try:\n        files = sorted(e.name for e in os.scandir(path) if e.name.endswith(\".dcm\"))\n        row[\"files\"] = files\n        row[\"n_slices\"] = len(files)\n        if not files:\n            return row\n        ds = pydicom.dcmread(os.path.join(path, files[len(files) // 2]),\n                             stop_before_pixels=True, force=True)\n        for t_ in HDR_TAGS:\n            v = getattr(ds, t_, None)\n            if v is None:\n                row[t_] = None\n            elif isinstance(v, (list, tuple)) or type(v).__name__ == \"MultiValue\":\n                row[t_] = \"|\".join(str(x) for x in v)\n            else:\n                row[t_] = str(v)\n    except Exception as exc:\n        row[\"err\"] = str(exc)[:120]\n    return row\n\n\ndef walk(split):\n    base = ROOT / split\n    items = []\n    if not base.is_dir():\n        return pd.DataFrame()\n    for study in os.scandir(base):\n        if study.is_dir():\n            for series in os.scandir(study.path):\n                if series.is_dir():\n                    items.append((split, study.name, series.name, series.path))\n    with ThreadPoolExecutor(max_workers=HDR_THREADS) as pool:\n        return pd.DataFrame(list(pool.map(probe, items)))\n\n\ndef annotate(df):\n    \"\"\"Recover fat suppression and pulse-sequence weighting from the header.\"\"\"\n    if not len(df):\n        return df\n    desc = (df[\"SeriesDescription\"].fillna(\"\") + \" \" + df[\"SequenceName\"].fillna(\"\"))\n    desc = desc.str.lower().str.replace(_SEP, \" \", regex=True)\n    # GE writes SAT_GEMS for spatial saturation, so ScanOptions is matched as exact\n    # tokens; a substring test on \"SAT\" fires on series that are not fat-suppressed.\n    opts = df[\"ScanOptions\"].fillna(\"\").str.upper().str.split(\"|\")\n    opts_fs = opts.apply(lambda ts: any(x.strip() in FATSAT_OPTS for x in ts))\n    df[\"fatsat\"] = desc.str.contains(_FATSAT_RX) | opts_fs\n\n    tr_ = pd.to_numeric(df[\"RepetitionTime\"], errors=\"coerce\")\n    te_ = pd.to_numeric(df[\"EchoTime\"], errors=\"coerce\")\n    gre = df[\"ScanningSequence\"].fillna(\"\").str.upper().str.contains(\"GR\")\n    t1 = desc.str.contains(_T1_RX)\n    t2 = desc.str.contains(_T2_RX)\n    pdw = desc.str.contains(_PD_RX)\n    df[\"weight\"] = np.where(t1 & ~t2 & ~pdw, \"T1\",\n                     np.where(t2 & ~pdw, \"T2\",\n                       np.where(pdw, \"PD\",\n                         np.where(gre, \"GRE\",\n                           np.where(tr_ < 800, \"T1\",\n                             np.where(te_ > 60, \"T2\",\n                               np.where(tr_ >= 800, \"PD\", \"UNK\")))))))\n    df[\"fluid\"] = np.isin(df[\"weight\"], [\"PD\", \"T2\"])\n    df[\"px\"] = pd.to_numeric(\n        df[\"PixelSpacing\"].fillna(\"\").str.split(\"|\").str[0].replace(\"\", np.nan),\n        errors=\"coerce\")\n    df[\"fov_mm\"] = pd.to_numeric(df[\"Rows\"], errors=\"coerce\") * df[\"px\"]\n    return df\n\n\ndef pick_slots(series_df, plane_map):\n    \"\"\"One series per slot per study.\n\n    Ties break toward the stack with the most slices: a denser stack samples the joint\n    better, and the slice sampler below has more to spread over.\n    \"\"\"\n    series_df = series_df.copy()\n    series_df[\"plane\"] = series_df[\"SeriesInstanceUID\"].map(plane_map)\n    out = {}\n    for study, g in series_df.groupby(\"StudyInstanceUID\"):\n        chosen = {}\n        for name, plane, fluid, fs in SLOTS:\n            sel = (g[\"plane\"] == plane) & (g[\"fatsat\"] == fs)\n            if fluid is not None:\n                sel &= (g[\"fluid\"] == fluid)\n            cand = g[sel]\n            if len(cand) == 0 and fluid is False:\n                # T1 slots are the scarcest; fall back to any non-fat-sat series in the\n                # plane before giving up on the slot entirely.\n                cand = g[(g[\"plane\"] == plane) & (~g[\"fatsat\"])]\n            if len(cand):\n                chosen[name] = cand.sort_values(\"n_slices\", ascending=False).iloc[0]\n        out[study] = chosen\n    return out\n\n\ndef lat_of(h):\n    \"\"\"Study -> 'L' / 'R' / None.\n\n    The tag is present on some studies and absent on others, and is sometimes an empty\n    string rather than missing, which is not the same as NaN.\n    \"\"\"\n    d = {}\n    for st, g in h.groupby(\"StudyInstanceUID\"):\n        v = [str(x).strip().upper() for x in g[\"Laterality\"].dropna()]\n        v = [x[0] for x in v if x and x[0] in (\"L\", \"R\")]\n        d[st] = v[0] if v else None\n    return d\n\n\n@stage(\"headers\")\ndef _headers():\n    global HTR, HTE, SLOTS_TR, SLOTS_TE, LAT_TR, LAT_TE, TEST_DF, TRAIN_SERIES\n    TEST_DF = pd.read_csv(ROOT / \"test.csv\")\n    TRAIN_SERIES = pd.read_csv(ROOT / \"train_series.csv\")\n    test_series = pd.read_csv(ROOT / \"test_series.csv\")\n    both = pd.concat([TRAIN_SERIES, test_series])\n    plane_map = dict(zip(both[\"SeriesInstanceUID\"], both[\"Anatomical_Plane\"]))\n\n    log(\"header pass: test\")\n    HTE = annotate(walk(\"test_series\"))\n    log(f\"  {len(HTE)} test series\")\n    log(\"header pass: train\")\n    HTR = annotate(walk(\"train_series\"))\n    log(f\"  {len(HTR)} train series\")\n\n    SLOTS_TE = pick_slots(HTE, plane_map)\n    SLOTS_TR = pick_slots(HTR, plane_map)\n    LAT_TE, LAT_TR = lat_of(HTE), lat_of(HTR)\n    cov = pd.Series([len(v) for v in SLOTS_TR.values()])\n    log(f\"train slots per study: mean {cov.mean():.2f}  min {cov.min()}  max {cov.max()}\")\n    lat_known = np.mean([v is not None for v in LAT_TR.values()])\n    log(f\"laterality known for {lat_known * 100:.1f}% of training studies\")\n\n\nif RUN_HEAVY:\n    _headers()"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"if STAGE_OK.get(\"headers\"):\n    fig, ax = plt.subplots(1, 3, figsize=(11.5, 3.2),\n                           gridspec_kw={\"width_ratios\": [1.1, 1, 1]})\n\n    # 1. how often each slot exists at all\n    names = [s[0] for s in SLOTS]\n    have = np.array([[1.0 if n in SLOTS_TR[st] else 0.0 for n in names]\n                     for st in SLOTS_TR])\n    frac = have.mean(0) * 100\n    ax[0].barh(np.arange(len(names)), frac, color=INK, alpha=.85, height=.65)\n    ax[0].set_yticks(np.arange(len(names)))\n    ax[0].set_yticklabels(names)\n    ax[0].set_xlabel(\"% of training studies that have this slot\")\n    ax[0].set_xlim(0, 105)\n    for i, v in enumerate(frac):\n        ax[0].text(v + 1.5, i, f\"{v:.0f}\", va=\"center\", fontsize=6.5, color=INK)\n    ax[0].set_title(\"the presence mask is not a formality\", loc=\"left\", fontsize=8)\n\n    # 2. how many slots a study has\n    cnt = have.sum(1)\n    ax[1].hist(cnt, bins=np.arange(-.5, N_SLOT + 1.5), color=ACC, alpha=.85, rwidth=.85)\n    ax[1].set_xlabel(\"slots present per study\")\n    ax[1].set_ylabel(\"studies\")\n    ax[1].set_title(f\"median {np.median(cnt):.0f} of {N_SLOT}\", loc=\"left\", fontsize=8)\n\n    # 3. the field of view, which is what fixes the crop\n    fov = HTR[\"fov_mm\"].dropna()\n    fov = fov[(fov > 40) & (fov < 500)]\n    ax[2].hist(fov, bins=60, color=INK, alpha=.75)\n    ax[2].axvline(CROP_MM, color=WARN, lw=1.6)\n    below = float((fov < CROP_MM).mean() * 100)\n    ax[2].text(CROP_MM + 4, ax[2].get_ylim()[1] * .88,\n               f\"crop {CROP_MM:.0f} mm\\nlarger than the image\\nfor {below:.1f}% of series\",\n               color=WARN, fontsize=6.8)\n    ax[2].set_xlabel(\"acquired field of view, Rows x PixelSpacing (mm)\")\n    ax[2].set_ylabel(\"series\")\n    ax[2].set_title(\"physical scale varies; the crop must fit under it\",\n                    loc=\"left\", fontsize=8)\n    fig.tight_layout()\n    plt.show()\n\n    print(f\"field of view: median {fov.median():.0f} mm, \"\n          f\"1st-99th percentile {fov.quantile(.01):.0f}-{fov.quantile(.99):.0f} mm; \"\n          f\"a {CROP_MM:.0f} mm crop is skipped for {below:.1f}% of series\")\n    print(f\"pixel spacing spans {HTR.px.quantile(.01):.2f} - \"\n          f\"{HTR.px.quantile(.99):.2f} mm  \"\n          f\"({HTR.px.quantile(.99) / max(HTR.px.quantile(.01), 1e-9):.1f}x)\")\n    print(f\"a {CROP_MM:.0f} mm crop at {CACHE_IMG} px is \"\n          f\"{CROP_MM / CACHE_IMG:.3f} mm/pixel; at 224 px it is \"\n          f\"{CROP_MM / 224:.3f} mm/pixel\")\n    print(HTR.groupby(\"weight\").size().sort_values(ascending=False).to_string())"},{"cell_type":"markdown","metadata":{},"source":"## 4. Geometry: order, scale, and which knee\n\nThree things about the pixels have to be fixed before an encoder sees them, and each one\nis a nuisance axis the model cannot observe for itself.\n\n**Slice order.** A DICOM file name here is a SOP Instance UID, assigned arbitrarily.\nSorting by it produces an order uncorrelated with anatomy, and anything downstream that\nassumes the file order means something is then operating on noise: the three channels of a\n\"2.5D\" input become three unrelated views, \"the middle of the stack\" becomes a random\nsubset, and reversing slice order to normalise laterality reverses nothing.\n\nThe physical order is recoverable exactly. Each slice carries its position in patient\ncoordinates and the in-plane axes, so projecting the position onto the slice normal gives a\nsigned through-plane coordinate that is monotonic along the stack:\n\n$$\\mathbf{n} = \\mathbf{r}_x \\times \\mathbf{r}_y, \\qquad k = \\mathbf{p}\\cdot\\mathbf{n}.$$\n\n`InstanceNumber` is the fallback, but only that: interleaved and multi-echo acquisitions\nneed not number slices in the order they occupy in space, and the number is unsigned, while\nthe projection is signed in patient coordinates — which is what laterality normalisation\nneeds.\n\n**Physical scale.** Pixel spacing varies several-fold across this corpus, so a fixed pixel\nresize hands the encoder the same anatomy at different sizes. Cropping a constant physical\nextent first and resizing after makes one millimetre the same number of pixels in every\nstudy. The crop must be smaller than the smallest field of view or it silently does nothing\n— the histogram above is what sets it.\n\n**Which knee.** Four of the twelve targets — the two menisci, the medial and lateral\ncompartments — are medial/lateral pairs, and medial is defined relative to the body's\nmidline. Unless laterality is normalised, those four labels are asking the model to learn\nfrom an axis it cannot see. The correction differs by plane, because the mirror acts on a\ndifferent image axis:\n\n* **Coronal and axial** — the medial-lateral direction lies in the image plane, so a left\n  knee is the horizontal mirror of a right knee. Flip the last axis.\n* **Sagittal** — the medial-lateral direction is the *slice* axis. Each individual slice is\n  unchanged by mirroring; what differs is the direction in which the stack traverses the\n  joint. Reverse the slice order, not the pixels.\n\nWhere `Laterality` is absent the volume is left alone, because a wrong flip is worse than\nno flip. Those studies keep an unresolved side axis, and the cost falls on the four\nside-specific targets as diluted supervision rather than as a corrupted input."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"ORDER_TAGS = [(0x0020, 0x0032), (0x0020, 0x0037), (0x0020, 0x0013)]\n\n\ndef order_slices(rec):\n    \"\"\"Return the series' files sorted along the through-plane axis.\"\"\"\n    files, d = rec[\"files\"], rec[\"dir\"]\n    keyed, ds = [], None\n    for f in files:\n        k = None\n        try:\n            ds = pydicom.dcmread(os.path.join(d, f), force=True, stop_before_pixels=True,\n                                 specific_tags=ORDER_TAGS)\n            iop = np.asarray(ds.ImageOrientationPatient, dtype=float)\n            ipp = np.asarray(ds.ImagePositionPatient, dtype=float)\n            k = float(np.dot(ipp, np.cross(iop[:3], iop[3:])))\n        except Exception:\n            try:\n                k = float(ds.InstanceNumber)\n            except Exception:\n                k = None\n        keyed.append((k, f))\n    if any(k is None for k, _ in keyed):\n        # A series with no usable geometry keeps its arbitrary order. That is worse than\n        # sorting and better than dropping the series; the count is logged.\n        return files, False\n    return [f for _, f in sorted(keyed, key=lambda t_: t_[0])], True\n\n\ndef read_slot(rec, n_slice, out_size):\n    \"\"\"`n_slice` physically spread slices from one series, at `out_size` pixels.\n\n    Returns uint8 [n_slice, out, out], normalised per-series to its 1st-99th percentile.\n    Percentiles rather than min/max because MR intensity has no absolute scale and one\n    bright vessel would otherwise compress the whole dynamic range.\n    \"\"\"\n    files = rec.get(\"ordered\") or rec[\"files\"]\n    d, px = rec[\"dir\"], rec[\"px\"]\n    n = len(files)\n    if n == 0:\n        return None\n    # Spread the samples over the central 60% of the stack: the outer slices of a knee\n    # series are mostly soft tissue outside the joint.\n    lo, hi = int(0.20 * (n - 1)), int(0.80 * (n - 1))\n    idx = np.unique(np.linspace(lo, hi, n_slice).astype(int)) if hi > lo \\\n        else np.array([n // 2])\n    while len(idx) < n_slice:\n        idx = np.append(idx, idx[-1])\n\n    planes = []\n    for i in idx[:n_slice]:\n        try:\n            ds = pydicom.dcmread(os.path.join(d, files[int(i)]), force=True)\n            a = ds.pixel_array.astype(np.float32)\n            a = a * float(getattr(ds, \"RescaleSlope\", 1) or 1) \\\n                + float(getattr(ds, \"RescaleIntercept\", 0) or 0)\n        except Exception:\n            a = None\n        planes.append(a)\n    shp = next((p.shape for p in planes if p is not None), None)\n    if shp is None:\n        return None\n    planes = [p if (p is not None and p.shape == shp) else np.zeros(shp, np.float32)\n              for p in planes]\n    vol = np.stack(planes)\n\n    # constant physical extent, then resize\n    if px and np.isfinite(px) and px > 0:\n        want = int(round(CROP_MM / px))\n        h, w = shp\n        if 16 < want < min(h, w):\n            cy, cx, half = h // 2, w // 2, want // 2\n            vol = vol[:, max(0, cy - half):cy + half, max(0, cx - half):cx + half]\n\n    lo_v, hi_v = np.percentile(vol, [1, 99])\n    vol = np.clip((vol - lo_v) / max(hi_v - lo_v, 1e-6), 0, 1)\n    t_ = torch.from_numpy(np.ascontiguousarray(vol)).unsqueeze(0)\n    t_ = F.interpolate(t_, size=(out_size, out_size), mode=\"bilinear\",\n                       align_corners=False)\n    # uint8, not float32: intensity is already normalised into [0, 1], so eight bits cost\n    # nothing that the bilinear resize has not already cost, and the cache is a quarter\n    # the size - which is the whole reason two groups of slices fit at all.\n    return (t_.squeeze(0) * 255).round().clamp(0, 255).to(torch.uint8)\n\n\ndef normalise_laterality(img, plane, lat):\n    \"\"\"Map every knee onto a left-knee convention.\"\"\"\n    if lat != \"R\":\n        return img\n    if plane in (\"Coronal\", \"Axial\"):\n        return torch.flip(img, dims=[-1])\n    return torch.flip(img, dims=[0])"},{"cell_type":"markdown","metadata":{},"source":"## 5. Reading once, training many times\n\nThe cost of this pipeline is dominated by reading, not by arithmetic. A study holds several\nseries and each series holds tens of slices, so the corpus is hundreds of thousands of\nheader reads and, once the slices are chosen, tens of thousands of pixel decodes.\n\nThat is affordable once. It is not affordable once per epoch, and fine-tuning needs the\nsame pixels every epoch. So the slot images are decoded a single time into memory and held\nas `uint8`:\n\n$$\\text{bytes} \\;=\\; N_{\\text{study}} \\times N_{\\text{slot}} \\times S \\times P^{2}$$\n\nwith $S$ slices kept per slot at $P$ pixels. The exponent on $P$ is what makes this a real\nconstraint rather than a detail — **the cache grows with the square of resolution and only\nlinearly with slices**. Coverage is the cheap axis; resolution is the expensive one.\n\nThat asymmetry decides the one genuinely open trade in the design. A knee sagittal stack\nholds tens of slices and the cruciate is visible on a handful of them, so three slices\nspread over the central 60% of the stack is a thin sample of the joint — but a coarser\npixel grid blurs a meniscal tear, which is about a millimetre across, and a feature of\nwidth $d$ survives resampling only if the pitch is at most $d/2$. At a 130 mm crop:\n\n| grid | mm / pixel | clears a 1 mm feature? | cache at 6 slices |\n|---|---|---|---|\n| 224 | 0.580 | no | 7.4 GiB |\n| 266 | 0.489 | yes | 10.5 GiB |\n| 336 | 0.387 | comfortably | 16.7 GiB |\n\nThat table is where an earlier version of this notebook stopped, having chosen 266 px on\nthe Nyquist argument and accepted the thinner sample of the joint that came with it. Then\nthe argument was tested: two arms, same encoder, same slices, 224 px against 266 px. They\nseparated by about a **thousandth** of macro AUC — twice, on two independent runs. Nyquist\nis right about what survives resampling and wrong about what limits this model.\n\nSo the trade is settled the other way, on evidence rather than on the inequality. The cache\nis built at **224 px** and holds **four** groups of three slices — twelve per slot, four\ntimes the previous coverage, at 14.8 GiB, for the same money. §8 then spends two arms\nmeasuring whether that coverage was worth buying, which costs nothing extra because both\narms read the same cache.\n\n### The size of that table is not known in advance\n\nSubmitting to a code competition re-runs this notebook against a hidden test set, and\n`test.csv` as shipped holds three placeholder studies. A cache sized against those three\nwould be sized wrong by orders of magnitude on the run that actually scores. Worse, the\nfailure mode is not a degraded number: exhausting memory is a `SIGKILL`, which no\n`try`/`except` in this notebook can catch, and it produces no submission at all.\n\nSo the layout is planned at run time, from the study count actually mounted and from the\nmemory the kernel actually reports free — not from a constant chosen on a laptop. When the\nbudget binds, the two axes give way in a fixed order, and the order follows from §5's\nargument rather than from convenience:\n\n1. **Coverage first.** Two groups of slices drop to one. The sample of the joint gets\n   thinner, which degrades smoothly.\n2. **Resolution last, and only if a single group still will not fit.** The pixel grid has a\n   hard threshold under it, so it is the thing worth defending; it steps down in multiples\n   of 14 to keep the patch tokenisation exact.\n\nIf the grid does shrink, the arms in §8 shrink with it rather than upsampling cached pixels\nback to a size that no longer carries the detail — and two arms that collapse onto the same\n(encoder, grid) are not run twice.\n\nTraining draws one group per step, which acts as augmentation along the stack. Inference\naverages the logits over both, so a prediction does not depend on which three slices a\nsingle draw happened to pick."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"RESERVE_GB = 7.0    # torch, the CUDA context, the frames, and the per-batch host copies\n\n\ndef available_gb():\n    \"\"\"Free memory as the kernel reports it, not as the documentation promises.\"\"\"\n    try:\n        with open(\"/proc/meminfo\") as f:\n            for line in f:\n                if line.startswith(\"MemAvailable:\"):\n                    return int(line.split()[1]) / 1024 ** 2\n    except Exception:\n        pass\n    return CACHE_BUDGET_GB + RESERVE_GB\n\n\ndef plan_cache(n_study):\n    \"\"\"Choose how many slices per slot the memory budget allows.\n\n    `n_study` is the train and test corpora together because both caches are held at\n    once. Completed arms write diagnostic predictions, while the ordinary submission\n    filename remains gated until the full run passes; both paths require the test cache\n    to remain resident throughout training.\n\n    The number that matters is not the one in the config: it is what the machine\n    actually has free. Submitting a code competition re-runs this notebook against a\n    hidden test set of unknown size, and a cache sized for the three-study placeholder\n    in `test.csv` would be sized wrong there by a factor of several hundred. Getting\n    this wrong is not a degraded score - it is a SIGKILL, which no try/except in this\n    notebook can catch, and a run that scores nothing at all.\n\n    Resolution is held fixed and coverage gives way, in that order and deliberately.\n    The pixel grid has a hard threshold under it - §5's Nyquist argument - while\n    dropping from two groups of slices to one degrades the sample of the joint\n    gracefully.\n    \"\"\"\n    budget = min(CACHE_BUDGET_GB, max(3.0, available_gb() - RESERVE_GB))\n    log(f\"memory: {available_gb():.1f} GB available, reserving {RESERVE_GB:.0f} GB \"\n        f\"-> cache budget {budget:.1f} GB for {n_study} studies\")\n\n    # Coverage gives way first, resolution second, and only if a single group of three\n    # slices still does not fit. A patch grid must stay a multiple of 14 to keep the\n    # DINOv2 tokenisation exact, so the search steps in 14s.\n    img = CACHE_IMG\n    while img > 154 and n_study * N_SLOT * img * img * GROUP > budget * 1024 ** 3:\n        img -= 14\n    per_slice = n_study * N_SLOT * img * img\n    afford = int(budget * 1024 ** 3 // max(per_slice, 1))\n    groups = max(1, min(N_GROUP_MAX, afford // GROUP))\n    if groups < N_GROUP_MAX or img != CACHE_IMG:\n        log(f\"  layout reduced: {groups} group(s) of {GROUP} at {img} px \"\n            f\"({CROP_MM / img:.3f} mm/pixel)\")\n    return groups, img\n\n\nN_STUDY_TOTAL = (len(SLOTS_TR) + len(SLOTS_TE)) if STAGE_OK.get(\"headers\") else 1\nif STAGE_OK.get(\"headers\"):\n    N_GROUP, CACHE_IMG = plan_cache(N_STUDY_TOTAL)\nelse:\n    N_GROUP = 1\nCACHE_SLICES = GROUP * N_GROUP\nlog(f\"cache layout: {N_GROUP} group(s) x {GROUP} slices = {CACHE_SLICES} per slot \"\n    f\"at {CACHE_IMG} px  ->  \"\n    f\"{N_STUDY_TOTAL * N_SLOT * CACHE_SLICES * CACHE_IMG ** 2 / 1024 ** 3:.1f} GB\")\n\n\ndef build_cache(slot_map, lat_map, tag, limit=None):\n    \"\"\"Decode every (study, slot) once into one uint8 array.\"\"\"\n    studies = sorted(slot_map)\n    if limit:\n        studies = studies[:limit]\n    sidx = {s: i for i, s in enumerate(studies)}\n    cache = np.zeros((len(studies), N_SLOT, CACHE_SLICES, CACHE_IMG, CACHE_IMG), np.uint8)\n    mask = np.zeros((len(studies), N_SLOT), np.float32)\n    log(f\"{tag}: cache {cache.shape} = {cache.nbytes / 1024 ** 3:.2f} GB\")\n\n    jobs = [(st, k, plane, slot_map[st][name])\n            for st in studies\n            for k, (name, plane, _, _) in enumerate(SLOTS)\n            if name in slot_map[st]]\n\n    # Ordering first, and as its own pass. It reads one header per slice of every chosen\n    # series - far more file opens than the decode that follows - and on a network mount\n    # that is latency rather than work, so it gets its own wider pool and its own budget.\n    t_ord = time.time()\n    log(f\"{tag}: ordering {len(jobs)} slot-series \"\n        f\"({sum(len(j[3]['files']) for j in jobs)} slice headers)\")\n    ok = done = 0\n    with ThreadPoolExecutor(max_workers=ORDER_THREADS) as pool:\n        for c0 in range(0, len(jobs), 1024):\n            block = jobs[c0:c0 + 1024]\n            for (_, _, _, rec), (files, good) in zip(\n                    block, pool.map(lambda j: order_slices(j[3]), block)):\n                rec[\"ordered\"] = files\n                ok += int(good)\n                done += 1\n            if time.time() - t_ord > ORDER_BUDGET_S:\n                log(f\"{tag}: ordering budget spent at {done}/{len(jobs)}; \"\n                    f\"the rest keep file order\")\n                break\n    log(f\"{tag}: ordered {ok}/{len(jobs)} by geometry in {time.time() - t_ord:.0f}s\")\n\n    log(f\"{tag}: decoding {len(jobs)} slot-series\")\n    done = 0\n    with ThreadPoolExecutor(max_workers=PIX_THREADS) as pool:\n        for c0 in range(0, len(jobs), 512):\n            block = jobs[c0:c0 + 512]\n            for (st, k, plane, _), img in zip(\n                    block, pool.map(lambda j: read_slot(j[3], CACHE_SLICES, CACHE_IMG),\n                                    block)):\n                done += 1\n                if img is None:\n                    continue\n                cache[sidx[st], k] = normalise_laterality(\n                    img, plane, lat_map.get(st)).numpy()\n                mask[sidx[st], k] = 1.0\n            if done % 4096 < 512:\n                log(f\"  {tag} {done}/{len(jobs)}\")\n            if time.time() - T0 > TIME_BUDGET:\n                raise TimeoutError(f\"{tag}: time budget reached during decode\")\n    gc.collect()\n    return studies, cache, mask\n\n\n@stage(\"cache\")\ndef _cache():\n    global ST_TE, CTE, MTE, ST_TR, CTR, MTR\n    lim = 240 if SMOKE else None\n    ST_TE, CTE, MTE = build_cache(SLOTS_TE, LAT_TE, \"test\")\n    ST_TR, CTR, MTR = build_cache(SLOTS_TR, LAT_TR, \"train\", limit=lim)\n    log(f\"cached {len(ST_TR)} train studies, {len(ST_TE)} test studies; \"\n        f\"slot presence {MTR.mean() * 100:.1f}%\")\n    log(f\"cache resident: {(CTR.nbytes + CTE.nbytes) / 1024 ** 3:.2f} GB; \"\n        f\"{available_gb():.1f} GB still free\")\n\n\nif STAGE_OK.get(\"headers\"):\n    _cache()"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"if STAGE_OK.get(\"cache\") and len(ST_TR):\n    # A study with as many slots as possible, so the figure shows what the head is given.\n    j = int(np.argmax(MTR.sum(1)))\n    names = [s[0] for s in SLOTS]\n    fig, axes = plt.subplots(N_GROUP, N_SLOT, figsize=(1.55 * N_SLOT, 1.7 * N_GROUP),\n                             squeeze=False)\n    for g in range(N_GROUP):\n        for k in range(N_SLOT):\n            ax_ = axes[g][k]\n            ax_.set_xticks([])\n            ax_.set_yticks([])\n            ax_.grid(False)\n            if MTR[j, k] < .5:\n                ax_.set_facecolor(\"#eceff2\")\n                ax_.text(.5, .5, \"absent\", ha=\"center\", va=\"center\", fontsize=6.5,\n                         color=\"#8a97a3\", transform=ax_.transAxes)\n            else:\n                # the three cached slices of this group, shown as an RGB composite -\n                # which is exactly what the encoder receives\n                rgb = CTR[j, k, g * GROUP:(g + 1) * GROUP].transpose(1, 2, 0)\n                ax_.imshow(rgb)\n            if g == 0:\n                ax_.set_title(names[k], fontsize=6.5, pad=3)\n            if k == 0:\n                ax_.set_ylabel(f\"group {g}\", fontsize=6.5)\n    fig.suptitle(\"one study as the model sees it: six protocol slots x \"\n                 f\"{N_GROUP} groups of {GROUP} neighbouring slices, \"\n                 \"normalised to a left knee and a 130 mm field\",\n                 fontsize=8, y=1.02)\n    fig.tight_layout()\n    plt.show()\n    print(f\"cache {CTR.nbytes / 1024 ** 3:.2f} GB train + \"\n          f\"{CTE.nbytes / 1024 ** 3:.2f} GB test\")"},{"cell_type":"markdown","metadata":{},"source":"## 6. Six views into twelve decisions\n\nA study arrives as up to six slot embeddings $x_s \\in \\mathbb{R}^{d}$ with a presence mask\n$m_s \\in \\{0,1\\}$. Pooling them identically would discard the reason the protocol has three\nplanes at all: each finding is read on particular sequences, and a mean over slots dilutes\nthe one that carries the evidence with five that do not.\n\nSo each diagnosis $o$ gets its own query $q_o$ and attends over the slots it needs:\n\n$$h_s = \\mathrm{GELU}(W\\,\\mathrm{LN}(x_s)) + e_s,\\qquad\n\\alpha_{os} = \\frac{\\exp\\!\\big(q_o^{\\top}h_s/\\sqrt{d_h}\\big)\\,m_s}\n{\\sum_{s'} \\exp\\!\\big(q_o^{\\top}h_{s'}/\\sqrt{d_h}\\big)\\,m_{s'}},\\qquad\nz_o = w_o^{\\top}\\!\\!\\sum_s \\alpha_{os} h_s + b_o$$\n\nwhere $e_s$ is a learned slot identity, so the query knows a coronal fat-suppressed series\nfrom a sagittal T1. Missing slots are masked out of the softmax rather than zero-filled —\nattention over a zero vector is still attention, and it would let an absent sequence dilute\na present one.\n\nThe aggregation is deliberately no deeper than this. The label is attached to the *study*,\nso nothing in the supervision says which slice or which region within a slot matters;\nattention parameters below the slot level would have nothing to learn from and would spend\ntheir capacity fitting noise.\n\n### Why the encoder is trained rather than frozen\n\nA frozen self-supervised encoder is the cheap option and it is bounded by something no\namount of work downstream can reach. Resolution, encoder size, slice coverage and slot\naggregation all change how much the model looks, how closely, and how it summarises what it\nsaw — but none of them changes the vocabulary it looks *with*. Every one of those axes runs\ninto the same ceiling, and the ceiling is the representation.\n\nThere is a concrete reason to expect that ceiling to bind here rather than sit harmlessly\nhigh: DINOv2 learned its features from natural images, where nothing resembles the signal a\ntorn meniscus makes on a proton-density sequence.\n\nSo the encoder is adapted, under two restraints.\n\n**Only the last blocks move.** The early blocks of a vision transformer are generic edge and\ntexture filters; the late blocks are where semantics live. There may not be enough\nsupervision here to improve the early ones, and there is certainly enough to damage them.\n\n**The encoder learns two orders of magnitude more slowly than the head.** The head is random\nat initialisation and has everything to learn; the encoder starts from a good solution and\nneeds only to be moved off it. A single learning rate would either leave the head untrained\nor destroy the encoder in the first few hundred steps."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"class SlotHead(nn.Module):\n    \"\"\"Per-diagnosis attention over the slot embeddings of one study.\"\"\"\n\n    def __init__(self, dim, n_slot, n_out, hidden=256, p=0.2):\n        super().__init__()\n        self.proj = nn.Sequential(nn.LayerNorm(dim), nn.Linear(dim, hidden), nn.GELU())\n        self.slot_emb = nn.Parameter(torch.randn(n_slot, hidden) * 0.02)\n        self.query = nn.Parameter(torch.randn(n_out, hidden) * 0.02)\n        self.drop = nn.Dropout(p)\n        self.out = nn.Linear(hidden, n_out)\n        self.hidden = hidden\n\n    def forward(self, x, mask, return_att=False):\n        h = self.proj(x) + self.slot_emb\n        att = torch.einsum(\"bsh,oh->bos\", h, self.query) / self.hidden ** 0.5\n        att = att.masked_fill(mask.unsqueeze(1) < 0.5, -1e4).softmax(-1)\n        ctx = self.drop(torch.einsum(\"bos,bsh->boh\", att, h))\n        z = (ctx * self.out.weight.unsqueeze(0)).sum(-1) + self.out.bias\n        return (z, att) if return_att else z\n\n\nclass Model(nn.Module):\n    \"\"\"Encoder plus head, trained end to end.\n\n    The bag is flattened for the encoder and folded back before the head, so the encoder\n    never sees the study structure and the head never sees pixels.\n    \"\"\"\n\n    def __init__(self, backbone, dim):\n        super().__init__()\n        self.backbone = backbone\n        self.head = SlotHead(dim, N_SLOT, len(TARGETS))\n        self.register_buffer(\"mean\", torch.tensor([0.485, 0.456, 0.406]).view(1, 3, 1, 1))\n        self.register_buffer(\"std\", torch.tensor([0.229, 0.224, 0.225]).view(1, 3, 1, 1))\n\n    def forward(self, imgs, mask, img_size=None, return_att=False):\n        B, S = imgs.shape[:2]\n        x = imgs.reshape(B * S, *imgs.shape[2:]).float().div_(255.0)\n        if img_size is not None and img_size != x.shape[-1]:\n            # The cache is held at the highest resolution any arm needs; the rest\n            # downsample from it, so every arm sees the same pixels through a different\n            # sampling grid rather than a different crop.\n            x = F.interpolate(x, size=(img_size, img_size), mode=\"bilinear\",\n                              align_corners=False)\n        x = (x - self.mean) / self.std\n        out = self.backbone(pixel_values=x).last_hidden_state\n        feat = torch.cat([out[:, 0], out[:, 1:].mean(1)], dim=1).reshape(B, S, -1)\n        return self.head(feat, mask, return_att)\n\n\ndef find_dinov2(variant):\n    \"\"\"Locate a mounted DINOv2 checkpoint directory by variant name.\"\"\"\n    base = Path(\"/kaggle/input\")\n    if not base.is_dir():\n        return None\n    hits = []\n    for root, dirs, files in os.walk(base):\n        dirs[:] = [d for d in dirs if d not in (\"train_series\", \"test_series\")]\n        if \"config.json\" in files and \"dinov2\" in root.lower():\n            hits.append(Path(root))\n    exact = [h for h in hits if f\"/{variant}/\" in str(h).lower()\n             or str(h).lower().rstrip(\"/\").endswith(variant)]\n    if exact:\n        return exact[0]\n    # Deliberately no fallback to \"some other DINOv2\". A loose match here would load\n    # whichever checkpoint happened to be mounted, so an arm asking for `large` would\n    # quietly train a second copy of `base` - and then be rank-averaged with the first,\n    # doubling one model's vote under a name that says otherwise. Under the safety\n    # contract, a missing exact encoder raises and the caller re-raises, aborting the run.\n    return None\n\n\ndef build_model(variant, unfreeze_last=UNFREEZE_LAST):\n    from transformers import AutoModel\n    p = find_dinov2(variant)\n    if p is None:\n        raise FileNotFoundError(f\"DINOv2 '{variant}' weights not attached\")\n    bb = AutoModel.from_pretrained(str(p))\n    n_layer = len(bb.encoder.layer)\n    for prm in bb.parameters():\n        prm.requires_grad = False\n    for blk in bb.encoder.layer[max(0, n_layer - unfreeze_last):]:\n        for prm in blk.parameters():\n            prm.requires_grad = True\n    for prm in bb.layernorm.parameters():\n        prm.requires_grad = True\n    trainable = sum(q.numel() for q in bb.parameters() if q.requires_grad)\n    log(f\"backbone {variant} from {p.name}: {n_layer} blocks, last {unfreeze_last} \"\n        f\"trainable ({trainable / 1e6:.1f}M params), hidden {bb.config.hidden_size}\")\n    return Model(bb, bb.config.hidden_size * 2)\n\n\ndef take_group(rows, g):\n    return rows[:, :, g * GROUP:(g + 1) * GROUP]\n\n\ndef augment(imgs):\n    \"\"\"Vertical flip and a small intensity scale, applied to a whole bag at once.\n\n    No horizontal flip: laterality was normalised onto a single convention upstream, and\n    flipping horizontally would reintroduce exactly the nuisance axis that removed.\n    \"\"\"\n    if torch.rand(1).item() < 0.5:\n        imgs = torch.flip(imgs, dims=[-2])\n    scale = 1.0 + (torch.rand(1, device=imgs.device) - 0.5) * 0.2\n    return (imgs.float() * scale).clamp(0, 255).to(imgs.dtype)\n\n\n@torch.no_grad()\ndef predict(model, cache, mask, idx, img_size, n_group=None, both=False):\n    \"\"\"Average the logits over the slice groups this arm is allowed to see.\n\n    Training draws one group per step, which is augmentation along the stack; inference\n    averages over all of an arm's groups, so a prediction does not depend on which three\n    slices a single draw happened to pick. `n_group` is per-arm so that two arms reading\n    the same cache can be given different amounts of the joint.\n\n    With `both`, the same forward passes are also aggregated by max, at no extra cost.\n    Mean and max are not interchangeable here and neither is right for all twelve\n    targets: a meniscal tear or a fracture is focal, visible on the groups that cross it\n    and absent from the rest, so averaging four groups dilutes the one that saw it by\n    four; an effusion or tricompartmental cartilage loss is diffuse and every group sees\n    it, where the mean is the better estimator and the max is just the noisiest group.\n    So the choice is made per target, on the holdout, in §8.\n    \"\"\"\n    ng = N_GROUP if n_group is None else min(n_group, N_GROUP)\n    model.eval()\n    mean_out, max_out = [], []\n    for b in range(0, len(idx), EVAL_BATCH):\n        sel = idx[b:b + EVAL_BATCH]\n        rows = torch.from_numpy(cache[sel]).to(DEV)\n        m = torch.from_numpy(mask[sel]).to(DEV)\n        zs = []\n        for g in range(ng):\n            with torch.autocast(\"cuda\", enabled=DEV.type == \"cuda\"):\n                zs.append(model(take_group(rows, g), m, img_size).float())\n        z = torch.stack(zs)                      # [groups, batch, targets]\n        mean_out.append(torch.sigmoid(z.mean(0)).cpu().numpy())\n        max_out.append(torch.sigmoid(z.max(0).values).cpu().numpy())\n    if not mean_out:\n        z0 = np.zeros((0, len(TARGETS)), np.float32)\n        return z0 if not both else (z0, z0)\n    mean_p = np.concatenate(mean_out)\n    if not both:\n        return mean_p\n    return mean_p, np.concatenate(max_out)\n\n\ndef macro_auc(y, p):\n    v = [roc_auc_score(y[:, j], p[:, j]) if len(set(y[:, j])) > 1 else np.nan\n         for j in range(y.shape[1])]\n    return float(np.nanmean(v)), np.array(v, dtype=float)"},{"cell_type":"markdown","metadata":{},"source":"## 7. Validating without fooling yourself\n\nTwo leaks are specific to this setup, and both inflate a holdout number without improving\nanything that will be scored.\n\n**Identical reports across studies.** Some reports in this corpus are byte-identical between\nstudies — short template reports, and at least one study filed with two reports. A\nreport-derived target vector is a deterministic function of its text, so splitting such a\ngroup across the train/holdout boundary scores the model against a target whose *source* it\nhas already fitted. The split is therefore grouped on a hash of the report text, not on the\nstudy identifier.\n\n**The annotated studies.** The 58 annotations are the only labels in this dataset read from\nthe images rather than from prose, so they are the most trustworthy check available — and\nthey are also, at triple weight, the most valuable training rows. Keeping them for training\nand then evaluating on all of them would be scoring the model against examples it has seen\nwith the answer attached. So they stay in training, and the annotation check reports only\nthe handful that happen to fall in the holdout. That number is quoted with its count\nattached, because a dozen studies cannot arbitrate between epochs.\n\n**What actually selects.** The holdout is the report-derived targets, binarised at their\nmidpoint. It is a proxy for a proxy, and it is what chooses the epoch and the arm — there is\nnothing better available at this sample size. The annotation check is reported alongside\nbecause it measures something genuinely different, not because it can decide anything.\n\nThe loss is a weighted binary cross-entropy over the twelve outputs,\n\n$$\\mathcal{L} = \\frac{1}{12B}\\sum_{b}\\sum_{o} w_{bo}\\;\n\\mathrm{BCE}\\big(z_{bo},\\,y_{bo}\\big),$$\n\nwith $y$ the graded report target in $(0,1)$ — not a hard label — and $w$ the per-finding\nconfidence from §2, rescaled to $[0.25, 1]$, with the annotated rows at 3. A study whose\nreport never mentions synovitis contributes to the synovitis head at a quarter of the\nweight of one that names it."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import hashlib\n\n\n@stage(\"targets\")\ndef _targets():\n    global Y, W, TR_IDX, VA_IDX, YV, GI, GOLD_Y\n    Y = np.zeros((len(ST_TR), len(TARGETS)), np.float32)\n    W = np.zeros_like(Y)\n    conf_cols = [t_ + \"__conf\" for t_ in TARGETS]\n    for i, st in enumerate(ST_TR):\n        if st in GOLD.index:\n            Y[i] = GOLD.loc[st].values\n            W[i] = 3.0\n        elif st in LAB.index:\n            r = LAB.loc[st]\n            Y[i] = r[TARGETS].values\n            W[i] = 0.25 + 0.75 * r[conf_cols].values\n    keep = np.where(W.sum(1) > 0)[0]\n\n    rep = train_df.set_index(\"StudyInstanceUID\")[\"Report\"].fillna(\"\")\n    grp = np.array([int(hashlib.md5(rep.get(s, s).encode()).hexdigest()[:8], 16) % 5\n                    for s in ST_TR])\n    VA_IDX = np.array([i for i in keep if grp[i] == 0])\n    TR_IDX = np.array([i for i in keep if grp[i] != 0])\n    if len(VA_IDX) == 0 or len(TR_IDX) < BATCH_STUDIES:\n        cut = max(1, len(keep) // 5)\n        VA_IDX, TR_IDX = keep[:cut], keep[cut:]\n    YV = (Y[VA_IDX] > 0.5).astype(int)\n\n    pos = {s: i for i, s in enumerate(ST_TR)}\n    va_set = set(VA_IDX.tolist())\n    GI = np.array([pos[s] for s in GOLD.index if s in pos and pos[s] in va_set],\n                  dtype=int)\n    GOLD_Y = (GOLD.loc[[ST_TR[i] for i in GI]].values.astype(int)\n              if len(GI) else None)\n\n    log(f\"supervised {len(keep)} of {len(ST_TR)} studies \"\n        f\"(annotated {sum(s in GOLD.index for s in ST_TR)})\")\n    log(f\"train {len(TR_IDX)} / holdout {len(VA_IDX)} studies, grouped on report text\")\n    log(f\"annotation check: {len(GI)} annotated studies fell in the holdout\")\n    log(f\"holdout positive rate per target: \"\n        f\"{np.round(YV.mean(0), 2).tolist()}\")\n\n\nif STAGE_OK.get(\"cache\"):\n    _targets()"},{"cell_type":"markdown","metadata":{},"source":"## 8. Five arms, two controlled comparisons, and how they are combined\n\n§1 established that the metric reads order and nothing else. That decides how to combine\nmodels: averaging probabilities lets the more confident arm dominate whatever its ranking\nis worth, while averaging **per-column ranks** combines exactly what the score reads. Every\nfile written below is converted to percentile ranks first.\n\nFive arms, all reading the same cache:\n\n| arm | encoder | slices / slot | what it varies |\n|---|---|---|---|\n| `s168g4` | DINOv2-S/14 | 12 | the **resolution safety check** — see below |\n| `s168g8` | DINOv2-S/14 | 24 | **coverage** |\n| `b168g8` | DINOv2-B/14 | 24 | encoder capacity |\n| `b168g8b` | DINOv2-B/14 | 24 | seed only |\n| `b168g8c` | DINOv2-B/14 | 24 | seed only |\n\n**The safety check matters more than the experiment.** The previous version measured a\ncoverage curve at 224 px — 3, 6 and 12 slices scoring 0.7536, 0.7726, 0.7866 — which was\nstill climbing when it ran out of memory. Buying more slices means paying in pixels, and\nthe evidence that pixels are free stops at 224: resolution was tested at 224 against 266\nand made no difference, but nothing has ever been measured *below* 224, and somewhere\nthere is a floor.\n\nSo `s168g4` holds slices at 12 and only drops the grid. It is directly comparable with the\nprevious version's `s224g4` (0.7866, same encoder, same slices, same epochs, same seed).\nIf it lands there, the coarser grid is free and the extra slices are pure gain. If it\ndrops, the floor has been found — and that is worth knowing and reporting either way,\nbecause it bounds how far this trade can be pushed by anyone who continues it.\n\n### Why the coverage comparison is free\n\n`s168g4` and `s168g8` differ in nothing but how many of the cached slice groups each may\ndraw from during training and average over at inference — same encoder, same grid, same\ncache, same initialisation, and **the same cost per epoch**, because training draws one\ngroup per step either way. Only the evaluation passes cost more. That is what makes it\naffordable to keep measuring this rather than assuming it, and it is why the notebook has\na coverage curve at all instead of a preference.\n\nA DINOv2-**large** arm was also tried and is not here: it scored below `base` on the\nholdout while costing two and a half hours, so capacity is not monotonic on this corpus.\nOnce coverage was corrected, capacity stopped separating the arms at all — `small` and\n`base` landed within a thousandth of each other at twelve slices. Reporting that is more\nuseful than quietly dropping it.\n\n### Mean or max over the slice groups, decided per finding\n\nAn arm with eight groups produces eight logits per target and has to reduce them to one.\nAveraging is the obvious choice and it is wrong for half the targets, for a reason that is\nanatomical rather than statistical.\n\nA meniscal tear, a fracture, a bone contusion are **focal**: they occupy a few slices, so\nthe groups that cross them see them and the rest see nothing. Averaging eight groups then\ndivides the evidence by eight, and does so most severely for exactly the small, rare\nfindings §1 says are most expensive to lose. An effusion, tricompartmental cartilage loss,\na large Baker's cyst are **diffuse**: every group sees them, the eight logits are estimates\nof the same quantity, and the mean is the better estimator while the max is just whichever\ngroup was noisiest.\n\nSo both reductions are computed — from the same forward passes, at no extra cost — and the\nchoice is made per target on the holdout. Twelve binary decisions fitted on roughly nine\nhundred studies is a real but small amount of freedom, and the arms print which targets\nchose the max, so the pattern can be checked against the anatomy rather than taken on\ntrust. If it does not look like the focal/diffuse split above, it is fitting noise.\n\nWith a single group mean and max coincide and the choice is a no-op, which is what made\nthe 3-slice control in the previous version clean.\n\n### Which arms actually get averaged\n\nAn equal-weight mean over every arm assumes every arm deserves an equal vote — an\nassumption this notebook tested and watched fail. So the arms that enter the submission\nare chosen by greedy forward selection *with replacement* on the holdout: at each step,\nadd whichever arm most improves the macro AUC of the running rank mean, stop when nothing\nimproves it by more than $10^{-4}$. Selection with replacement is what turns this from a\nfilter into a weighting — an arm picked twice is an arm weighted twice — and a weak arm is\nsimply never picked.\n\nThat selection reads the same report-derived holdout that chose the epochs, so it can\noverfit it. With five candidates and roughly nine hundred studies the risk is small but\nreal, which is why the equal-weight mean over all arms is written alongside as\n`submission_equalweight.csv` and both numbers are printed. If they disagree by more than\nthe holdout can resolve, believe neither.\n\nArms run cheapest-first and write diagnostic per-arm and partial-ensemble files. This\nsafety derivative exposes the ordinary `submission.csv` name only after every expected\narm and stage completes and the final matrix passes validation; a time-budget expiry is\ntherefore a hard failure rather than a partially successful submission."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"def write_submission(pred, studies, path):\n    \"\"\"One submission file from one prediction matrix, as per-column ranks.\"\"\"\n    sub = pd.DataFrame(pd.DataFrame(pred).rank(pct=True).values, columns=TARGETS)\n    sub.insert(0, \"StudyInstanceUID\", studies)\n    sub = TEST_DF[[\"StudyInstanceUID\"]].merge(sub, on=\"StudyInstanceUID\", how=\"left\")\n    sub[TARGETS] = sub[TARGETS].fillna(0.5)\n    sub.to_csv(path, index=False)\n    return sub\n\n\ndef rank_mean(preds):\n    return np.mean([pd.DataFrame(p).rank(pct=True).values for p in preds], axis=0)\n\n\ndef train_arm(cfg):\n    \"\"\"Fine-tune one configuration; return its holdout history and test predictions.\"\"\"\n    # An arm can ask for no more resolution than the cache was actually built at; if\n    # §5's planner had to shrink the grid, the arms shrink with it rather than\n    # upsampling cached pixels back to a size that no longer carries the detail.\n    cfg = dict(cfg, img=min(cfg[\"img\"], CACHE_IMG))\n    batch = int(cfg.get(\"batch\", BATCH_STUDIES))\n    ng = min(int(cfg.get(\"groups\", N_GROUP)), N_GROUP)\n    log(f\"=== {cfg['name']}: {cfg['variant']}, batch {batch}, {cfg['img']} px, \"\n        f\"{CROP_MM / cfg['img']:.3f} mm/pixel, \"\n        f\"{CROP_MM / cfg['img'] * 14:.2f} mm per patch token, \"\n        f\"{ng * GROUP} slices/slot ===\")\n    torch.manual_seed(int(cfg.get(\"seed\", SEED)))\n    np.random.seed(int(cfg.get(\"seed\", SEED)))\n    model = build_model(cfg[\"variant\"]).to(DEV)\n    opt = torch.optim.AdamW([\n        {\"params\": [p for p in model.backbone.parameters() if p.requires_grad],\n         \"lr\": cfg[\"lr_bb\"]},\n        {\"params\": model.head.parameters(), \"lr\": LR_HEAD},\n    ], weight_decay=WEIGHT_DECAY)\n    steps = max(EPOCHS * (len(TR_IDX) // batch), 1)\n    sched = torch.optim.lr_scheduler.OneCycleLR(\n        opt, max_lr=[cfg[\"lr_bb\"], LR_HEAD], total_steps=steps, pct_start=0.15)\n    scaler = torch.amp.GradScaler(\"cuda\", enabled=DEV.type == \"cuda\")\n\n    hist, best, best_state, best_per, best_val = [], -1.0, None, None, None\n    for ep in range(EPOCHS):\n        model.train()\n        perm = np.random.permutation(TR_IDX)\n        tot = nstep = 0\n        for b in range(0, len(perm) - batch + 1, batch):\n            sel = perm[b:b + batch]\n            rows = torch.from_numpy(CTR[sel]).to(DEV)\n            g = int(torch.randint(ng, (1,)).item())\n            imgs = augment(take_group(rows, g))\n            m = torch.from_numpy(MTR[sel]).to(DEV)\n            y = torch.from_numpy(Y[sel]).to(DEV)\n            w = torch.from_numpy(W[sel]).to(DEV)\n            with torch.autocast(\"cuda\", enabled=DEV.type == \"cuda\"):\n                loss = (F.binary_cross_entropy_with_logits(\n                    model(imgs, m, cfg[\"img\"]), y, reduction=\"none\") * w).mean()\n            opt.zero_grad(set_to_none=True)\n            scaler.scale(loss).backward()\n            scaler.step(opt)\n            scaler.update()\n            sched.step()\n            tot += float(loss.item())\n            nstep += 1\n            if time.time() - T0 > TIME_BUDGET:\n                raise TimeoutError(\"time budget reached inside training batch\")\n\n        pv = predict(model, CTR, MTR, VA_IDX, cfg[\"img\"], ng)\n        d, per = macro_auc(YV, pv)\n        g_auc = float(\"nan\")\n        if GOLD_Y is not None and len(GI) > 3:\n            g_auc = macro_auc(GOLD_Y,\n                              predict(model, CTR, MTR, GI, cfg[\"img\"], ng))[0]\n        hist.append({\"epoch\": ep + 1, \"loss\": tot / max(nstep, 1), \"holdout\": d,\n                     \"annot\": g_auc})\n        log(f\"  epoch {ep + 1}/{EPOCHS}  loss {tot / max(nstep, 1):.4f}  \"\n            f\"holdout {d:.4f}  annot(n={len(GI)}) {g_auc:.4f}  \"\n            f\"[{(time.time() - T0) / 3600:.2f} h]\")\n\n        # Selection reads the holdout alone; see §7 for why the annotation check cannot\n        # arbitrate between epochs.\n        if d > best:\n            best, best_per, best_val = d, per, pv\n            best_state = {k: v.detach().cpu().clone()\n                          for k, v in model.state_dict().items()}\n        if time.time() - T0 > TIME_BUDGET:\n            raise TimeoutError(\"time budget reached after epoch evaluation\")\n\n    if best_state is not None:\n        model.load_state_dict(best_state)\n\n    # Choose mean- or max-over-groups per target, on the holdout. Free: both come out of\n    # the same forward passes. With one group the two are identical and the choice is a\n    # no-op, which is what makes `s224g1` a clean control.\n    vm, vx = predict(model, CTR, MTR, VA_IDX, cfg[\"img\"], ng, both=True)\n    _, per_mean = macro_auc(YV, vm)\n    _, per_max = macro_auc(YV, vx)\n    use_max = np.nan_to_num(per_max) > np.nan_to_num(per_mean)\n    best_val = np.where(use_max, vx, vm)\n    gain = float(np.nanmean(np.where(use_max, per_max, per_mean)) -\n                 np.nanmean(per_mean))\n    log(f\"  group aggregation: max chosen for \"\n        f\"{[t for t, u in zip(TARGETS, use_max) if u] or 'no target'}\"\n        f\"  (+{gain:.4f} holdout)\")\n\n    tm, tx = predict(model, CTE, MTE, np.arange(len(ST_TE)), cfg[\"img\"], ng, both=True)\n    tp = np.where(use_max, tx, tm)\n    best_per = np.where(use_max, per_max, per_mean)\n    best = float(np.nanmean(best_per))\n    log(f\"  {cfg['name']}: holdout {best:.4f}\")\n    del model, opt, sched, scaler, best_state\n    gc.collect()\n    if DEV.type == \"cuda\":\n        torch.cuda.empty_cache()\n    return {\"best\": best, \"per_target\": best_per, \"hist\": hist, \"test\": tp,\n            \"val\": best_val, \"slices\": ng * GROUP,\n            \"use_max\": [t for t, u in zip(TARGETS, use_max) if u]}\n\n\nRESULTS, TEST_PREDS = {}, {}\n\n\n@stage(\"train\")\ndef _train():\n    seen = set()\n    for cfg in ARMS:\n        left = TIME_BUDGET - (time.time() - T0)\n        if left < 900:\n            raise TimeoutError(f\"{cfg['name']}: only {left / 60:.0f} min left\")\n        # If the planner shrank the grid or the slice budget, two arms can collapse onto\n        # the same configuration. Running the second one would spend an hour reproducing\n        # the first one's predictions and then average them with itself.\n        #\n        # The key has to name every axis an arm can vary, not just the grid. An earlier\n        # version keyed on (encoder, resolution) alone, which silently swallowed the\n        # entire coverage experiment - `s224g1` and `s224g4` differ in neither - and the\n        # second seed with it. A deduplication rule that is coarser than the experiment\n        # deletes the experiment.\n        key = (cfg[\"variant\"], min(cfg[\"img\"], CACHE_IMG),\n               min(int(cfg.get(\"groups\", N_GROUP)), N_GROUP),\n               int(cfg.get(\"seed\", SEED)), int(cfg.get(\"batch\", BATCH_STUDIES)))\n        if key in seen:\n            raise RuntimeError(f\"{cfg['name']}: collapsed onto prior arm {key}\")\n        seen.add(key)\n        try:\n            r = train_arm(cfg)\n        except Exception:\n            import traceback\n            traceback.print_exc()\n            log(f\"arm {cfg['name']} failed; aborting safety candidate\")\n            gc.collect()\n            if DEV.type == \"cuda\":\n                torch.cuda.empty_cache()\n            raise\n        RESULTS[cfg[\"name\"]] = r\n        TEST_PREDS[cfg[\"name\"]] = r[\"test\"]\n        write_submission(r[\"test\"], ST_TE, f\"submission_{cfg['name']}.csv\")\n        # Refresh a diagnostic partial ensemble after each completed arm. It is never\n        # promoted to the ordinary submission filename by an incomplete run.\n        ens = rank_mean(list(TEST_PREDS.values()))\n        write_submission(ens, ST_TE, \"submission_partial_rankmean.csv\")\n        log(f\"  diagnostic partial rank mean now has {len(TEST_PREDS)} arm(s)\")\n    if not TEST_PREDS:\n        raise RuntimeError(\"no arm completed\")\n\n\nSELECTED = []\n\n\n@stage(\"ensemble\")\ndef _ensemble():\n    \"\"\"Choose which arms to average, on the holdout, instead of averaging all of them.\n\n    An equal-weight mean over every arm assumes every arm deserves an equal vote. That\n    assumption was tested and failed: a DINOv2-large arm scored well below base here, and\n    averaging it in at equal weight drags the mean toward a model already known to be\n    worse. Greedy forward selection with replacement (Caruana et al.) fixes both halves\n    of the problem at once - a bad arm is simply never picked, and a good arm picked\n    twice is a good arm weighted twice.\n\n    Selection reads the same report-derived holdout that chose the epochs, so it can\n    overfit it; with four candidates and ~900 studies that risk is small, but it is real,\n    which is why a pick has to improve the holdout by more than EPS to be accepted, and\n    why the unselected equal-weight mean is written alongside for comparison.\n    \"\"\"\n    global SELECTED\n    names = [n for n in TEST_PREDS if RESULTS[n].get(\"val\") is not None]\n    if not names:\n        raise RuntimeError(\"no validation-backed arm available for ensemble\")\n    EPS = 1e-4\n    chosen, best = [], -1.0\n    for _ in range(2 * len(names)):\n        pick, pick_s = None, best + EPS\n        for n in names:\n            s, _ = macro_auc(YV, rank_mean([RESULTS[k][\"val\"] for k in chosen + [n]]))\n            if s > pick_s:\n                pick, pick_s = n, s\n        if pick is None:\n            break\n        chosen.append(pick)\n        best = pick_s\n    if not chosen:\n        chosen = [max(names, key=lambda n: RESULTS[n][\"best\"])]\n        best, _ = macro_auc(YV, RESULTS[chosen[0]][\"val\"])\n    SELECTED = chosen\n\n    flat, _ = macro_auc(YV, rank_mean([RESULTS[k][\"val\"] for k in names]))\n    from collections import Counter\n    log(f\"ensemble: {dict(Counter(chosen))}\")\n    log(f\"  holdout macro AUC  selected {best:.4f}   equal-weight-all {flat:.4f}   \"\n        f\"best single {max(RESULTS[n]['best'] for n in names):.4f}\")\n    write_submission(rank_mean([RESULTS[k][\"test\"] for k in names]), ST_TE,\n                     \"submission_equalweight.csv\")\n    sub = write_submission(rank_mean([RESULTS[k][\"test\"] for k in chosen]), ST_TE,\n                           \"_submission_candidate.csv\")\n    log(f\"  validated candidate matrix = rank mean of {len(chosen)} picks over \"\n        f\"{len(set(chosen))} distinct arm(s); {sub.shape}\")\n\n\nif STAGE_OK.get(\"targets\"):\n    _train()\n\nif STAGE_OK.get(\"train\"):\n    _ensemble()"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"if RESULTS:\n    fig, ax = plt.subplots(1, 3, figsize=(12, 3.4),\n                           gridspec_kw={\"width_ratios\": [1, 1.35, .9]})\n    cols = {\"s168g4\": \"#8a97a3\", \"s168g8\": ACC, \"b168g8\": WARN,\n            \"b168g8b\": \"#6b8f3a\", \"b168g8c\": \"#9a6fb0\"}\n\n    for name, r in RESULTS.items():\n        h = pd.DataFrame(r[\"hist\"])\n        c = cols.get(name, INK)\n        ax[0].plot(h.epoch, h.holdout, \"-o\", ms=3, color=c, label=f\"{name} holdout\")\n        ax[0].plot(h.epoch, h.annot, \":\", lw=1, color=c, alpha=.6)\n    ax[0].set_xlabel(\"epoch\")\n    ax[0].set_ylabel(\"macro AUC\")\n    ax[0].legend(fontsize=6.5, frameon=False)\n    ax[0].set_title(\"solid: report-derived holdout (selects)\\n\"\n                    \"dotted: annotation check (reports only)\", loc=\"left\", fontsize=8)\n\n    w = 0.8 / max(len(RESULTS), 1)\n    xs = np.arange(len(TARGETS))\n    for i, (name, r) in enumerate(RESULTS.items()):\n        ax[1].bar(xs + i * w, r[\"per_target\"], width=w, color=cols.get(name, INK),\n                  alpha=.85, label=name)\n    ax[1].axhline(0.5, color=INK, lw=.8, ls=\":\")\n    ax[1].set_xticks(xs + w * (len(RESULTS) - 1) / 2)\n    ax[1].set_xticklabels(TARGETS, rotation=55, ha=\"right\", fontsize=6.5)\n    ax[1].set_ylabel(\"holdout AUC\")\n    ax[1].set_ylim(0.4, 1.0)\n    ax[1].legend(fontsize=6.5, frameon=False)\n    ax[1].set_title(\"per target — §1 says the weakest one is the expensive one\",\n                    loc=\"left\", fontsize=8)\n\n    names = list(TEST_PREDS)\n    if len(names) > 1:\n        R = np.array([pd.DataFrame(TEST_PREDS[n]).rank(pct=True).values.ravel()\n                      for n in names])\n        C = np.corrcoef(R)\n        im = ax[2].imshow(C, cmap=\"Blues\", vmin=max(0, C.min() - .02), vmax=1)\n        ax[2].set_xticks(range(len(names)))\n        ax[2].set_xticklabels(names, fontsize=6.5)\n        ax[2].set_yticks(range(len(names)))\n        ax[2].set_yticklabels(names, fontsize=6.5)\n        ax[2].grid(False)\n        for a in range(len(names)):\n            for b in range(len(names)):\n                ax[2].text(b, a, f\"{C[a, b]:.2f}\", ha=\"center\", va=\"center\", fontsize=6.5,\n                           color=\"white\" if C[a, b] > .8 else INK)\n        ax[2].set_title(\"rank correlation between arms\\nlower is why the mean helps\",\n                        loc=\"left\", fontsize=8)\n    else:\n        ax[2].axis(\"off\")\n    fig.tight_layout()\n    plt.show()\n\n    summary = pd.DataFrame({n: {\"best holdout macro AUC\": r[\"best\"],\n                                \"slices/slot\": r.get(\"slices\"),\n                                \"epochs run\": len(r[\"hist\"]),\n                                \"picked\": SELECTED.count(n)}\n                            for n, r in RESULTS.items()}).T\n    print(summary.round(4).to_string())\n    per = pd.DataFrame({n: r[\"per_target\"] for n, r in RESULTS.items()}, index=TARGETS)\n    print(\"\\nper-target holdout AUC\")\n    print(per.round(3).to_string())\n\n    # The controlled coverage comparison, stated as a number rather than a claim.\n    curve = [(RESULTS[n][\"slices\"], RESULTS[n][\"best\"])\n             for n in (\"s168g4\", \"s168g8\") if n in RESULTS]\n    if len(curve) >= 2:\n        curve.sort()\n        print(\"\\ncoverage curve, everything else held fixed \"\n              \"(same encoder, grid, cache and cost per epoch):\")\n        for i, (sl, sc) in enumerate(curve):\n            d = f\"   {sc - curve[i - 1][1]:+.4f} vs previous\" if i else \"\"\n            print(f\"   {sl:2d} slices/slot   holdout {sc:.4f}{d}\")\n        a, b = RESULTS[\"s168g4\"], RESULTS[\"s168g8\"]\n        pt = pd.Series(b[\"per_target\"] - a[\"per_target\"], index=TARGETS).sort_values()\n        print(\"  per target, most helped by more slices:\")\n        print(\"   \", \", \".join(f\"{k} {v:+.3f}\" for k, v in pt.tail(4)[::-1].items()))\n        print(\"  least:\")\n        print(\"   \", \", \".join(f\"{k} {v:+.3f}\" for k, v in pt.head(3).items()))"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"EXPECTED_STAGES = {\"headers\", \"cache\", \"targets\", \"train\", \"ensemble\"}\nEXPECTED_ARMS = {\"s168g4\", \"s168g8\", \"b168g8\", \"b168g8b\", \"b168g8c\"}\n\nassert not SMOKE, \"SMOKE mode cannot produce a scored candidate\"\nassert EPOCHS == 16, EPOCHS\nassert abs(TIME_BUDGET - 7.6 * 3600) < 1e-9, TIME_BUDGET\nassert EXPECTED_STAGES.issubset(STAGE_OK), STAGE_OK\nassert all(STAGE_OK[name] is True for name in EXPECTED_STAGES), STAGE_OK\nassert set(RESULTS) == EXPECTED_ARMS, set(RESULTS)\nassert set(TEST_PREDS) == EXPECTED_ARMS, set(TEST_PREDS)\nfor _name in sorted(EXPECTED_ARMS):\n    assert len(RESULTS[_name][\"hist\"]) == EPOCHS, (_name, len(RESULTS[_name][\"hist\"]))\n    assert np.asarray(RESULTS[_name][\"test\"]).shape == (len(ST_TE), len(TARGETS))\nassert SELECTED and set(SELECTED).issubset(EXPECTED_ARMS), SELECTED\n\n_candidate_path = Path(\"_submission_candidate.csv\")\nassert _candidate_path.is_file(), \"final candidate matrix was not written\"\nassert not Path(\"submission.csv\").exists(), \"ordinary output name appeared before final gate\"\nsub = pd.read_csv(_candidate_path)\nassert len(sub) == len(_bench), (len(sub), len(_bench))\nassert list(sub.columns) == [\"StudyInstanceUID\"] + TARGETS, list(sub.columns)\nassert sub[\"StudyInstanceUID\"].notna().all()\nassert sub[\"StudyInstanceUID\"].is_unique\nassert sub[\"StudyInstanceUID\"].astype(str).tolist() == _bench[\"StudyInstanceUID\"].astype(str).tolist()\n_values = sub[TARGETS].apply(pd.to_numeric, errors=\"raise\").to_numpy(dtype=float)\nassert np.isfinite(_values).all()\nassert ((_values >= 0.0) & (_values <= 1.0)).all()\nassert np.unique(_values).size > 1, \"prediction matrix is globally constant\"\n_constant_targets = [t for t in TARGETS if sub[t].nunique(dropna=False) == 1]\nlog(f\"final gate passed: {sub.shape}, range [{_values.min():.3f}, {_values.max():.3f}], \"\n    f\"constant targets on this test set: {_constant_targets or 'none'}\")\nlog(f\"done in {(time.time() - T0) / 3600:.2f} h; stages {STAGE_OK}\")\n_audit_signal.alarm(0)\nos.replace(_candidate_path, \"submission.csv\")\n"},{"cell_type":"markdown","metadata":{},"source":"## 9. What this notebook is confident about, and what it is not\n\n**Confident.** The report reader is measured on two gauges that fail in different ways, and\nthe one that runs on all 4 407 studies — the silence rate — needs no labels at all, so it\ncannot be talked into agreeing. The geometry is not a matter of opinion: slice order is\nrecoverable exactly from patient coordinates, physical scale exactly from pixel spacing,\nand both are wrong by default if left alone. The cache arithmetic is arithmetic. Rank\naveraging follows from the metric.\n\n**Not confident.** Two things, named rather than buried.\n\n*The annotated subset cannot settle anything on its own.* Fifty-eight studies put a ±0.16\ninterval around a per-target AUC for a rare finding, and ±0.10 for a balanced one — the\ndrawn intervals bear that out. Every number reported against them carries that interval,\nand the rules in §2.1 were kept on the argument plus the coverage gauge, not on a point\nestimate.\n\n*Report labels have a ceiling this pipeline cannot cross.* The 58 annotations were read\nfrom the images by someone who was not the reporting radiologist, and the two disagree on a visible fraction of findings. The exact report examples are\nomitted from this derivative because narrative documentation must not redistribute corpus\ntext. That\ndisagreement is inside the training targets of all 4 407 studies, invisibly. It is the\nsingle largest thing standing between this notebook and a better score, and no amount of\nencoder capacity addresses it.\n\n### Three things that were believed and then measured\n\nThe useful part of running the same pipeline repeatedly is that it keeps disagreeing with\nthe reasoning that built it. All three of these were arguments this notebook made in\nearlier versions, in its own voice, before the arithmetic arrived:\n\n| claim | how it was tested | outcome |\n|---|---|---|\n| a 1 mm tear needs sub-0.5 mm pixels, so 266 px beats 224 px | same encoder, same slices, two arms | **wrong** — separated by ~0.001 macro AUC, twice |\n| capacity is the lever, so larger is better | small → base → large, held otherwise fixed | **half wrong** — base beat small by ~0.011, large beat neither |\n| every arm deserves an equal vote in the rank mean | greedy selection against equal weighting | equal weighting carries a known-weak arm; selection does not |\n| *(untested for three revisions)* how much of the joint is sampled | 3 → 6 → 12 slices, everything else identical | **the largest effect in the notebook: +0.033 macro AUC** |\n\nThe last row is the point. Resolution was argued for at length and bought a thousandth;\ncapacity bought a hundredth and then reversed; the axis nobody had measured bought three\nhundredths — thirty times the effect that the most carefully argued change produced. And\nonce coverage was fixed, capacity stopped mattering: at twelve slices, `small` and `base`\nland within a thousandth of each other, which means the earlier \"capacity is the lever\"\nresult was an artefact of a model starved of anatomy rather than of parameters.\n\nWhere the gain lands is the part worth reading. More slices help the **focal** findings —\nthe cruciate, the menisci, bone contusion — and do essentially nothing for the diffuse ones\n(the compartmental osteoarthritis targets, effusion). That is exactly what the anatomy\npredicts: a tear occupies a few slices and is either sampled or missed, while cartilage\nloss and joint fluid are visible on every group. The same split shows up independently in\nthe mean-versus-max choice, which the arms make per target without being told about it.\n\nNone of that was foreseeable from the reasoning. Each argument was sound and each\nconclusion was still wrong, which is the ordinary condition of this kind of work and the\nreason the arms are built so a comparison costs one decode pass rather than three.\n\n**A note on what did not change.** Every rewrite above touched the imaging side. Not one of\nthem moved the holdout as far as the report reader did. If you fork this notebook, the\nreport reader is where the score is.\n\n### Where the next point of score is\n\nNot in the encoder. §1's arithmetic says a target left near chance costs about $0.029$ of\nthe final score by itself, and the coverage gauge says exactly which targets sit closest to\nthat: the ones the reports are most often silent about. `Synovitis` is named in fewer than\none report in eight and annotated on nearly half the studies; `Fracture` and `Lateral OA`\nare close behind. Those three are where the vocabulary is thin, where the labels are\nweakest, and — by §1 — where a fixed unit of work buys the most.\n\nTwo concrete leads for anyone continuing:\n\n1. **More slices, and then attend over them rather than averaging.** The coverage curve\n   above says whether twelve is enough; if it is still climbing, the cheapest next move is\n   a smaller grid and more slices, since resolution has twice been shown not to matter\n   here. Beyond that, the head sees six slot embeddings and the slice groups are reduced\n   at the logit level afterwards. Letting the head attend over (slot × group) tokens would\n   let a focal finding be *found* on the group that shows it rather than diluted across\n   the rest — which is what the max-reduction is crudely approximating. That costs encoder\n   passes during training, which is why it is not here.\n2. **Read the reports with a language model rather than a lexicon.** The silence rate says\n   where morphology defeats a regular expression, and it is concentrated in particular\n   languages on particular findings. That is a bounded, measurable target.\n\n---\n\n*Reproducibility: every target is derived in-line from `train.csv`, so there is no external\nlabel table to go stale. `FEATURES` toggles the four non-obvious extraction rules,\n`SLOT_SCHEME=public` switches to the delivered protocol flags, and `SMOKE=1` runs the whole\ncontrol path on a subset. Feedback on the report reader is worth more than feedback on the\nmodel — that is where the score is.*"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11.13"},"release_provenance":{"upstream_owner":"prvsiyan","upstream_script_version_id":340808178,"co_attribution":"pilkwang/rsna-knee-baseline-v1 V14","validated_source":"beicicc/rsna-knee-research-prvsiyan-v9-safe-reproduction","validated_source_script_version_id":340921674}},"nbformat":4,"nbformat_minor":4}