{"cells":[{"cell_type":"markdown","id":"3ccba25e","metadata":{},"source":"# Read the report then the knee silver baseline\n\nThe companion notebook (Osteoarthritis is almost never written as OA) shows that the reports carry the\nlabels but exist only in training: `test.csv` is study identifiers only. So the deployable model must read\npixels. This notebook closes that loop end to end:\n\n1. Mine SILVER labels from the 4349 report only training studies with the assertion aware rule labeler.\n2. Extract image features per study: frozen DINOv2 ViT-S/14 embeddings (weights attached offline) plus\n   classical intensity and gradient statistics, one canonical series per plane.\n3. Train a per label head on the silver labels, holding out the 58 gold studies entirely as the anchor.\n4. Predict the test set from images alone and write submission.csv.\n\nThis is deliberately a baseline, not a leaderboard chase: the point is a clean, reproducible, images only\npipeline that turns report derived labels into a real submission. Every validation number is computed on\nthe 58 gold, which the model never trains on. No leaderboard position is claimed."},{"cell_type":"code","execution_count":null,"id":"7037d692","metadata":{},"outputs":[],"source":"import os, re, glob, math, time, warnings\nimport numpy as np, pandas as pd\nimport matplotlib as mpl, matplotlib.pyplot as plt\nwarnings.filterwarnings(\"ignore\")\nmpl.rcParams[\"axes.unicode_minus\"] = False   # ASCII hyphen in figures, never U+2212\nmpl.rcParams[\"figure.dpi\"] = 120; mpl.rcParams[\"font.size\"] = 11\nmpl.rcParams[\"axes.grid\"] = True; mpl.rcParams[\"grid.alpha\"] = 0.25\n\ndef resolve_input():\n    for c in [\"/kaggle/input/rsna-knee-abnormality-detection\",\n              \"/kaggle/input/competitions/rsna-knee-abnormality-detection\"]:\n        if os.path.exists(os.path.join(c, \"train.csv\")): return c\n    if os.path.isdir(\"/kaggle/input\"):\n        for dp, dirs, files in os.walk(\"/kaggle/input\"):\n            if \"train.csv\" in files: return dp\n            if dp.count(\"/\") - \"/kaggle/input\".count(\"/\") >= 3: dirs[:] = []\n    return \"D:/kaggle/rsna_data\"\nINPUT = resolve_input()\nBASE = \"train_series\" if os.path.isdir(os.path.join(INPUT, \"train_series\")) else None\nTBASE = \"test_series\" if os.path.isdir(os.path.join(INPUT, \"test_series\")) else None\nprint(\"INPUT =\", INPUT, \"| train_series dir:\", bool(BASE), \"| test_series dir:\", bool(TBASE))\n\ndef find_weights():\n    for dp, dirs, files in os.walk(\"/kaggle/input\"):\n        for f in files:\n            if f == \"dinov2_vits14.pth\": return os.path.join(dp, f)\n    p = \"D:/kaggle/rsna_data/dinov2_vits14.pth\"\n    return p if os.path.exists(p) else None\nWEIGHTS = find_weights()\nprint(\"DINOv2 weights:\", WEIGHTS)\n\nLABELS = [\"ACL\",\"MCL\",\"Medial Meniscus\",\"Lateral Meniscus\",\"Medial OA\",\"Lateral OA\",\n          \"PF OA\",\"Effusion\",\"Synovitis\",\"Baker's\",\"Contusion\",\"Fracture\"]\ntrain = pd.read_csv(os.path.join(INPUT, \"train.csv\"))\ntrain_series = pd.read_csv(os.path.join(INPUT, \"train_series.csv\"))\ntest = pd.read_csv(os.path.join(INPUT, \"test.csv\"))\ntest_series = pd.read_csv(os.path.join(INPUT, \"test_series.csv\"))\nprint(\"train\", train.shape, \"| test\", test.shape, \"| test_series\", test_series.shape)"},{"cell_type":"markdown","id":"46f6e813","metadata":{},"source":"## 1. Silver labels from the reports (train only)"},{"cell_type":"code","execution_count":null,"id":"3a60277e","metadata":{},"outputs":[],"source":"# RSNA Knee report labeler: language-agnostic, assertion-aware (ConText), no machine translation.\n# Lexicon built from EXTERNAL multilingual radiology terminology (standard finding names), not by\n# fitting to which tokens appear in the 58 gold. Emits soft per-label scores + abstain flags.\n# Pure stdlib + numpy + pandas. Importable; used identically by the notebook and by local validation.\nimport re, unicodedata\nimport numpy as np\n\nLABELS = [\"ACL\",\"MCL\",\"Medial Meniscus\",\"Lateral Meniscus\",\"Medial OA\",\"Lateral OA\",\n          \"PF OA\",\"Effusion\",\"Synovitis\",\"Baker's\",\"Contusion\",\"Fracture\"]\n\ndef norm(t):\n    t = str(t).lower()\n    t = unicodedata.normalize(\"NFKD\", t)\n    t = \"\".join(c for c in t if not unicodedata.combining(c))\n    return re.sub(r\"\\s+\", \" \", t)\n\n# ---- assertion cues (multilingual), matched inside a bounded scope ----\nNEG_PRE = [  # negation appearing BEFORE the finding\n    \"no \",\"not \",\"non \",\"without\",\"w/o\",\"absence of\",\"absent\",\"negative for\",\"no evidence\",\n    \"no sign\",\"free of\",\"rule out\",\"ruled out\",\"r/o\",\n    \"sin \",\"ausencia\",\"ausente\",\"descarta\",\"no hay\",\"no se observ\",\"no se identific\",\"sin signos\",\n    \"sin evidencia\",\"no existe\",\"no presenta\",\n    \"kein\",\"keine\",\"ohne\",\"nicht\",\"negativ\",\n    \"geen\",\"zonder\",\"niet\",\"negatief\",\n    \"pas de\",\"pas d\",\"aucun\",\"absence\",\n    \"senza\",\"assenza\",\"nessun\",\"non \",\n    \"sem \",\"nao \",\"yok\",\"olmayan\",\"izlenmedi\",\"saptanmadi\",\"gozlenmedi\",\n]\nNORMAL_POST = [  # normality / intactness AFTER the structure implies negative\n    \"intact\",\"normal\",\"unremarkable\",\"preserved\",\"conserved\",\"within normal\",\n    \"integr\",\"conservad\",\"normale\",\"normales\",\"sin alteracion\",\"sin particularidad\",\n    \"unauffallig\",\"regelrecht\",\"erhalten\",\"intakt\",\n    \"normaal\",\"ongestoord\",\n    \"sans particularite\",\"integre\",\n    \"normaldir\",\"dogal\",\"olagan\",\n]\nUNC = [  # uncertainty / hedge\n    \"possible\",\"possibly\",\"cannot exclude\",\"can not exclude\",\"suspicious\",\"suspected\",\"probable\",\n    \"likely\",\"questionable\",\"equivocal\",\"may represent\",\"concerning for\",\n    \"posible\",\"no se puede excluir\",\"sospech\",\"probable\",\"dudos\",\"sugestiv\",\"sugiere\",\n    \"moglich\",\"nicht auszuschliessen\",\"verdacht\",\"v.a\",\"dd \",\n    \"mogelijk\",\"verdenking\",\"waarschijnlijk\",\n    \"possible\",\"probable\",\"suspect\",\n    \"supheli\",\"olabilir\",\"suphesi\",\n]\nADVERSATIVE = [\"but \",\"however\",\"although\",\"though\",\"pero \",\"aunque\",\"sin embargo\",\"jedoch\",\"aber \",\n               \"echter\",\"maar \",\"mais \",\"cependant\",\"ma \",\"tuttavia\",\"porem\",\"mas \"]\n\n# ---- indication / clinical-question headers to DROP (state suspicion, not findings) ----\nINDIC_HDR = [\"indication\",\"clinical\",\"history\",\"reason\",\"question\",\"antecedent\",\"motivo\",\"clinica\",\n             \"vraagstelling\",\"inlichting\",\"fragestellung\",\"klinische\",\"anamnes\",\"indicac\",\"hikaye\",\n             \"klinik\",\"diagnostische vraag\",\"clinical inlichtingen\"]\nIMPRESS_HDR = [\"impression\",\"conclusion\",\"assessment\",\"opinion\",\"impresion\",\"conclusion\",\"besluit\",\n               \"beurteilung\",\"zusammenfassung\",\"conclusie\",\"sonuc\",\"diagnostico\",\"impressie\",\"fazit\"]\n\ndef split_sentences(text):\n    return [s.strip() for s in re.split(r\"[.\\n;:]+\", text) if s.strip()]\n\ndef split_scopes(clause):\n    # terminate scope at adversative conjunctions so \"no effusion but meniscus tear\" scopes correctly\n    parts = [clause]\n    for adv in ADVERSATIVE:\n        out = []\n        for p in parts:\n            out.extend(p.split(adv))\n        parts = out\n    return [p.strip() for p in parts if p.strip()]\n\ndef assertion(scope, term):\n    \"\"\"Return AFFIRMED / NEGATED / UNCERTAIN for a term mention inside its scope.\"\"\"\n    i = scope.find(term)\n    if i < 0:\n        return None\n    before = scope[max(0, i-55):i]\n    after = scope[i+len(term): i+len(term)+40]\n    if any(u in scope for u in UNC):\n        # uncertainty only if the hedge is near the term (same scope already bounds it)\n        if any(u in before or u in after for u in UNC):\n            return \"UNCERTAIN\"\n    if any(n in before for n in NEG_PRE):\n        return \"NEGATED\"\n    if any(x in after for x in NORMAL_POST):\n        return \"NEGATED\"\n    return \"AFFIRMED\"\n\n# ---- concept lexicon (external radiology vocabulary) ----\n# TEAR words for anatomy-based classes (ACL/MCL/menisci)\nTEAR = [norm(x) for x in [\"tear\",\"torn\",\"rupture\",\"ruptured\",\"rupt\",\"disrupt\",\"disruption\",\"rotura\",\n    \"roto\",\"rota\",\"desgarro\",\"ruptura\",\"desinsercion\",\"riss\",\"gerissen\",\"ruptur\",\"scheur\",\"gescheurd\",\n    \"ruptuur\",\"dechirure\",\"dechir\",\"rottura\",\"lacerazione\",\"yirtik\",\"yirtig\",\"rexis\",\"ρηξη\"]]\n# meniscus signals that should DEMOTE from a true tear\nMENISC_DEMOTE = [norm(x) for x in [\"grade 1\",\"grade i\",\"grade 2\",\"grade ii\",\"grado 1\",\"grado 2\",\n    \"intrasubstance\",\"intrasustancia\",\"mucoid\",\"mixoide\",\"degenerac\",\"degenerative signal\",\n    \"no surface\",\"does not contact\",\"sin contactar\",\"without surface\",\"gebni\"]]\n\ndef clause_has(scope, terms):\n    return any(t in scope for t in terms)\n\n# structure regex fragments per compartment/ligament (accent-folded)\nS_ACL = [norm(x) for x in [\"acl\",\"anterior cruciate\",\"cruzado anterior\",\"lca\",\"vorderes kreuzband\",\"vkb\",\n    \"voorste kruisband\",\"ligament croise anterieur\",\"crociato anteriore\",\"on capraz bag\",\"kreuzband vorder\"]]\nS_MCL = [norm(x) for x in [\"mcl\",\"medial collateral\",\"colateral medial\",\"colateral interno\",\"lli\",\"lcm\",\n    \"ligamento lateral interno\",\"innenband\",\"mediales seitenband\",\"mediale collaterale\",\"medial kollateral\",\n    \"ligament collateral medial\",\"ligament lateral interne\",\"ic bag\",\"medial band\"]]\nS_MMEN = [norm(x) for x in [\"medial meniscus\",\"menisco interno\",\"menisco medial\",\"innenmeniskus\",\n    \"medialer meniskus\",\"mediale meniscus\",\"binnenmeniscus\",\"menisque interne\",\"menisque medial\",\n    \"menisco mediale\",\"medial meniskus\",\"ic menisk\"]]\nS_LMEN = [norm(x) for x in [\"lateral meniscus\",\"menisco externo\",\"menisco lateral\",\"aussenmeniskus\",\n    \"lateraler meniskus\",\"laterale meniscus\",\"buitenmeniscus\",\"menisque externe\",\"menisque lateral\",\n    \"menisco laterale\",\"lateral meniskus\",\"dis menisk\"]]\n\n# finding-based single concepts\nC_EFFUSION = [norm(x) for x in [\"effusion\",\"joint effusion\",\"joint fluid\",\"synovial fluid\",\"derrame\",\n    \"derrame articular\",\"erguss\",\"gelenkerguss\",\"epanchement\",\"versamento\",\"effusie\",\"gewrichtsvocht\",\n    \"hydrops\",\"hidrops\",\"intraarticular fluid\",\"intra-articular fluid\",\"liquido articular\",\n    \"liquido intraarticular\",\"liquido intra-articular\",\"eklem sivisi\",\"hemartros\"]]\nC_SYNOV = [norm(x) for x in [\"synovitis\",\"synovial thickening\",\"synovial hypertroph\",\"sinovitis\",\n    \"synovialitis\",\"synovite\",\"sinovite\",\"synoviale verdikking\",\"engrosamiento sinovial\",\n    \"proliferacion sinovial\",\"hypertrophie synoviale\",\"hoffitis\",\"sinovial hipertrofia\"]]\nC_BAKER = [norm(x) for x in [\"baker\",\"popliteal cyst\",\"quiste de baker\",\"quiste poplite\",\"quiste popliteo\",\n    \"bakerzyste\",\"poplitealzyste\",\"bakercyste\",\"poplitea cyste\",\"kyste de baker\",\"kyste poplite\",\n    \"cisti di baker\",\"cisti poplitea\",\"baker kisti\",\"popliteal kist\"]]\nC_CONTUS = [norm(x) for x in [\"contusion\",\"bone bruise\",\"bone marrow edema\",\"bone marrow oedema\",\n    \"marrow edema\",\"marrow oedema\",\"contusion osea\",\"edema oseo\",\"edema medular\",\"edema de medula\",\n    \"knochenmarkodem\",\"botcontusie\",\"botoedeem\",\"beenmergoedeem\",\"contusion osseuse\",\"oedeme osseux\",\n    \"oedeme medullaire\",\"subchondral edema\",\"subchondral bone marrow edema\",\"bone edema\",\"kemik odem\",\n    \"kontuzyon\",\"edema oseo subcondral\"]]\nC_FRACT = [norm(x) for x in [\"fracture\",\"fractura\",\"fraktur\",\"fractuur\",\"frattura\",\"fratura\",\"avulsion\",\n    \"insufficiency fracture\",\"stress fracture\",\"subchondral fracture\",\"fisura osea\",\"kirik\",\"fissur oss\"]]\n\n# OA consequence vocabulary (compartment assigned separately)\nOA_ANY = [norm(x) for x in [\"osteoarthr\",\"artrosis\",\"artrose\",\"gonartros\",\"gonarthros\",\"arthrose\",\n    \"osteophyt\",\"osteofit\",\"osteofyt\",\"joint space narrow\",\"joint space loss\",\"jsn\",\"joint-space narrow\",\n    \"pinzamiento\",\"estrechamiento\",\"reduccion de la interlinea\",\"interlinea articular reducida\",\n    \"chondral loss\",\"cartilage loss\",\"cartilage thinning\",\"chondrosis\",\"condrosis\",\"condromalacia\",\n    \"condropatia\",\"chondromalacia\",\"chondropath\",\"ulcera condral\",\"ulceras condral\",\"condral defect\",\n    \"chondral defect\",\"degenerativ\",\"degeneratief\",\"mucoid degeneration\",\"subchondral cyst\",\n    \"subchondral sclerosis\",\"esclerosis subcondral\",\"kellgren\",\"degenerative change\",\"kraakbeenverlies\"]]\nOA_TRICOMP = [norm(x) for x in [\"tricompartmental\",\"three compartment\",\"three-compartment\",\"pancompartmental\",\n    \"pangonarthros\",\"incipient oa of all\",\"oa of all three\",\"tricompartimental\",\"all three compartment\"]]\nCTX_MED = [norm(x) for x in [\"medial\",\"interno\",\"interna\",\"medialis\",\"binnen\",\"femorotibial medial\",\n    \"femorotibial interna\",\"medial compartment\",\"compartimento medial\",\"mediale femorotibiaal\",\"mediaal\"]]\nCTX_LAT = [norm(x) for x in [\"lateral\",\"externo\",\"externa\",\"buiten\",\"femorotibial lateral\",\n    \"femorotibial externa\",\"lateral compartment\",\"compartimento lateral\",\"laterale femorotibiaal\"]]\nCTX_PF = [norm(x) for x in [\"patellofemoral\",\"femoropatel\",\"retropatel\",\"patelofemoral\",\"rotulian\",\n    \"patellar\",\"patela \",\"rotula\",\"troclea\",\"trochlea\",\"femoro-patellaire\",\"femoropatellar\",\"patello-femoral\"]]\nOA_GENERIC = [norm(x) for x in [\"gonartros\",\"gonarthros\",\"osteoarthr\",\"artrosis\",\"artrose\",\"arthrose\"]]\n\ndef score_report(report, use_negation=True, use_oa_vocab=True):\n    \"\"\"Return dict label -> evidence score e (real). Higher = more likely positive.\n    use_negation/use_oa_vocab toggles support the ablation ladder.\"\"\"\n    raw = norm(report)\n    lines = [l.strip() for l in re.split(r\"[\\n]+\", raw) if l.strip()]\n    # section weighting: mark impression lines x1.5, drop indication lines\n    weighted = []  # (text, weight)\n    in_indic = False\n    for l in lines:\n        head = l[:40]\n        if any(h in head for h in INDIC_HDR) and (\":\" in l[:40] or len(l) < 60):\n            in_indic = True\n            continue\n        if any(h in head for h in IMPRESS_HDR):\n            in_indic = False\n            weighted.append((l, 1.5)); continue\n        # a findings/results header ends indication\n        if any(h in head for h in [\"finding\",\"result\",\"hallazgo\",\"bevinding\",\"befund\",\"bulgular\",\"resultat\"]):\n            in_indic = False; weighted.append((l, 1.0)); continue\n        if in_indic:\n            continue\n        weighted.append((l, 1.0))\n    if not weighted:\n        weighted = [(raw, 1.0)]\n\n    acc = {l: 0.0 for l in LABELS}\n    W_AFF, W_UNC, W_NEG = 1.0, 0.4, 1.0\n\n    def add(label, state, w):\n        if state is None:\n            return\n        if not use_negation:\n            # crude keyword-presence baseline: any detected mention counts as positive\n            acc[label] += W_AFF * w\n            return\n        if state == \"AFFIRMED\": acc[label] += W_AFF * w\n        elif state == \"UNCERTAIN\": acc[label] += W_UNC * w\n        elif state == \"NEGATED\": acc[label] -= W_NEG * w\n\n    for text, w in weighted:\n        for clause in split_sentences(text):\n            for scope in split_scopes(clause):\n                # anatomy-based classes: structure + tear in scope\n                for label, struct in ((\"ACL\",S_ACL),(\"MCL\",S_MCL),\n                                      (\"Medial Meniscus\",S_MMEN),(\"Lateral Meniscus\",S_LMEN)):\n                    if clause_has(scope, struct) and clause_has(scope, TEAR):\n                        # demote meniscus grade-1/2/intrasubstance/mucoid to uncertain\n                        st = None\n                        for s in struct:\n                            if s in scope:\n                                st = assertion(scope, s); break\n                        if st is None:\n                            continue\n                        if label in (\"Medial Meniscus\",\"Lateral Meniscus\") and st == \"AFFIRMED\" \\\n                           and clause_has(scope, MENISC_DEMOTE):\n                            st = \"UNCERTAIN\"\n                        add(label, st, w)\n                # finding-based single concepts\n                for label, terms in ((\"Effusion\",C_EFFUSION),(\"Synovitis\",C_SYNOV),\n                                     (\"Baker's\",C_BAKER),(\"Contusion\",C_CONTUS),(\"Fracture\",C_FRACT)):\n                    for t in terms:\n                        if t in scope:\n                            add(label, assertion(scope, t), w); break\n                # OA classes\n                if use_oa_vocab:\n                    if clause_has(scope, OA_TRICOMP):\n                        for lab in (\"Medial OA\",\"Lateral OA\",\"PF OA\"):\n                            acc[lab] += W_AFF * w\n                    for o in OA_ANY:\n                        if o in scope:\n                            st = assertion(scope, o)\n                            hit_ctx = False\n                            if clause_has(scope, CTX_MED): add(\"Medial OA\", st, w); hit_ctx=True\n                            if clause_has(scope, CTX_LAT): add(\"Lateral OA\", st, w); hit_ctx=True\n                            if clause_has(scope, CTX_PF):  add(\"PF OA\", st, w); hit_ctx=True\n                            if not hit_ctx and any(g in scope for g in OA_GENERIC):\n                                # bare knee OA -> low-weight signal to all three\n                                for lab in (\"Medial OA\",\"Lateral OA\",\"PF OA\"):\n                                    add(lab, st, 0.5*w)\n                            break\n                else:\n                    # naive OA: only the literal words \"osteoarthritis/oa/artrosis\" with compartment\n                    for o in [norm(x) for x in [\"osteoarthr\",\"artrosis\",\"oa \"]]:\n                        if o in scope:\n                            st = assertion(scope, o)\n                            if clause_has(scope, CTX_MED): add(\"Medial OA\", st, w)\n                            if clause_has(scope, CTX_LAT): add(\"Lateral OA\", st, w)\n                            if clause_has(scope, CTX_PF):  add(\"PF OA\", st, w)\n                            break\n    return acc\n\ndef sigmoid(x): return 1.0 / (1.0 + np.exp(-x))\n\ndef evidence_frame(reports, use_negation=True, use_oa_vocab=True):\n    \"\"\"Return a DataFrame of raw evidence scores (columns=LABELS) for a series of reports.\"\"\"\n    import pandas as pd\n    rows = [score_report(r, use_negation, use_oa_vocab) for r in reports]\n    return pd.DataFrame(rows)[LABELS]\n"},{"cell_type":"code","execution_count":null,"id":"62a87179","metadata":{},"outputs":[],"source":"is_labeled = train[LABELS].notna().any(axis=1)\ngold = train[is_labeled].copy().reset_index(drop=True)\nreport_only = train[~is_labeled].copy().reset_index(drop=True)\nprint(\"gold (held out):\", len(gold), \"| report-only (silver train):\", len(report_only))\n\nS_ro = evidence_frame(report_only[\"Report\"].tolist(), True, True).values.astype(float)\nsilver_bin = (S_ro > 0).astype(int)          # binary silver target for the head\ny_gold = gold[LABELS].fillna(0).astype(int).values\nprint(\"silver positive counts per label:\", dict(zip(LABELS, silver_bin.sum(0).tolist())))"},{"cell_type":"markdown","id":"234f1007","metadata":{},"source":"## 2. Image features (DINOv2 embedding + classical), one canonical series per plane"},{"cell_type":"code","execution_count":null,"id":"965b9167","metadata":{},"outputs":[],"source":"# Classical per-slice DICOM image features (offline, no pretrained weights). Robust fallback + plumbing.\nimport numpy as np\n\ndef window_slice(ds):\n    import pydicom  # noqa\n    a = ds.pixel_array.astype(np.float32)\n    if str(ds.get(\"PhotometricInterpretation\", \"\")) == \"MONOCHROME1\":\n        a = a.max() - a\n    lo, hi = np.percentile(a, [1, 99])\n    a = np.clip((a - lo) / max(hi - lo, 1e-6), 0, 1)\n    return a\n\ndef resize2d(a, s=128):\n    # simple bilinear-ish via numpy (avoid cv2/PIL dependency drift)\n    from numpy import linspace\n    h, w = a.shape\n    yi = np.clip((linspace(0, h - 1, s)).round().astype(int), 0, h - 1)\n    xi = np.clip((linspace(0, w - 1, s)).round().astype(int), 0, w - 1)\n    return a[yi][:, xi]\n\ndef slice_features(a):\n    \"\"\"a is a 0..1 windowed 2D array. Return a small fixed feature vector.\"\"\"\n    a = resize2d(a, 128)\n    gx = np.abs(np.diff(a, axis=1)); gy = np.abs(np.diff(a, axis=0))\n    gm = np.sqrt(gx[:-1, :]**2 + gy[:, :-1]**2) if gx.size and gy.size else np.array([0.0])\n    p = np.percentile(a, [5, 25, 50, 75, 95])\n    feats = [\n        a.mean(), a.std(),\n        p[0], p[1], p[2], p[3], p[4],\n        (a > 0.8).mean(),   # bright fraction (fluid/effusion on fluid-sensitive)\n        (a < 0.1).mean(),   # dark fraction\n        gm.mean(), gm.std(), np.percentile(gm, 95),\n    ]\n    return np.asarray(feats, dtype=np.float32)\n\nFEAT_NAMES = [\"mean\",\"std\",\"p5\",\"p25\",\"p50\",\"p75\",\"p95\",\"bright\",\"dark\",\"grad_mean\",\"grad_std\",\"grad_p95\"]\n\n"},{"cell_type":"code","execution_count":null,"id":"2bd9b50d","metadata":{},"outputs":[],"source":"import torch, timm\nDEV = \"cpu\"\nif torch.cuda.is_available():\n    try:\n        torch.zeros(4, device=\"cuda\").sum().item(); DEV = \"cuda\"\n    except Exception as e:\n        print(\"cuda smoke failed:\", e)\nprint(\"device:\", DEV)\n\nMODEL = None; EMB_DIM = 0\nif WEIGHTS is not None:\n    try:\n        MODEL = timm.create_model(\"vit_small_patch14_dinov2.lvd142m\", pretrained=False, num_classes=0)\n        sd = torch.load(WEIGHTS, map_location=\"cpu\", weights_only=True)\n        miss, unexp = MODEL.load_state_dict(sd, strict=False)\n        MODEL.eval().to(DEV)\n        EMB_DIM = MODEL.num_features\n        CFG = MODEL.default_cfg\n        INP = CFG[\"input_size\"][1]; MEAN = np.array(CFG[\"mean\"]); STD = np.array(CFG[\"std\"])\n        print(f\"DINOv2 loaded: missing={len(miss)} unexpected={len(unexp)} emb_dim={EMB_DIM} input={INP}\")\n    except Exception as e:\n        print(\"DINOv2 load FAILED, classical-only fallback:\", repr(e)); MODEL = None\nN_CLASSICAL = len(FEAT_NAMES)\nPLANES = [\"Sagittal\", \"Axial\", \"Coronal\"]\nPER_PLANE = EMB_DIM + N_CLASSICAL\nFEAT_LEN = len(PLANES) * PER_PLANE\nprint(\"feature length per study:\", FEAT_LEN)"},{"cell_type":"code","execution_count":null,"id":"cb7b4a7b","metadata":{},"outputs":[],"source":"def pick_series(rows, plane):\n    sub = rows[rows.Anatomical_Plane == plane]\n    if len(sub) == 0: return None\n    sub = sub.sort_values([\"Fluid_Sensitive\", \"Fat_Suppression\"], ascending=False)\n    return sub.iloc[0][\"SeriesInstanceUID\"]\n\ndef read_slices(series_dir, K=2):\n    files = sorted(glob.glob(os.path.join(series_dir, \"*.dcm\")))\n    if not files: return []\n    mid = len(files) // 2\n    lo = max(0, mid - K // 2)\n    return files[lo:lo + K] or files[:K]\n\n@torch.no_grad()\ndef embed_batch(imgs):\n    if MODEL is None or not imgs: return None\n    xs = []\n    for a in imgs:\n        yi = np.clip(np.round(np.linspace(0, a.shape[0]-1, INP)).astype(int), 0, a.shape[0]-1)\n        xi = np.clip(np.round(np.linspace(0, a.shape[1]-1, INP)).astype(int), 0, a.shape[1]-1)\n        b = a[yi][:, xi]\n        x = np.stack([b, b, b], 0)\n        x = (x - MEAN.reshape(3,1,1)) / STD.reshape(3,1,1)\n        xs.append(x)\n    t = torch.tensor(np.stack(xs), dtype=torch.float32, device=DEV)\n    out = MODEL(t).float().cpu().numpy()\n    return out\n\ndef study_feature(study_id, series_df, base, K=2):\n    import pydicom\n    rows = series_df[series_df.StudyInstanceUID == study_id]\n    vec = np.zeros(FEAT_LEN, dtype=np.float32)\n    for pi, pl in enumerate(PLANES):\n        sid = pick_series(rows, pl)\n        if sid is None: continue\n        files = read_slices(os.path.join(INPUT, base, study_id, sid), K=K)\n        if not files: continue\n        imgs, classic = [], []\n        for f in files:\n            try:\n                ds = pydicom.dcmread(f)\n                a = window_slice(ds)\n                imgs.append(a); classic.append(slice_features(a))\n            except Exception:\n                pass\n        if not imgs: continue\n        off = pi * PER_PLANE\n        if MODEL is not None:\n            emb = embed_batch(imgs)\n            if emb is not None: vec[off:off+EMB_DIM] = emb.mean(0)\n        vec[off+EMB_DIM:off+PER_PLANE] = np.mean(classic, 0)\n    return vec\n\ndef build_features(study_ids, series_df, base, tag, K=2):\n    t0 = time.time(); X = np.zeros((len(study_ids), FEAT_LEN), dtype=np.float32)\n    for i, sid in enumerate(study_ids):\n        X[i] = study_feature(sid, series_df, base, K=K)\n        if (i+1) % 500 == 0:\n            print(f\"  {tag}: {i+1}/{len(study_ids)}  ({(time.time()-t0)/60:.1f} min)\", flush=True)\n    print(f\"  {tag}: done {len(study_ids)} in {(time.time()-t0)/60:.1f} min\", flush=True)\n    return X"},{"cell_type":"markdown","id":"0cb301d5","metadata":{},"source":"## 3. Extract features (train silver, held-out gold, test)"},{"cell_type":"code","execution_count":null,"id":"b60d34b1","metadata":{},"outputs":[],"source":"K_SLICES = 2\nif BASE is None:\n    print(\"no train_series DICOM dir present (local smoke); skipping heavy extraction\")\n    X_silver = np.random.rand(len(report_only), FEAT_LEN).astype(np.float32)\n    X_gold   = np.random.rand(len(gold), FEAT_LEN).astype(np.float32)\n    X_test   = np.random.rand(len(test), FEAT_LEN).astype(np.float32)\nelse:\n    X_silver = build_features(report_only[\"StudyInstanceUID\"].tolist(), train_series, BASE, \"silver\", K=K_SLICES)\n    X_gold   = build_features(gold[\"StudyInstanceUID\"].tolist(), train_series, BASE, \"gold\", K=K_SLICES)\n    X_test   = build_features(test[\"StudyInstanceUID\"].tolist(), test_series, TBASE, \"test\", K=K_SLICES)\nprint(\"shapes:\", X_silver.shape, X_gold.shape, X_test.shape)"},{"cell_type":"markdown","id":"cceb73d4","metadata":{},"source":"## 4. Train per-label head on silver, validate on the 58 gold"},{"cell_type":"code","execution_count":null,"id":"e60dce69","metadata":{},"outputs":[],"source":"from sklearn.linear_model import LogisticRegression\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.metrics import roc_auc_score\n\nscaler = StandardScaler().fit(X_silver)\nXs = scaler.transform(X_silver); Xg = scaler.transform(X_gold); Xt = scaler.transform(X_test)\n\ntest_pred = np.full((len(test), 12), 0.5, dtype=float)\ngold_pred = np.full((len(gold), 12), 0.5, dtype=float)\nfor j, lab in enumerate(LABELS):\n    yj = silver_bin[:, j]\n    if len(set(yj)) < 2:\n        continue\n    clf = LogisticRegression(max_iter=2000, C=1.0, class_weight=\"balanced\")\n    clf.fit(Xs, yj)\n    gold_pred[:, j] = clf.predict_proba(Xg)[:, 1]\n    test_pred[:, j] = clf.predict_proba(Xt)[:, 1]\n\naucs = []\nfor j, lab in enumerate(LABELS):\n    yj = y_gold[:, j]\n    a = roc_auc_score(yj, gold_pred[:, j]) if len(set(yj)) > 1 else float(\"nan\")\n    aucs.append(a)\n    print(f\"  {lab:18s} n+={int(yj.sum()):3d}  gold AUC={a:.3f}\")\nmacro = np.nanmean(aucs)\nprint(f\"\\nIMAGE MODEL macro AUC on held-out 58 gold = {macro:.3f}\")\nprint(\"(trained only on 4349 silver studies; the 58 gold were never used for training)\")"},{"cell_type":"markdown","id":"5363fc54","metadata":{},"source":"## 4b. Diagnostics: what the images-only head learned\n\nRich visuals on the held-out 58 gold. Note on wording: \"silver\" here means weak-supervision silver labels\nmined from the reports, not a medal. Every number below is computed live from the gold predictions."},{"cell_type":"code","execution_count":null,"id":"1e4ce631","metadata":{},"outputs":[],"source":"# DICOM montage of one held-out gold study across planes (grounds what the model consumes)\nimport pydicom\ntry:\n    done = False\n    if BASE is not None and len(gold):\n        sid = gold[\"StudyInstanceUID\"].iloc[0]\n        rws = train_series[train_series.StudyInstanceUID == sid]\n        fig, axes = plt.subplots(1, 3, figsize=(12, 4.2))\n        for i, pl in enumerate([\"Sagittal\", \"Axial\", \"Coronal\"]):\n            axes[i].set_axis_off(); axes[i].set_title(pl)\n            serid = pick_series(rws, pl)\n            if serid is None: continue\n            files = read_slices(os.path.join(INPUT, BASE, sid, serid), K=1)\n            if not files: continue\n            try:\n                a = window_slice(pydicom.dcmread(files[0]))\n                axes[i].imshow(resize2d(a, 300), cmap=\"gray\", vmin=0, vmax=1)\n            except Exception: pass\n        fig.suptitle(\"One held-out gold study across planes (identifiers not shown)\")\n        plt.tight_layout(); plt.show(); done = True\n    if not done: print(\"montage renders on Kaggle (DICOM not present here)\")\nexcept Exception as e:\n    print(\"montage skipped:\", repr(e))"},{"cell_type":"code","execution_count":null,"id":"f419b665","metadata":{},"outputs":[],"source":"# per-label held-out gold AUC bar, sorted, with chance line\nfig, ax = plt.subplots(figsize=(9, 5))\nav = np.array(aucs, dtype=float); oi = np.argsort(np.nan_to_num(av, nan=0.0))\nax.barh(range(12), av[oi], color=\"#3d7ea8\")\nax.axvline(0.5, ls=\"--\", color=\"k\", lw=1)\nax.set_yticks(range(12)); ax.set_yticklabels([f\"{LABELS[j]} (n+={int(y_gold[:,j].sum())})\" for j in oi])\nax.set_xlabel(\"held-out gold ROC AUC\"); ax.set_xlim(0, 1)\nfor i, j in enumerate(oi):\n    if np.isfinite(av[j]): ax.text(av[j] + 0.01, i, f\"{av[j]:.2f}\", va=\"center\", fontsize=8)\nax.set_title(f\"Images-only head, per-label AUC on the 58 gold (macro {macro:.3f})\")\nplt.tight_layout(); plt.show()"},{"cell_type":"code","execution_count":null,"id":"fee9f1bd","metadata":{},"outputs":[],"source":"# per-label ROC small multiples on the 58 gold\nfrom sklearn.metrics import roc_curve\nfig, axes = plt.subplots(3, 4, figsize=(14, 10)); axes = axes.ravel()\nfor j, lab in enumerate(LABELS):\n    ax = axes[j]; yj = y_gold[:, j]\n    if len(set(yj)) > 1:\n        fpr, tpr, _ = roc_curve(yj, gold_pred[:, j]); ax.plot(fpr, tpr, color=\"#c0603a\", lw=2)\n    ax.plot([0, 1], [0, 1], ls=\"--\", color=\"k\", lw=0.8)\n    ax.set_title(f\"{lab}  AUC {av[j]:.2f}\", fontsize=9)\n    ax.set_xlim(0, 1); ax.set_ylim(0, 1); ax.set_xticks([0, 1]); ax.set_yticks([0, 1])\nfig.suptitle(\"Per-label ROC on the 58 held-out gold (images-only head)\")\nplt.tight_layout(); plt.show()"},{"cell_type":"code","execution_count":null,"id":"1355e668","metadata":{},"outputs":[],"source":"# image-feature PCA of the gold studies, colored by Effusion (the co-occurrence hub)\nfrom sklearn.decomposition import PCA\nimport matplotlib.patches as mpatches\ntry:\n    P = PCA(n_components=2).fit(Xs)\n    Zg = P.transform(Xg); j = LABELS.index(\"Effusion\")\n    fig, ax = plt.subplots(figsize=(6.6, 5.6))\n    ax.scatter(Zg[:, 0], Zg[:, 1], c=y_gold[:, j], cmap=\"coolwarm\", s=45, edgecolor=\"k\", linewidth=0.3)\n    ax.set_title(\"Held-out gold in image-feature PCA space, colored by Effusion\")\n    ax.set_xlabel(\"PC1\"); ax.set_ylabel(\"PC2\")\n    ax.legend(handles=[mpatches.Patch(color=plt.cm.coolwarm(0.0), label=\"Effusion 0\"),\n                       mpatches.Patch(color=plt.cm.coolwarm(1.0), label=\"Effusion 1\")], loc=\"best\")\n    plt.tight_layout(); plt.show()\nexcept Exception as e:\n    print(\"PCA skipped:\", repr(e))"},{"cell_type":"code","execution_count":null,"id":"1f0652ae","metadata":{},"outputs":[],"source":"# series coverage by plane, train vs test (why per-plane routing is available at inference)\nfig, axes = plt.subplots(1, 2, figsize=(12, 4.2))\nfor ax, (name, ser) in zip(axes, [(\"train\", train_series), (\"test\", test_series)]):\n    vc = ser[\"Anatomical_Plane\"].value_counts()\n    ax.bar(range(len(vc)), vc.values, color=\"#3d7ea8\")\n    ax.set_xticks(range(len(vc))); ax.set_xticklabels(list(vc.index), rotation=20)\n    ax.set_title(f\"{name} series by plane\")\n    for i, v in enumerate(vc.values): ax.text(i, v, str(int(v)), ha=\"center\", va=\"bottom\", fontsize=8)\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"6e2d0e51","metadata":{},"source":"## 5. Write submission.csv"},{"cell_type":"code","execution_count":null,"id":"2d92ccc8","metadata":{},"outputs":[],"source":"sub = pd.DataFrame({\"StudyInstanceUID\": test[\"StudyInstanceUID\"].values})\nfor j, lab in enumerate(LABELS):\n    sub[lab] = np.clip(test_pred[:, j], 1e-6, 1 - 1e-6).round(5)\nsub = sub[[\"StudyInstanceUID\"] + LABELS]\nsub.to_csv(\"submission.csv\", index=False)\nprint(\"wrote submission.csv:\", sub.shape)\nprint(sub.head().to_string())\nassert list(sub.columns) == [\"StudyInstanceUID\"] + LABELS, \"column order must match sample_submission\"\nprint(\"column order OK\")"},{"cell_type":"markdown","id":"603cf9c2","metadata":{},"source":"## Notes and limits\n\n- The head is trained only on the 4349 report only studies (silver targets); the 58 gold are a clean\n  held out anchor, so the reported macro AUC is not inflated by training on the evaluation set.\n- Silver labels are a noisy teacher (macro AUC about 0.73 versus gold in the companion notebook), so the\n  image model is bounded by that teacher plus the information the images carry; this is a baseline.\n- Features are frozen DINOv2 embeddings plus classical statistics over two central slices per plane; more\n  slices, per finding plane routing, and fine tuning are the obvious next steps.\n- Reports are never read at inference (test has none). The pipeline is images only, internet off, with the\n  DINOv2 weights attached as a public dataset. No leaderboard position is claimed."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python"}},"nbformat":4,"nbformat_minor":5}