{"cells":[{"cell_type":"markdown","id":"7a214bd1e71b","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:10px solid #2a78d6;border-radius:8px;padding:26px 30px;margin:6px 0 20px 0;\"><div style=\"color:#c04b1c;font-size:12px;font-weight:700;letter-spacing:.16em;text-transform:uppercase;margin-bottom:10px;\">Baseline submission</div><h1 style=\"color:#0b0b0b;font-size:31px;line-height:1.22;margin:0;font-weight:700;letter-spacing:-.01em;\">RSNA Knee Abnormality Detection: a metadata-only baseline</h1><div style=\"color:#6a6862;font-size:13px;margin-top:12px;\">Twelve findings per study &middot; macro-averaged ROC AUC &middot; rule-extracted labels &middot; metadata only, no pixel is read</div></div>\n\n> **How this was made.** This notebook was written automatically by Claude Code running a\n> custom multi-agent LLM harness. Two rounds of adversarial verifier agents executed it end\n> to end against a local mirror of the Kaggle input mount and simulated the hidden rerun\n> against study ids it had never seen, with patients deliberately duplicated across\n> studies to attack the fold grouping; three real defects went back to the builder script\n> and were fixed there. It is published to Kaggle and submitted to the competition with\n> the Kaggle CLI. **Code cells are collapsed by default** so this reads as a report; click\n> any *Show code* toggle to expand one. It has no leaderboard score yet and claims none.\n> Full detail is in the last section.\n\nThis notebook is a **plumbing proof and a floor to beat**, not an attempt at the metric.\nIt exists to prove one thing end to end on the hidden test set: that a submission can be\nproduced from data the test studies actually carry, scored on all twelve labels, without\ntouching a single pixel.\n\nEverything here is deliberately weak. The labels come from a hand-written multilingual\nkeyword extractor with a negation guard, which is exactly the artifact a proper LLM\nextraction has to beat. The features come from DICOM headers and the series manifest,\nwhich describe the scanner and the protocol rather than the knee. A model built on those\ntwo things can only pick up site and protocol correlations with disease prevalence. That\nis enough to clear 0.5 macro AUC and nothing more.\n\nRead it as a scaffold. Every section is separated so the next version can replace one of\nthem without disturbing the rest."},{"cell_type":"markdown","id":"8a9c3af54ffe","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">1</span><span style=\"vertical-align:middle;\">What the competition actually asks</span></h2></div>\n\nPredict, per study, the probability of twelve findings: ACL, MCL, Medial Meniscus,\nLateral Meniscus, Medial OA, Lateral OA, PF OA, Effusion, Synovitis, Baker's,\nContusion, Fracture. The score is the macro average of the twelve per-label ROC\nAUCs, so a rare label counts exactly as much as a common one.\n\nTwo facts shape every decision below.\n\n**Labels are scarce.** Only a small subset of the training studies carry per-condition\nlabels. The rest ship a free-text radiology report, and the host says outright that\nyou may wish to derive labels from it. So this is a label-extraction problem before\nit is an imaging problem, and model architecture is downstream of label quality.\n\n**Reports exist at training time only.** `test.csv` has one column,\n`StudyInstanceUID`. There is no report at inference. Text is a label source and an\nauxiliary training signal; a pipeline that reads report text at prediction time\nscores nothing. That constraint is why section 4 below is limited to the series\nmanifest and DICOM headers."},{"cell_type":"markdown","id":"47c6e2a0fb82","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">2</span><span style=\"vertical-align:middle;\">Setup</span></h2></div>\n\nSeeds are fixed everywhere. Thread counts are pinned before NumPy and LightGBM are\nimported, because setting them afterwards has no effect. The input root is detected\nrather than hardcoded, so the same notebook runs on the Kaggle mount and on a local\nmirror of it."},{"cell_type":"code","execution_count":null,"id":"fbddc109f9df","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"import os\n\n# Pin thread counts BEFORE numpy / lightgbm import. Setting them afterwards is a no-op.\nfor _v in (\"OMP_NUM_THREADS\", \"OPENBLAS_NUM_THREADS\", \"MKL_NUM_THREADS\",\n           \"NUMEXPR_NUM_THREADS\", \"VECLIB_MAXIMUM_THREADS\"):\n    os.environ.setdefault(_v, \"4\")\n\nimport random\nimport re\nimport time\nimport unicodedata\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\n\nSEED = 42\nrandom.seed(SEED)\nnp.random.seed(SEED)\n\n# Bounds on the DICOM header walk. Header tags are constant within a series, so two\n# slices per series is enough; the second one only guards against a corrupt first file.\nMAX_SLICES_PER_SERIES = 2\nMAX_SERIES_PER_STUDY = 16\nPROBE_STUDIES = 24\nTRAIN_FEATURE_BUDGET_S = 2400.0\nMIN_TRAIN_STUDIES = 1200\nN_FOLDS = 5\nMIN_POSITIVES = 25\n\nLABELS = [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\", \"Medial OA\",\n          \"Lateral OA\", \"PF OA\", \"Effusion\", \"Synovitis\", \"Baker's\",\n          \"Contusion\", \"Fracture\"]\n\ntry:\n    import pydicom\n    HAVE_PYDICOM = True\nexcept Exception as exc:\n    HAVE_PYDICOM = False\n    print(\"pydicom unavailable:\", exc)\n\nprint(\"pandas\", pd.__version__, \"| numpy\", np.__version__, \"| pydicom\", HAVE_PYDICOM)"},{"cell_type":"code","execution_count":null,"id":"c840d3015305","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"import matplotlib as mpl\nimport matplotlib.pyplot as plt\n\n# ---- Ink and surface -------------------------------------------------------\n# The same design system as the companion EDA notebook, so the two read as one\n# series. Three text tiers, every one of them at or above the WCAG AA 4.5:1 floor\n# against SURFACE. Ratios are computed below, not eyeballed: 19.2, 7.7, 5.4, 4.8.\n# GRID and AXIS carry no text and are deliberately left light; darkening a\n# gridline only adds noise, and a contrast floor does not apply to it.\nSURFACE = \"#fcfcfb\"   # chart surface\nINK = \"#0b0b0b\"       # titles, tick labels                     19.2:1\nINK_2 = \"#52514e\"     # value annotations, detail line           7.7:1\nMUTED = \"#6a6862\"     # axis labels, captions                    5.4:1\nGRID = \"#e1e0d9\"      # hairline gridlines, non-text\nAXIS = \"#c3c2b7\"      # spines and the random-baseline rule, non-text\n\n# ---- Marks -----------------------------------------------------------------\n# A bar does not have to clear a contrast floor; the sentence next to it does.\n# PRIMARY is 4.3:1 and ACCENT is 3.1:1, which is fine as a fill and unreadable\n# as a caption, so annotations that belong to an accent mark are drawn in\n# ACCENT_TEXT instead.\nPRIMARY = \"#2a78d6\"       # series 1\nACCENT = \"#eb6834\"        # series 2, the thing the chart is about\nACCENT_TEXT = \"#c04b1c\"   # accent annotations                   4.8:1\nPRIMARY_DEEP = \"#256abf\"  # fill behind SURFACE-coloured text    5.3:1\n\nmpl.rcParams.update({\n    \"figure.facecolor\": SURFACE,\n    \"figure.dpi\": 118,\n    \"savefig.facecolor\": SURFACE,\n    \"axes.facecolor\": SURFACE,\n    \"axes.edgecolor\": AXIS,\n    \"axes.linewidth\": 0.9,\n    \"axes.spines.top\": False,\n    \"axes.spines.right\": False,\n    \"axes.labelcolor\": MUTED,\n    \"axes.labelsize\": 12.5,\n    \"axes.titlesize\": 13.5,\n    \"axes.titlecolor\": INK,\n    \"axes.grid\": False,\n    \"grid.color\": GRID,\n    \"grid.linewidth\": 0.8,\n    \"grid.linestyle\": \"-\",\n    \"xtick.color\": MUTED,\n    \"ytick.color\": MUTED,\n    \"xtick.labelsize\": 12,\n    \"ytick.labelsize\": 12,\n    \"xtick.major.size\": 0,\n    \"ytick.major.size\": 0,\n    \"legend.frameon\": False,\n    \"legend.fontsize\": 12,\n    \"lines.linewidth\": 2.4,\n    \"lines.markersize\": 9,\n    \"font.family\": \"sans-serif\",\n    \"font.sans-serif\": [\"DejaVu Sans\", \"Segoe UI\", \"Helvetica\", \"Arial\"],\n    \"font.size\": 12.5,\n    \"text.color\": INK,\n})\n\n\ndef _rel_lum(colour):\n    \"\"\"WCAG relative luminance of a colour.\n\n    Args:\n        colour: Anything matplotlib can parse as a colour.\n\n    Returns:\n        Relative luminance in [0, 1].\n    \"\"\"\n    total = 0.0\n    for weight, chan in zip((0.2126, 0.7152, 0.0722), mpl.colors.to_rgb(colour)):\n        lin = chan / 12.92 if chan <= 0.03928 else ((chan + 0.055) / 1.055) ** 2.4\n        total += weight * lin\n    return total\n\n\ndef contrast(fg, bg):\n    \"\"\"WCAG 2.x contrast ratio between two colours.\n\n    Args:\n        fg: Foreground colour.\n        bg: Background colour.\n\n    Returns:\n        A ratio between 1.0 and 21.0. AA wants 4.5 for normal-size text.\n    \"\"\"\n    lo, hi = sorted((_rel_lum(fg), _rel_lum(bg)))\n    return (hi + 0.05) / (lo + 0.05)\n\n\ndef style(ax, grid_axis=\"x\"):\n    \"\"\"Apply the recessive grid and spine treatment to one axes.\n\n    Args:\n        ax: A matplotlib axes.\n        grid_axis: \"x\", \"y\" or \"none\". Gridlines run along this axis only,\n            behind the marks.\n\n    Returns:\n        The same axes, for chaining.\n    \"\"\"\n    for side in (\"top\", \"right\"):\n        ax.spines[side].set_visible(False)\n    for side in (\"left\", \"bottom\"):\n        ax.spines[side].set_color(AXIS)\n    ax.grid(False)\n    if grid_axis == \"x\":\n        ax.xaxis.grid(True, zorder=0)\n    elif grid_axis == \"y\":\n        ax.yaxis.grid(True, zorder=0)\n    ax.set_axisbelow(True)\n    ax.tick_params(length=0)\n    return ax\n\n\ndef _wrap(text, width_in, fontsize):\n    \"\"\"Wrap a header or caption string to the figure's own width.\n\n    The inline backend saves with a tight bounding box, so a text run wider than\n    the figure silently widens the exported PNG and strands the plot in one\n    corner. Wrapping to the figure width prevents that.\n\n    Args:\n        text: The string to wrap.\n        width_in: Figure width in inches.\n        fontsize: Point size the string will be drawn at.\n\n    Returns:\n        The string with newlines inserted.\n    \"\"\"\n    import textwrap\n    chars = max(24, int((width_in - 0.15) * 72.0 / (fontsize * 0.53)))\n    return \"\\n\".join(textwrap.wrap(text, chars)) if text else \"\"\n\n\ndef finish(fig, finding, detail=\"\", caption=\"\", headroom=0.0, footroom=0.0):\n    \"\"\"Add the finding title, optional detail line and caption, then show.\n\n    Titles state the finding, not the mechanism. Reserved space is computed in\n    inches from the wrapped line count, so a tall figure and a short one get the\n    same visual header band.\n\n    Args:\n        fig: The figure to close out.\n        finding: The headline. One sentence, stating what the chart shows.\n        detail: Optional second line in secondary ink.\n        caption: Optional footer in muted ink, for provenance and caveats.\n        headroom: Extra inches reserved above the axes, for figures whose panels\n            carry their own titles.\n        footroom: Extra inches reserved below the axes, for anything the tight\n            bounding box does not see.\n\n    Returns:\n        None.\n    \"\"\"\n    w, h = fig.get_size_inches()\n    finding = _wrap(finding, w, 17.5)\n    detail = _wrap(detail, w, 13.0)\n    caption = _wrap(caption, w, 11.25)\n    n_find = finding.count(\"\\n\") + 1\n    n_det = (detail.count(\"\\n\") + 1) if detail else 0\n    n_cap = (caption.count(\"\\n\") + 1) if caption else 0\n\n    head_in = 0.32 + 0.31 * n_find + (0.10 + 0.25 * n_det) + headroom\n    foot_in = ((0.18 + 0.21 * n_cap) if caption else 0.13) + footroom\n\n    # tight_layout silently does nothing on any figure whose gridspec was given an\n    # explicit hspace, which is every stacked pair below, and only emits a\n    # UserWarning when it declines. So the header band cannot be left to it: run it\n    # for the tick-label margins, then clamp top and bottom directly. The warning is\n    # suppressed here and only here, because the clamp below is the answer to it and\n    # a reader of the published notebook should not be shown a warning that has\n    # already been handled.\n    import warnings\n    try:\n        with warnings.catch_warnings():\n            warnings.simplefilter(\"ignore\", UserWarning)\n            fig.tight_layout(rect=(0.0, foot_in / h, 1.0, 1.0 - head_in / h))\n    except Exception:\n        pass\n    sp = fig.subplotpars\n    top = min(sp.top, 1.0 - head_in / h)\n    bottom = max(sp.bottom, foot_in / h)\n    if top > bottom + 0.05:\n        fig.subplots_adjust(top=top, bottom=bottom)\n\n    # When tight_layout declines, the axes box moves but its tick labels and axis\n    # label still hang below it, unmeasured, and the caption is pinned to the\n    # figure edge. Measure the real overhang and lift the axes if the two would\n    # meet. Only ever lifts, never lowers.\n    if caption:\n        try:\n            fig.canvas.draw()\n            rend = fig.canvas.get_renderer()\n            low = min(a.get_tightbbox(rend).y0 for a in fig.axes) / fig.bbox.height\n            shift = min(foot_in / h - low, 0.25)\n            sp = fig.subplotpars\n            if shift > 0.002 and sp.top > sp.bottom + shift + 0.05:\n                fig.subplots_adjust(bottom=sp.bottom + shift)\n        except Exception:\n            pass\n\n    fig.text(0.008, 1.0 - 0.20 / h, finding, ha=\"left\", va=\"top\",\n             fontsize=17.5, fontweight=\"bold\", color=INK, linespacing=1.25)\n    if detail:\n        fig.text(0.008, 1.0 - (0.25 + 0.31 * n_find) / h, detail, ha=\"left\",\n                 va=\"top\", fontsize=13, color=INK_2, linespacing=1.3)\n    if caption:\n        fig.text(0.008, 0.09 / h, caption, ha=\"left\", va=\"bottom\",\n                 fontsize=11.25, color=MUTED, linespacing=1.3)\n    plt.show()\n\n\ndef label_bars(ax, values, ypos, fmt=\"{:.2f}\", pad=0.012):\n    \"\"\"Write a value label just past the end of each horizontal bar.\n\n    Annotations wear secondary ink, never the series colour, so the number stays\n    readable whatever the bar is filled with.\n\n    Args:\n        ax: Axes holding the bars.\n        values: Sequence of bar lengths.\n        ypos: Sequence of bar centre positions.\n        fmt: Format string applied to each value.\n        pad: Horizontal offset in data units.\n\n    Returns:\n        None.\n    \"\"\"\n    for v, y in zip(values, ypos):\n        if v is None or (isinstance(v, float) and not np.isfinite(v)):\n            continue\n        ax.text(v + pad, y, fmt.format(v), va=\"center\", ha=\"left\",\n                fontsize=12, color=INK_2)\n\n\nprint(\"design system ready | text contrast on the chart surface \"\n      \"(WCAG AA floor 4.5):\",\n      \", \".join(\"{} {:.1f}\".format(n, contrast(c, SURFACE)) for n, c in\n                [(\"INK\", INK), (\"INK_2\", INK_2), (\"MUTED\", MUTED),\n                 (\"ACCENT_TEXT\", ACCENT_TEXT)]))"},{"cell_type":"code","execution_count":null,"id":"8e304cf53038","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"def detect_input_root():\n    \"\"\"Find the directory holding train.csv, without hardcoding a mount path.\n\n    Returns:\n        Path to the competition input root.\n\n    Raises:\n        FileNotFoundError: If no candidate directory contains train.csv.\n    \"\"\"\n    # Kaggle does not always mount a competition at /kaggle/input/<slug>. For this one it\n    # is /kaggle/input/competitions/<slug>, one level deeper, which a shallow scan of\n    # /kaggle/input misses: it finds the `competitions` directory and no train.csv inside\n    # it, then gives up. Version 1 of this notebook died exactly that way. So both shapes\n    # are named explicitly and the fallback walks two levels rather than one.\n    candidates = [\n        Path(\"/kaggle/input/rsna-knee-abnormality-detection\"),\n        Path(\"/kaggle/input/competitions/rsna-knee-abnormality-detection\"),\n        Path(\"/mnt/data/Github/Kaggle/Competitions/rsna_knee_abnormality_detection\"\n             \"/data/kaggle_mirror/rsna-knee-abnormality-detection\"),\n        Path(\"/mnt/data/Github/Kaggle/Competitions/rsna_knee_abnormality_detection/data/raw\"),\n        Path(\"data/raw\"),\n        Path(\".\"),\n    ]\n    for c in candidates:\n        try:\n            if (c / \"train.csv\").is_file():\n                return c\n        except OSError:\n            continue\n\n    kin = Path(\"/kaggle/input\")\n    seen = []\n    if kin.is_dir():\n        for depth1 in sorted(p for p in kin.iterdir() if p.is_dir()):\n            seen.append(str(depth1))\n            try:\n                if (depth1 / \"train.csv\").is_file():\n                    return depth1\n                for depth2 in sorted(p for p in depth1.iterdir() if p.is_dir()):\n                    seen.append(str(depth2))\n                    if (depth2 / \"train.csv\").is_file():\n                        return depth2\n            except OSError:\n                continue\n\n    # Name what was checked AND what is actually mounted. A rerun failure is only\n    # debuggable from the log, and \"not found\" on its own says nothing.\n    raise FileNotFoundError(\n        \"no input root containing train.csv was found. Checked: \"\n        + \", \".join(str(c) for c in candidates)\n        + \". Directories present under /kaggle/input: \"\n        + (\", \".join(seen) if seen else \"none\")\n    )\n\n\ndef find_subdir(root, names):\n    \"\"\"Return the first existing subdirectory of root among names, else None.\n\n    Args:\n        root: Directory to look inside.\n        names: Candidate directory names, tried in order.\n\n    Returns:\n        A Path, or None when none of the names is a directory.\n    \"\"\"\n    for n in names:\n        p = root / n\n        try:\n            if p.is_dir():\n                return p\n        except OSError:\n            continue\n    return None\n\n\nROOT = detect_input_root()\nTRAIN_DCM = find_subdir(ROOT, [\"train_series\", \"train_images\"])\nTEST_DCM = find_subdir(ROOT, [\"test_series\", \"test_images\"])\nWORK = Path(\"/kaggle/working\") if Path(\"/kaggle/working\").is_dir() else Path(\".\")\n\ntrain = pd.read_csv(ROOT / \"train.csv\")\ntest = pd.read_csv(ROOT / \"test.csv\")\n\n\ndef read_optional(path, columns):\n    \"\"\"Read a CSV that may be absent, returning an empty typed frame instead.\n\n    Args:\n        path: CSV path.\n        columns: Column names for the empty fallback frame.\n\n    Returns:\n        A DataFrame, empty with the given columns when the file is missing.\n    \"\"\"\n    try:\n        if path.is_file():\n            return pd.read_csv(path)\n    except Exception as exc:\n        print(\"could not read\", path.name, \":\", exc)\n    return pd.DataFrame(columns=columns)\n\n\nSERIES_COLS = [\"StudyInstanceUID\", \"SeriesInstanceUID\", \"Fluid_Sensitive\",\n               \"Fat_Suppression\", \"Anatomical_Plane\"]\ntrain_series = read_optional(ROOT / \"train_series.csv\", SERIES_COLS)\ntest_series = read_optional(ROOT / \"test_series.csv\", SERIES_COLS)\nsample_sub = read_optional(ROOT / \"sample_submission.csv\", [\"StudyInstanceUID\"] + LABELS)\n\nprint(\"input root      :\", ROOT)\nprint(\"train dicom dir :\", TRAIN_DCM)\nprint(\"test dicom dir  :\", TEST_DCM)\nprint(\"writing to      :\", WORK)\nprint()\nprint(\"train studies   :\", len(train), \"| train series :\", len(train_series))\nprint(\"test studies    :\", len(test), \"| test series  :\", len(test_series))\n\ngold_mask = train[LABELS].notna().all(axis=1) if set(LABELS).issubset(train.columns) else pd.Series(False, index=train.index)\nN_GOLD = int(gold_mask.sum())\nprint(\"gold-labelled   :\", N_GOLD, \"of\", len(train),\n      \"({:.2f}%)\".format(100.0 * N_GOLD / max(1, len(train))))"},{"cell_type":"markdown","id":"851fead069a6","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">3</span><span style=\"vertical-align:middle;\">Labels from the reports</span></h2></div>\n\nThe reports are genuinely multilingual. English, Spanish, French, Dutch, German,\nTurkish, Greek and Bulgarian all appear, and a Croatian report turned up too. Any\nkeyword extractor has to clear that bar in every language present, and negation is\nwhere such extractors fall over: *no meniscal tear* and *meniscal tear* differ by one\ntoken in every one of them.\n\nThe extractor below works clause by clause. It folds case and diacritics, splits on\nsentence punctuation and line breaks, then applies two kinds of rule:\n\n*Paired rules* for the four structures that need an anatomy word and a pathology\nword together. ACL, MCL, medial meniscus and lateral meniscus each carry a list of\nspecific surface forms across languages, plus a fallback that pairs a generic organ\nterm with a side qualifier so `μηνισκ` beside `εσω` reads as medial meniscus.\n\n*Direct rules* for the five findings where the term itself is the finding: effusion,\nsynovitis, Baker's cyst, contusion, fracture.\n\nOsteoarthritis is handled separately. A cartilage or osteophyte term fires the\ncompartment named nearest to it, with `tricompartmental` and `gonarthrosis` firing\nall three. Phrases like *medial patellar facet* are masked out before the medial and\nlateral test runs, so a patellofemoral sentence does not silently claim the medial\ncompartment as well.\n\nThe negation guard is a word-boundary regex over just under sixty cues, spanning\nevery language the corpus turned up, checked in a 70-character window before the\nfinding term and a 45-character window after it. Both directions matter: Turkish,\nGreek and Bulgarian put the negation after the noun. The cell below prints the cue\ncount it actually loaded."},{"cell_type":"code","execution_count":null,"id":"b47baec8d2a7","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"# --------------------------------------------------------------------------- #\n# Multilingual rule-based label extractor.\n# Weak on purpose. This is the artifact an LLM extraction has to beat.\n# --------------------------------------------------------------------------- #\n\n_CHAR_MAP = str.maketrans({\"ı\": \"i\", \"İ\": \"i\", \"ß\": \"ss\",\n                           \"ø\": \"o\", \"Ø\": \"o\",\n                           \"đ\": \"d\", \"Đ\": \"d\"})\n\nTEAR = [\n    \"tear\", \"torn\", \"tearing\", \"rupture\", \"ruptured\", \"disruption\", \"discontinuity\",\n    \"rotura\", \"ruptura\", \"desgarro\", \"roto\", \"rota\",\n    \"dechirure\", \"dechire\", \"lesion meniscale\",\n    \"scheur\", \"ruptuur\", \"gescheurd\",\n    \"riss\", \"rissbildung\", \"ruptur\", \"zerreissung\", \"einriss\",\n    \"yirtik\", \"yirtig\", \"kopma\", \"butunluk kaybi\",\n    \"ρηξη\", \"ρηξις\", \"ρηγμα\",\n    \"руптура\", \"разкъсв\",\n    \"разрив\", \"puknuce\", \"prekid\",\n]\nSPRAIN = [\n    \"sprain\", \"esguince\", \"entorse\", \"verstauchung\", \"verstuiking\",\n    \"distorsiyon\", \"burkulma\", \"διαστρεμμα\",\n    \"навяхван\",\n    \"injury\", \"lesion\", \"letsel\", \"verletzung\", \"zedelenme\",\n]\nOA_FIND = [\n    \"osteoarthritis\", \"osteoarthrosis\", \"arthrosis\", \"arthritic\",\n    \"artrosis\", \"artrose\", \"arthrose\", \"gonartrose\", \"gonartroz\", \"gonarthrose\",\n    \"gonartro\", \"artroz\",\n    \"chondrosis\", \"chondral\", \"chondropathy\", \"chondropathie\", \"chondropatie\",\n    \"chondromalacia\", \"kondromalazi\", \"condropatia\", \"condral\", \"chondropatia\",\n    \"kraakbeenlijden\", \"kraakbeenverlies\", \"kraakbeenschade\",\n    \"knorpelschaden\", \"knorpeldefekt\", \"knorpelverlust\", \"knorpelbelag\",\n    \"cartilage loss\", \"cartilage thinning\", \"cartilage fissuring\",\n    \"cartilage defect\", \"cartilage fissure\", \"chondral defect\", \"chondral loss\",\n    \"osteophyt\", \"osteofit\", \"osteofyt\", \"osteophyte\", \"spurring\", \"osteofyte\",\n    \"joint space narrowing\", \"gelenkspaltverschmalerung\",\n    \"kikirdak kaybi\", \"kikirdak incelme\", \"kondral\",\n    \"ulcera condral\", \"hrskavice\", \"denudacija\",\n    \"χονδροπαθ\", \"οστεοαρθρ\",\n    \"οστεοφυτ\", \"αρθριτ\",\n    \"артроз\", \"хондропат\",\n    \"остеофит\", \"хрущялн\",\n]\nMEDIAL_Q = [\"medial\", \"interno\", \"interna\", \"interne\", \"mediaal\", \"mediale\",\n            \"innen\", \"medyal\", \"medijaln\", \"εσω\",\n            \"медиал\", \"вътреш\"]\nLATERAL_Q = [\"lateral\", \"externo\", \"externa\", \"externe\", \"buiten\", \"aussen\",\n             \"lateraln\", \"εξω\", \"латерал\",\n             \"външ\"]\nPF_Q = [\"patellofemoral\", \"patelofemoral\", \"femoropatellar\", \"femoropatelar\",\n        \"femoropatellair\", \"retropatellar\", \"retrorotulian\", \"patellar facet\",\n        \"trochlea\", \"troclea\", \"trochlee\", \"rotula\", \"rotulian\", \"patella\",\n        \"patellaire\", \"patellar\", \"diz kapagi\", \"patelofemoraln\",\n        \"επιγονατιδ\", \"τροχιλ\",\n        \"пател\", \"ретропател\"]\nTRICOMP = [\"tricompartmental\", \"tricompartimental\", \"three compartments\",\n           \"three compartmens\", \"all compartments\", \"gonartrose\", \"gonartroz\",\n           \"gonarthrose\", \"gonartrosis\", \"gonartro\", \"pangonartro\"]\n\nSPECIFIC = {\n    \"ACL\": [\n        \"acl\", \"anterior cruciate\", \"lca\", \"ligamento cruzado anterior\",\n        \"ligament croise anterieur\", \"croise anterieur\",\n        \"voorste kruisband\", \"vkb\", \"vorderes kreuzband\", \"vorderen kreuzband\",\n        \"vordere kreuzband\", \"on capraz bag\", \"anterior capraz bag\",\n        \"προσθιο χιαστ\",\n        \"προσθιου χιαστ\",\n        \"προσθιος χιαστ\",\n        \"предна кръстна\",\n        \"предната кръстна\",\n        \"prednji ukrizeni\",\n    ],\n    \"MCL\": [\n        \"mcl\", \"medial collateral\", \"lcm\", \"ligamento colateral medial\",\n        \"ligamento colateral interno\", \"ligament collateral medial\",\n        \"collateral medial\", \"mediale collaterale\", \"mediaal collateraal\",\n        \"innenband\", \"mediales kollateralband\", \"medialen kollateralband\",\n        \"mediale kollateralband\", \"medyal kollateral\", \"ic yan bag\",\n        \"εσω πλαγιο\",\n        \"медиален колатерален\",\n        \"вътрешна колатерална\",\n    ],\n    \"Medial Meniscus\": [\n        \"medial meniscus\", \"meniscus medialis\", \"menisco medial\", \"menisco interno\",\n        \"mediale meniscus\", \"binnenmeniscus\", \"innenmeniskus\", \"meniskus medialis\",\n        \"menisque interne\", \"menisque medial\", \"medial menisk\", \"medyal menisk\",\n        \"εσω μηνισκ\",\n        \"медиалния менискус\",\n        \"медиален мениск\",\n        \"вътрешния мениск\",\n        \"medial and lateral menisc\", \"medial ve lateral menisk\",\n        \"menisco medial y lateral\", \"menisco interno y externo\",\n        \"menisco interno y lateral\", \"mediale en laterale meniscus\",\n        \"innen und aussenmeniskus\", \"medijalnog meniskusa\",\n    ],\n    \"Lateral Meniscus\": [\n        \"lateral meniscus\", \"meniscus lateralis\", \"menisco lateral\", \"menisco externo\",\n        \"laterale meniscus\", \"buitenmeniscus\", \"aussenmeniskus\", \"meniskus lateralis\",\n        \"menisque externe\", \"menisque lateral\", \"lateral menisk\",\n        \"εξω μηνισκ\",\n        \"латералния менискус\",\n        \"латерален мениск\",\n        \"външния мениск\",\n        \"medial and lateral menisc\", \"medial ve lateral menisk\",\n        \"menisco medial y lateral\", \"menisco interno y externo\",\n        \"menisco interno y lateral\", \"mediale en laterale meniscus\",\n        \"innen und aussenmeniskus\", \"lateralnog meniskusa\",\n    ],\n}\n\nGENERIC = {\n    \"ACL\": [\"cruciate\", \"cruzado\", \"croise\", \"kruisband\", \"kreuzband\",\n            \"χιαστ\", \"кръстн\",\n            \"capraz bag\", \"ukrizen\"],\n    \"MCL\": [\"collateral\", \"colateral\", \"kollateral\", \"collaterale\",\n            \"πλαγι\", \"колатерал\",\n            \"yan bag\"],\n    \"Medial Meniscus\": [\"menisc\", \"menisk\", \"μηνισκ\",\n                        \"мениск\"],\n    \"Lateral Meniscus\": [\"menisc\", \"menisk\", \"μηνισκ\",\n                         \"мениск\"],\n}\nSIDE_Q = {\n    \"ACL\": [\"anterior\", \"anterieur\", \"voorste\", \"vorder\",\n            \"προσθι\", \"предн\", \"prednj\"],\n    \"MCL\": MEDIAL_Q,\n    \"Medial Meniscus\": MEDIAL_Q,\n    \"Lateral Meniscus\": LATERAL_Q,\n}\n\nDIRECT = {\n    \"Effusion\": [\n        \"effusion\", \"joint fluid\", \"hemarthrosis\", \"haemarthrosis\", \"hydrops\",\n        \"derrame\", \"epanchement\", \"gewrichtsvocht\", \"vocht in het gewricht\",\n        \"erguss\", \"gelenkerguss\", \"gelenkserguss\",\n        \"eklem ici sivi\", \"sivi artisi\", \"efuzyon\", \"eklem mesafesinde sivi\",\n        \"eklem ici serbest sivi\", \"eklem sivisi\",\n        \"ενδαρθρικ\",\n        \"αρθρικο υγρο\",\n        \"συλλογη υγρου\",\n        \"ставен излив\",\n        \"излив\", \"хидропс\",\n        \"zglobni izljev\", \"izljev\",\n    ],\n    \"Synovitis\": [\n        \"synovitis\", \"synovial thickening\", \"synovial hypertrophy\",\n        \"thickened synovial\", \"hypertrophy of the synovium\",\n        \"synovial proliferation\", \"proliferation of the synovium\",\n        \"sinovitis\", \"synovite\", \"synovitiden\", \"synovialitis\",\n        \"verdikking van het synovium\", \"verdikkingen van het synovium\",\n        \"synoviale verdikking\", \"synovialisverdickung\", \"synovialverdickung\",\n        \"sinovit\", \"hoffitis\", \"sinovyal kalinlasma\", \"sinovyal proliferasyon\",\n        \"συνοβιτ\", \"υμενιτ\",\n        \"синовит\", \"sinovij\", \"sinovije\",\n    ],\n    \"Baker's\": [\n        \"baker\", \"popliteal cyst\", \"poplitealcyst\", \"popliteal cysts\",\n        \"quiste popliteo\", \"quistes popliteos\", \"quiste de baker\",\n        \"kyste de baker\", \"kyste poplite\", \"popliteale cyste\", \"popliteale cyst\",\n        \"bakercyste\", \"bakerzyste\", \"poplitealzyste\", \"popliteazyste\",\n        \"baker kisti\", \"popliteal kist\",\n        \"κυστη baker\", \"κυστη του baker\",\n        \"киста на бейкър\",\n        \"бейкърова киста\",\n        \"poplitealna cista\",\n    ],\n    \"Contusion\": [\n        \"contusion\", \"contusiones\", \"contusie\", \"kontusyon\", \"kontuzyon\",\n        \"bone bruise\", \"bone bruising\", \"knochenprellung\", \"prellung\",\n        \"botcontusie\", \"μωλωπ\", \"θλαση\",\n        \"контузи\", \"kontuzij\", \"nagnjecen\",\n    ],\n    \"Fracture\": [\n        \"fracture\", \"fractur\", \"fractura\", \"fraktur\", \"fractuur\", \"breuk\",\n        \"kirik\", \"kirig\", \"καταγμα\",\n        \"καταγματ\",\n        \"фрактура\", \"счупван\",\n        \"avulsion\", \"avulsie\", \"avulsiyon\", \"prijelom\",\n    ],\n}\n\n# Negation cues, word-boundary anchored. A bare \"no \" substring matches inside the\n# Spanish word \"cuerno\", which silently killed most Spanish positives before this\n# was anchored.\nNEG_SRC = [\n    r\"\\bno\\b\", r\"\\bnot\\b\", r\"\\bnon\\b\", r\"\\bwithout\\b\", r\"\\babsen\\w*\",\n    r\"\\bnegative for\\b\", r\"\\bunremarkable\\b\", r\"\\bintact\\w*\", r\"\\bnormal\\w*\",\n    r\"\\bpreserved\\b\", r\"\\bfree of\\b\", r\"\\bexcluded\\b\",\n    r\"\\bno hay\\b\", r\"\\bsin\\b\", r\"\\bausen\\w*\", r\"\\bno se\\b\", r\"\\bconservad\\w*\",\n    r\"\\bintegr\\w*\", r\"\\bindemne\\b\",\n    r\"\\bpas de\\b\", r\"\\baucun\\w*\", r\"\\bsans\\b\",\n    r\"\\bgeen\\b\", r\"\\bzonder\\b\", r\"\\bonopvallend\\w*\", r\"\\bnormaal\\b\",\n    r\"\\bvrij van\\b\", r\"\\bintacte\\b\",\n    r\"\\bkein\\w*\", r\"\\bohne\\b\", r\"\\bunauffallig\\w*\", r\"\\bregelrecht\\w*\",\n    r\"\\bnicht\\b\", r\"\\bintakt\\w*\",\n    r\"\\byok\\w*\", r\"\\bizlenmemis\\w*\", r\"\\bizlenmedi\\w*\", r\"\\bsaptanmamis\\w*\",\n    r\"\\bgozlenmemis\\w*\", r\"\\bgorulmemis\\w*\", r\"\\bkorunmus\\w*\",\n    r\"\\bmevcut degil\\b\", r\"\\bnormaldir\\b\", r\"\\bdogaldir\\b\",\n    r\"\\bδεν\\b\", r\"\\bχωρις\\b\",\n    r\"\\bφυσιολογικ\\w*\",\n    r\"\\bακεραι\\w*\",\n    r\"\\bбез\\b\", r\"\\bняма\\b\",\n    r\"\\bне се\\b\", r\"\\bнормал\\w*\",\n    r\"\\bзапазен\\w*\",\n    r\"\\bсъхранен\\w*\",\n    r\"\\bsenza\\b\", r\"\\bsem\\b\", r\"\\bnao\\b\", r\"\\bbez\\b\", r\"\\buredn\\w*\",\n]\nNEG_RE = re.compile(\"|\".join(NEG_SRC))\n\nNEG_BACK = 70\nNEG_FWD = 45\nQUAL_WINDOW = 90\nPAIR_WINDOW = 160\nCLAUSE_SPLIT = re.compile(r\"[.;:\\n\\r•·]+|\\s-\\s|\\s>\\s|\\s\\*\\s\")\nPF_MASK = re.compile(r\"(medial|lateral)\\s+(patellar|patella|facet|trochlea|trochlear|retinac)\\w*\")\n\n\ndef fold_text(text):\n    \"\"\"Lowercase, normalise script-specific letters, and strip diacritics.\n\n    Turkish dotless i, the German sharp s and the Croatian barred d have no\n    combining-mark decomposition, so they are mapped explicitly before NFKD.\n\n    Args:\n        text: Any report text.\n\n    Returns:\n        A lowercase, accent-free string safe for substring matching.\n    \"\"\"\n    t = str(text).lower().translate(_CHAR_MAP)\n    t = unicodedata.normalize(\"NFKD\", t)\n    return \"\".join(c for c in t if not unicodedata.combining(c))\n\n\ndef _find_any(clause, terms):\n    \"\"\"Return the earliest index at which any term occurs, or -1.\n\n    Args:\n        clause: Folded clause text.\n        terms: Iterable of folded surface forms.\n\n    Returns:\n        Character index of the earliest hit, or -1 when none matches.\n    \"\"\"\n    best = -1\n    for t in terms:\n        i = clause.find(t)\n        if i >= 0 and (best < 0 or i < best):\n            best = i\n    return best\n\n\ndef _negated(clause, pos, span):\n    \"\"\"Test whether a negation cue sits near a matched finding term.\n\n    Args:\n        clause: Folded clause text.\n        pos: Start index of the finding term.\n        span: Length of the finding term.\n\n    Returns:\n        True when a cue appears in the preceding or following window.\n    \"\"\"\n    back = clause[max(0, pos - NEG_BACK):pos]\n    fwd = clause[pos + span:pos + span + NEG_FWD]\n    return bool(NEG_RE.search(back) or NEG_RE.search(fwd))\n\n\ndef _fire_direct(clause, terms):\n    \"\"\"Test a direct finding term, retrying later occurrences past a negation.\n\n    Args:\n        clause: Folded clause text.\n        terms: Surface forms that are themselves the finding.\n\n    Returns:\n        True when at least one occurrence is unnegated.\n    \"\"\"\n    for t in terms:\n        start = 0\n        while True:\n            i = clause.find(t, start)\n            if i < 0:\n                break\n            if not _negated(clause, i, len(t)):\n                return True\n            start = i + 1\n    return False\n\n\ndef _anatomy_index(clause, label):\n    \"\"\"Locate the anatomy mention for a paired label.\n\n    Falls back to a generic organ term paired with a side qualifier, which is what\n    catches Greek and Bulgarian phrasing that names the compartment rather than the\n    structure.\n\n    Args:\n        clause: Folded clause text.\n        label: One of the four paired label names.\n\n    Returns:\n        Character index of the anatomy mention, or -1.\n    \"\"\"\n    i = _find_any(clause, SPECIFIC[label])\n    if i >= 0:\n        return i\n    g = _find_any(clause, GENERIC[label])\n    if g >= 0 and _find_any(clause, SIDE_Q[label]) >= 0:\n        return g\n    return -1\n\n\ndef extract_labels(report):\n    \"\"\"Extract twelve binary findings from one free-text radiology report.\n\n    Args:\n        report: Report text in any of the languages present in the corpus.\n\n    Returns:\n        Dict mapping each of the twelve label names to 0 or 1.\n    \"\"\"\n    out = {lab: 0 for lab in LABELS}\n    if not isinstance(report, str) or not report.strip():\n        return out\n    text = fold_text(report)\n    for raw in CLAUSE_SPLIT.split(text):\n        clause = raw.strip()\n        if len(clause) < 3:\n            continue\n        for lab in (\"Effusion\", \"Synovitis\", \"Baker's\", \"Contusion\", \"Fracture\"):\n            if not out[lab] and _fire_direct(clause, DIRECT[lab]):\n                out[lab] = 1\n        for lab in SPECIFIC:\n            if out[lab]:\n                continue\n            ai = _anatomy_index(clause, lab)\n            if ai < 0:\n                continue\n            finds = TEAR + SPRAIN if lab in (\"ACL\", \"MCL\") else TEAR\n            for t in finds:\n                fi = clause.find(t)\n                if fi < 0 or abs(fi - ai) > PAIR_WINDOW:\n                    continue\n                if not _negated(clause, fi, len(t)):\n                    out[lab] = 1\n                    break\n        oi = _find_any(clause, OA_FIND)\n        if oi >= 0 and not _negated(clause, oi, 8):\n            if _find_any(clause, TRICOMP) >= 0:\n                out[\"Medial OA\"] = 1\n                out[\"Lateral OA\"] = 1\n                out[\"PF OA\"] = 1\n            win = clause[max(0, oi - QUAL_WINDOW):oi + QUAL_WINDOW]\n            if _find_any(win, PF_Q) >= 0:\n                out[\"PF OA\"] = 1\n            masked = PF_MASK.sub(\" \", win)\n            if _find_any(masked, MEDIAL_Q) >= 0:\n                out[\"Medial OA\"] = 1\n            if _find_any(masked, LATERAL_Q) >= 0:\n                out[\"Lateral OA\"] = 1\n    return out\n\n\nprint(\"extractor ready:\", len(SPECIFIC), \"paired labels,\",\n      len(DIRECT), \"direct labels,\", len(NEG_SRC), \"negation cues\")"},{"cell_type":"code","execution_count":null,"id":"438660df48f7","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"t0 = time.time()\nreport_col = \"Report\" if \"Report\" in train.columns else None\nif report_col is None:\n    raise KeyError(\"train.csv has no Report column; the label extractor cannot run\")\n\nrule = pd.DataFrame([extract_labels(r) for r in train[report_col]], index=train.index)\nrule = rule[LABELS].astype(np.int8)\nprint(\"extracted {} reports in {:.1f}s\".format(len(rule), time.time() - t0))\n\nprev = pd.DataFrame({\n    \"rule_positives\": rule.sum().astype(int),\n    \"rule_prevalence_%\": (rule.mean() * 100).round(1),\n})\nprint()\nprint(prev.to_string())"},{"cell_type":"markdown","id":"1b57c3224974","metadata":{},"source":"### Auditing the extractor against the gold studies\n\nThe handful of studies that carry real labels are the only ground truth available.\nThey are positive-enriched rather than a random sample, and there are far too few of\nthem to rank models. What they can support is exactly this: checking, per label,\nwhether the extractor agrees with a radiologist.\n\nThree numbers per label, all computed below. Agreement is the fraction of gold\nstudies where the extractor and the label agree. Precision is the fraction of the\nextractor's positives that are real. Recall is the fraction of real positives the\nextractor found.\n\nOne thing this notebook deliberately does not do is tune the rule lists against these\nstudies. Bugs were fixed against them, such as a bare `no ` cue that was matching\ninside the Spanish word `cuerno`, and languages were added when a report turned out\nto be in a script the lists did not cover. Thresholds were not moved to chase the\nagreement number. Fitting a rule set to a few dozen studies produces a number that\ndoes not survive contact with the leaderboard."},{"cell_type":"code","execution_count":null,"id":"19fbaad99919","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"gold = train.loc[gold_mask, LABELS].astype(int)\npred_gold = rule.loc[gold_mask]\n\nrows = []\nfor lab in LABELS:\n    yt = gold[lab].values\n    yp = pred_gold[lab].values\n    tp = int(((yt == 1) & (yp == 1)).sum())\n    fp = int(((yt == 0) & (yp == 1)).sum())\n    fn = int(((yt == 1) & (yp == 0)).sum())\n    rows.append({\n        \"label\": lab,\n        \"gold_pos\": int(yt.sum()),\n        \"rule_pos\": int(yp.sum()),\n        \"agreement\": float((yt == yp).mean()),\n        \"precision\": tp / (tp + fp) if (tp + fp) else np.nan,\n        \"recall\": tp / (tp + fn) if (tp + fn) else np.nan,\n    })\naudit = pd.DataFrame(rows)\n\nprint(\"Extractor vs gold labels, n = {} studies\".format(len(gold)))\nprint(audit.round(3).to_string(index=False))\nprint()\nprint(\"mean agreement {:.3f} | mean precision {:.3f} | mean recall {:.3f}\".format(\n    audit[\"agreement\"].mean(), audit[\"precision\"].mean(), audit[\"recall\"].mean()))"},{"cell_type":"code","execution_count":null,"id":"2edeb6222bb1","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"# Two panels, stacked, sharing one label order so a finding can be tracked straight\n# down the figure. Agreement gets its own panel because it is a different kind of\n# number from the other two: it counts negatives as successes, so a label the\n# extractor never fires on still scores well on it.\norder = audit.sort_values(\"agreement\").reset_index(drop=True)\nnames = order[\"label\"].tolist()\nypos = np.arange(len(order))\n\nfig = plt.figure(figsize=(14.0, 12.6))\ngs = fig.add_gridspec(2, 1, height_ratios=[1.0, 1.22], hspace=0.30)\n\nax0 = fig.add_subplot(gs[0, 0])\nagree = order[\"agreement\"].values.astype(float)\nax0.barh(ypos, np.nan_to_num(agree), height=0.68, color=PRIMARY, edgecolor=\"none\",\n         zorder=3)\nlabel_bars(ax0, agree, ypos, fmt=\"{:.2f}\", pad=0.012)\nax0.set_xlim(0, 1.16)\nax0.set_xticks([0, 0.25, 0.5, 0.75, 1.0])\nax0.set_yticks(ypos)\nax0.set_yticklabels(names, fontsize=12.5, color=INK)\nax0.set_ylim(-0.7, len(order) - 0.3)\nax0.set_xlabel(\"fraction of the {} gold studies where extractor and label agree\"\n               .format(len(gold)))\nax0.set_title(\"Agreement\", loc=\"left\", fontsize=13.5, color=INK, pad=10)\nstyle(ax0, \"x\")\nax0.spines[\"left\"].set_visible(False)\n\nax1 = fig.add_subplot(gs[1, 0])\nh, gap = 0.34, 0.05  # the gap keeps a surface line between the two bars of a group\nprec = order[\"precision\"].values.astype(float)\nrec = order[\"recall\"].values.astype(float)\nax1.barh(ypos + (h + gap) / 2, np.nan_to_num(prec), height=h, color=PRIMARY,\n         edgecolor=\"none\", zorder=3, label=\"precision: of what it flagged, how much was real\")\nax1.barh(ypos - (h + gap) / 2, np.nan_to_num(rec), height=h, color=ACCENT,\n         edgecolor=\"none\", zorder=3, label=\"recall: of the real positives, how many it found\")\nlabel_bars(ax1, prec, ypos + (h + gap) / 2, fmt=\"{:.2f}\", pad=0.012)\nlabel_bars(ax1, rec, ypos - (h + gap) / 2, fmt=\"{:.2f}\", pad=0.012)\nax1.set_xlim(0, 1.16)\nax1.set_xticks([0, 0.25, 0.5, 0.75, 1.0])\nax1.set_yticks(ypos)\nax1.set_yticklabels(names, fontsize=12.5, color=INK)\nax1.set_ylim(-0.9, len(order) - 0.3)\nax1.set_xlabel(\"rate\")\nax1.set_title(\"Precision and recall\", loc=\"left\", fontsize=13.5, color=INK, pad=10)\nstyle(ax1, \"x\")\nax1.spines[\"left\"].set_visible(False)\n# Legend below the axis so it can never sit on top of a bar.\nax1.legend(loc=\"upper center\", bbox_to_anchor=(0.5, -0.085), ncol=2,\n           labelcolor=INK_2)\n\nfinish(fig, \"The keyword extractor is mediocre, and unevenly so\",\n       detail=\"Mean agreement {:.2f}, mean precision {:.2f}, mean recall {:.2f} \"\n              \"across the twelve findings.\".format(\n                  audit[\"agreement\"].mean(), audit[\"precision\"].mean(),\n                  audit[\"recall\"].mean()),\n       headroom=0.34, footroom=0.75,\n       caption=\"Weak by design. This is the floor the LLM extraction has to beat, per \"\n               \"label, on these same {} studies. A blank bar means the extractor never \"\n               \"fired on that finding, so the rate is undefined.\".format(len(gold)))"},{"cell_type":"markdown","id":"3ac1a106ee76","metadata":{},"source":"### Reading the audit honestly\n\nThe extractor is mediocre and the table says so. Recall on synovitis is poor because\nradiologists describe it a dozen ways and only some of them are a keyword. Precision\non lateral OA is poor because a cartilage sentence mentioning the lateral condyle is\nnot the same as lateral compartment osteoarthritis, and because the graders appear to\nrequire more than mild fissuring before they call it disease. Effusion over-fires\nbecause every trace of fluid gets a mention.\n\nSeveral gold studies disagree with their own report text. One report states\n*complete tear of the anterior cruciate ligament* and carries ACL negative; another\nstates *no acute fracture* and carries fracture positive. Whatever the labelling\nprotocol was, it is not a literal reading of the report, so a perfect extractor would\nnot reach perfect agreement here.\n\nThese are the numbers the LLM extraction has to beat, per label, on the same studies."},{"cell_type":"markdown","id":"dedf734e229b","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">4</span><span style=\"vertical-align:middle;\">Features the test set actually carries</span></h2></div>\n\nTest studies ship `test_series.csv` and the DICOM files. Nothing else. So the feature\nset is the series manifest plus DICOM header tags, and the same function builds it for\ntrain and test. A train/test feature mismatch is the most common way a code\ncompetition notebook dies at rerun time, so there is exactly one feature function and\none fill policy, applied identically to both.\n\nHeaders are read with `stop_before_pixels=True`. That matters for more than speed:\nthe corpus mixes uncompressed, JPEG Lossless, JPEG 2000 and implicit VR transfer\nsyntaxes, and stopping before the pixel data means no decoder is needed and no\ntransfer syntax can fail the read.\n\nReads are bounded. At most **2 slices per series** and at most **16 series per study**\nare opened, which is plenty for header tags that are constant within a series. Slice\ncounts come from the directory listing, which is one cheap call, not from opening\nevery file. Every read is wrapped in try/except and failures are counted rather than\nraised.\n\nOne guard is worth naming. If header reads succeed on one side and fail on the other,\nthe model fits a feature block at training time that does not exist at prediction\ntime, and every prediction collapses toward the same row. That was reproduced on\npurpose during development. So the header coverage rate is measured on both sides, and\nif the gap exceeds 25 points the entire header block is dropped and the model falls\nback to the series manifest, which both sides always have."},{"cell_type":"code","execution_count":null,"id":"439e0f59171c","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"PLANES = [\"Sagittal\", \"Coronal\", \"Axial\"]\n\n# One column order, declared once. The function has two branches, one for a real\n# manifest and one for an absent or empty one, and both return through this list.\n# Without it the branches emit the same column SET in a different ORDER, which\n# survives every set-based check and only bites whoever later feeds the assembled\n# frame to a model positionally.\nSERIES_FEATURE_COLUMNS = (\n    [\"n_series\"]\n    + [\"n_plane_\" + p.lower() for p in PLANES + [\"Other\"]]\n    + [\"frac_plane_\" + p.lower() for p in PLANES]\n    + [\"n_fluid_sensitive\", \"n_fat_suppression\", \"frac_fluid_sensitive\",\n       \"frac_fat_suppression\", \"fluid_equals_fat\"]\n)\n\n\ndef series_features(study_ids, series_df):\n    \"\"\"Build per-study features from the series manifest.\n\n    Args:\n        study_ids: Sequence of StudyInstanceUID values to produce rows for.\n        series_df: Frame with StudyInstanceUID, SeriesInstanceUID,\n            Fluid_Sensitive, Fat_Suppression and Anatomical_Plane columns.\n\n    Returns:\n        DataFrame indexed by study id, one row per requested study, columns in\n        SERIES_FEATURE_COLUMNS order, zero-filled where a study has no manifest\n        rows at all.\n    \"\"\"\n    idx = pd.Index(study_ids, name=\"StudyInstanceUID\")\n    out = pd.DataFrame(index=idx)\n    if series_df is None or len(series_df) == 0:\n        for c in SERIES_FEATURE_COLUMNS:\n            out[c] = 1 if c == \"fluid_equals_fat\" else (0.0 if c.startswith(\"frac_\") else 0)\n        return out[SERIES_FEATURE_COLUMNS]\n\n    sdf = series_df.copy()\n    for c in (\"Fluid_Sensitive\", \"Fat_Suppression\"):\n        col = sdf[c] if c in sdf.columns else pd.Series(0, index=sdf.index)\n        sdf[c] = pd.to_numeric(col, errors=\"coerce\").fillna(0).astype(int)\n    raw_plane = (sdf[\"Anatomical_Plane\"] if \"Anatomical_Plane\" in sdf.columns\n                 else pd.Series(\"\", index=sdf.index))\n    plane = raw_plane.astype(str).str.strip().str.title()\n    sdf[\"_plane\"] = plane.where(plane.isin(PLANES), \"Other\")\n\n    g = sdf.groupby(\"StudyInstanceUID\")\n    out[\"n_series\"] = g.size().reindex(idx).fillna(0).astype(int)\n    counts = pd.crosstab(sdf[\"StudyInstanceUID\"], sdf[\"_plane\"])\n    for p in PLANES + [\"Other\"]:\n        col = counts[p] if p in counts.columns else pd.Series(0, index=counts.index)\n        out[\"n_plane_\" + p.lower()] = col.reindex(idx).fillna(0).astype(int)\n    denom = out[\"n_series\"].replace(0, np.nan)\n    for p in PLANES:\n        out[\"frac_plane_\" + p.lower()] = (out[\"n_plane_\" + p.lower()] / denom).fillna(0.0)\n    out[\"n_fluid_sensitive\"] = g[\"Fluid_Sensitive\"].sum().reindex(idx).fillna(0).astype(int)\n    out[\"n_fat_suppression\"] = g[\"Fat_Suppression\"].sum().reindex(idx).fillna(0).astype(int)\n    out[\"frac_fluid_sensitive\"] = (out[\"n_fluid_sensitive\"] / denom).fillna(0.0)\n    out[\"frac_fat_suppression\"] = (out[\"n_fat_suppression\"] / denom).fillna(0.0)\n    # Fluid_Sensitive and Fat_Suppression have identical global value counts, so they\n    # may be the same column. Recorded per study so the model can answer that.\n    same = ((sdf[\"Fluid_Sensitive\"] == sdf[\"Fat_Suppression\"]).astype(int)\n            .groupby(sdf[\"StudyInstanceUID\"]).min())\n    out[\"fluid_equals_fat\"] = same.reindex(idx).fillna(1).astype(int)\n    return out[SERIES_FEATURE_COLUMNS]\n\n\nseries_by_study = {}\nfor name, sdf in ((\"train\", train_series), (\"test\", test_series)):\n    if len(sdf) and \"SeriesInstanceUID\" in sdf.columns:\n        series_by_study[name] = (sdf.groupby(\"StudyInstanceUID\")[\"SeriesInstanceUID\"]\n                                    .apply(lambda s: sorted(map(str, s.tolist()))).to_dict())\n    else:\n        series_by_study[name] = {}\n\ntrain_sf = series_features(train[\"StudyInstanceUID\"].tolist(), train_series)\ntest_sf = series_features(test[\"StudyInstanceUID\"].tolist(), test_series)\nprint(\"series features:\", train_sf.shape[1], \"columns\")\nprint(train_sf.head(3).to_string())"},{"cell_type":"code","execution_count":null,"id":"6ff85d30cf57","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"NUM_TAGS = [\n    \"SliceThickness\", \"SpacingBetweenSlices\", \"RepetitionTime\", \"EchoTime\",\n    \"MagneticFieldStrength\", \"FlipAngle\", \"EchoTrainLength\", \"PixelBandwidth\",\n    \"NumberOfAverages\", \"ImagingFrequency\", \"Rows\", \"Columns\", \"BitsStored\",\n    \"PercentSampling\", \"PercentPhaseFieldOfView\",\n]\nCAT_TAGS = [\n    \"PatientSex\", \"Laterality\", \"Manufacturer\", \"ManufacturerModelName\",\n    \"SoftwareVersions\", \"ReceiveCoilName\", \"PatientPosition\", \"BodyPartExamined\",\n    \"MRAcquisitionType\",\n]\nSET_TAGS = [\"SeriesDescription\", \"ScanningSequence\", \"SequenceVariant\",\n            \"FrameOfReferenceUID\"]\nAGG3 = [\"RepetitionTime\", \"EchoTime\", \"SliceThickness\"]\n\n\ndef _tag_number(ds, keyword):\n    \"\"\"Read one DICOM tag as a float, tolerating multi-valued and absent tags.\n\n    Args:\n        ds: A pydicom Dataset read with stop_before_pixels=True.\n        keyword: DICOM keyword to read.\n\n    Returns:\n        A float, or NaN when the tag is missing or not numeric.\n    \"\"\"\n    try:\n        v = getattr(ds, keyword, None)\n        if v is None:\n            return np.nan\n        if isinstance(v, (list, tuple)) or (hasattr(v, \"__len__\") and not isinstance(v, str)):\n            v = list(v)\n            if not v:\n                return np.nan\n            v = v[0]\n        return float(v)\n    except Exception:\n        return np.nan\n\n\ndef _tag_text(ds, keyword):\n    \"\"\"Read one DICOM tag as a stripped string, or an empty string.\n\n    Args:\n        ds: A pydicom Dataset.\n        keyword: DICOM keyword to read.\n\n    Returns:\n        A string; empty when the tag is missing or unreadable.\n    \"\"\"\n    try:\n        v = getattr(ds, keyword, None)\n        if v is None:\n            return \"\"\n        if isinstance(v, (list, tuple)) or (hasattr(v, \"__len__\") and not isinstance(v, str)):\n            v = \"|\".join(str(x) for x in list(v))\n        return str(v).strip()\n    except Exception:\n        return \"\"\n\n\ndef _pick_slice_paths(series_dir):\n    \"\"\"List a series directory once and choose a bounded, deterministic sample.\n\n    Args:\n        series_dir: Directory holding one series worth of .dcm files.\n\n    Returns:\n        Tuple of (total file count, list of at most MAX_SLICES_PER_SERIES paths).\n    \"\"\"\n    try:\n        names = sorted(e.name for e in os.scandir(series_dir) if e.is_file())\n    except OSError:\n        return 0, []\n    if not names:\n        return 0, []\n    picks = [0]\n    if len(names) > 1 and MAX_SLICES_PER_SERIES > 1:\n        picks.append(len(names) // 2)\n    picks = sorted(set(picks))[:MAX_SLICES_PER_SERIES]\n    return len(names), [series_dir / names[i] for i in picks]\n\n\ndef study_header_features(study_uid, series_uids, dcm_root):\n    \"\"\"Build per-study DICOM header features under a bounded number of reads.\n\n    Reads at most MAX_SLICES_PER_SERIES slices from at most MAX_SERIES_PER_STUDY\n    series, always with stop_before_pixels=True so no transfer syntax decoder is\n    involved. Every read is guarded; failures are counted, never raised.\n\n    Args:\n        study_uid: StudyInstanceUID of the study.\n        series_uids: Series ids from the manifest, possibly empty.\n        dcm_root: Root directory holding <study>/<series>/<sop>.dcm, or None.\n\n    Returns:\n        Dict of features. Numeric aggregates are NaN when nothing could be read;\n        raw categorical values are empty strings.\n    \"\"\"\n    feat = {\"hdr_available\": 0, \"hdr_read_fail\": 0, \"hdr_series_missing\": 0,\n            \"n_slices_total\": 0.0, \"n_series_on_disk\": 0}\n    for k in NUM_TAGS:\n        feat[\"hdr_\" + k] = np.nan\n    for k in AGG3:\n        feat[\"hdr_\" + k + \"_min\"] = np.nan\n        feat[\"hdr_\" + k + \"_max\"] = np.nan\n    for k in CAT_TAGS:\n        feat[\"raw_\" + k] = \"\"\n    for k in SET_TAGS:\n        feat[\"n_distinct_\" + k] = 0\n    feat[\"slices_per_series_max\"] = np.nan\n    feat[\"slices_per_series_min\"] = np.nan\n    feat[\"patient_id\"] = \"\"\n    feat[\"pixel_spacing\"] = np.nan\n    feat[\"fov_mm\"] = np.nan\n\n    if dcm_root is None or not HAVE_PYDICOM:\n        return feat\n\n    study_dir = dcm_root / str(study_uid)\n    if not study_dir.is_dir():\n        return feat\n\n    uids = list(series_uids) if series_uids else []\n    if not uids:\n        try:\n            uids = sorted(e.name for e in os.scandir(study_dir) if e.is_dir())\n        except OSError:\n            uids = []\n    feat[\"n_series_on_disk\"] = len(uids)\n\n    numeric = {k: [] for k in NUM_TAGS}\n    spacing = []\n    slice_counts = []\n    cats = {k: [] for k in CAT_TAGS}\n    sets = {k: set() for k in SET_TAGS}\n    pids = []\n\n    for uid in uids[:MAX_SERIES_PER_STUDY]:\n        sdir = study_dir / str(uid)\n        if not sdir.is_dir():\n            feat[\"hdr_series_missing\"] += 1\n            continue\n        n_files, paths = _pick_slice_paths(sdir)\n        if n_files:\n            slice_counts.append(n_files)\n        for p in paths:\n            try:\n                ds = pydicom.dcmread(str(p), stop_before_pixels=True, force=True)\n            except Exception:\n                feat[\"hdr_read_fail\"] += 1\n                continue\n            feat[\"hdr_available\"] = 1\n            for k in NUM_TAGS:\n                v = _tag_number(ds, k)\n                if np.isfinite(v):\n                    numeric[k].append(v)\n            ps = _tag_number(ds, \"PixelSpacing\")\n            if np.isfinite(ps):\n                spacing.append(ps)\n            for k in CAT_TAGS:\n                t = _tag_text(ds, k)\n                if t:\n                    cats[k].append(t)\n            for k in SET_TAGS:\n                t = _tag_text(ds, k)\n                if t:\n                    sets[k].add(t)\n            pid = _tag_text(ds, \"PatientID\")\n            if pid:\n                pids.append(pid)\n\n    for k in NUM_TAGS:\n        if numeric[k]:\n            feat[\"hdr_\" + k] = float(np.mean(numeric[k]))\n    for k in AGG3:\n        if numeric[k]:\n            feat[\"hdr_\" + k + \"_min\"] = float(np.min(numeric[k]))\n            feat[\"hdr_\" + k + \"_max\"] = float(np.max(numeric[k]))\n    if spacing:\n        feat[\"pixel_spacing\"] = float(np.mean(spacing))\n    if spacing and numeric[\"Rows\"]:\n        feat[\"fov_mm\"] = float(np.mean(spacing) * np.mean(numeric[\"Rows\"]))\n    if slice_counts:\n        feat[\"n_slices_total\"] = float(np.sum(slice_counts))\n        feat[\"slices_per_series_max\"] = float(np.max(slice_counts))\n        feat[\"slices_per_series_min\"] = float(np.min(slice_counts))\n    for k in CAT_TAGS:\n        if cats[k]:\n            feat[\"raw_\" + k] = pd.Series(cats[k]).mode().iloc[0]\n    for k in SET_TAGS:\n        feat[\"n_distinct_\" + k] = len(sets[k])\n    if pids:\n        feat[\"patient_id\"] = pd.Series(pids).mode().iloc[0]\n    return feat\n\n\ndef build_header_frame(study_ids, series_map, dcm_root, cache=None, deadline=None):\n    \"\"\"Run the header walk over many studies, honouring a wall-clock deadline.\n\n    Args:\n        study_ids: Sequence of StudyInstanceUID values, processed in order.\n        series_map: Dict from study id to list of series ids.\n        dcm_root: Root DICOM directory, or None.\n        cache: Optional dict of already-computed feature rows, keyed by study id.\n        deadline: Optional time.time() value past which the walk stops early.\n\n    Returns:\n        Tuple of (DataFrame indexed by the studies actually processed, count of\n        studies skipped because the deadline was reached).\n    \"\"\"\n    cache = cache or {}\n    rows = {}\n    skipped = 0\n    for i, sid in enumerate(study_ids):\n        if sid in cache:\n            rows[sid] = cache[sid]\n            continue\n        if deadline is not None and time.time() > deadline:\n            skipped = len(study_ids) - i\n            break\n        rows[sid] = study_header_features(sid, series_map.get(sid, []), dcm_root)\n    df = pd.DataFrame.from_dict(rows, orient=\"index\")\n    df.index.name = \"StudyInstanceUID\"\n    return df, skipped\n\n\nprint(\"header feature function ready |\", len(NUM_TAGS), \"numeric tags,\",\n      len(CAT_TAGS), \"categorical tags,\", len(SET_TAGS), \"set-cardinality tags\")"},{"cell_type":"markdown","id":"a28bd817d207","metadata":{},"source":"### Runtime projection before the work\n\nThe visible `test.csv` is a placeholder. At rerun it holds roughly 1,300 studies this\nnotebook has never seen. Before extracting anything at scale, a short probe times the\nheader walk on a few studies and projects the full cost, for both the visible test set\nand the rerun-sized one. If the projected training cost exceeds the budget, the\ntraining studies are subsampled deterministically, with the gold studies always kept."},{"cell_type":"code","execution_count":null,"id":"176231bffa61","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"test_ids = test[\"StudyInstanceUID\"].astype(str).tolist()\ntrain_ids = train[\"StudyInstanceUID\"].astype(str).tolist()\n\nprobe_ids = test_ids[:min(PROBE_STUDIES, len(test_ids))]\nt0 = time.time()\nprobe_cache = {}\nfor sid in probe_ids:\n    probe_cache[sid] = study_header_features(sid, series_by_study[\"test\"].get(sid, []), TEST_DCM)\nprobe_s = time.time() - t0\nper_study = probe_s / max(1, len(probe_ids))\n\nRERUN_TEST_STUDIES = 1300  # the host's stated hidden test size, used for projection only\nprint(\"probed {} test studies in {:.2f}s -> {:.3f}s per study\".format(\n    len(probe_ids), probe_s, per_study))\nprint()\nprint(\"projected header-walk cost\")\nprint(\"  visible test  ({:5d} studies): {:6.1f}s\".format(len(test_ids), per_study * len(test_ids)))\nprint(\"  rerun test    (~{:4d} studies): {:6.1f}s\".format(\n    RERUN_TEST_STUDIES, per_study * RERUN_TEST_STUDIES))\nprint(\"  full train    ({:5d} studies): {:6.1f}s\".format(len(train_ids), per_study * len(train_ids)))\nprint(\"  budget for train features      : {:6.1f}s\".format(TRAIN_FEATURE_BUDGET_S))\ncap = MAX_SERIES_PER_STUDY * MAX_SLICES_PER_SERIES\nprint()\nprint(\"hard bound on header reads: {} per study, so at most {:,} for a {}-study \"\n      \"rerun and {:,} for the full train set\".format(\n          cap, cap * RERUN_TEST_STUDIES, RERUN_TEST_STUDIES, cap * len(train_ids)))\n\nif per_study <= 0 or per_study * len(train_ids) <= TRAIN_FEATURE_BUDGET_S:\n    n_train_use = len(train_ids)\n    reason = \"full train set fits the budget\"\nelse:\n    n_train_use = max(MIN_TRAIN_STUDIES, int(TRAIN_FEATURE_BUDGET_S / per_study))\n    n_train_use = min(n_train_use, len(train_ids))\n    reason = \"subsampled to fit the budget\"\n\n# Gold studies go first so they survive any truncation, then a deterministic shuffle.\ngold_ids = train.loc[gold_mask, \"StudyInstanceUID\"].astype(str).tolist()\nrest = [s for s in train_ids if s not in set(gold_ids)]\nrng = np.random.default_rng(SEED)\nrng.shuffle(rest)\ntrain_order = gold_ids + rest\ntrain_use = train_order[:n_train_use]\nprint()\nprint(\"training studies to featurise: {} ({})\".format(len(train_use), reason))"},{"cell_type":"code","execution_count":null,"id":"20861875fb10","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"t0 = time.time()\ntest_hdr, test_skipped = build_header_frame(\n    test_ids, series_by_study[\"test\"], TEST_DCM, cache=probe_cache)\nprint(\"test headers : {} studies in {:.1f}s (skipped {})\".format(\n    len(test_hdr), time.time() - t0, test_skipped))\n\nt0 = time.time()\ndeadline = t0 + TRAIN_FEATURE_BUDGET_S * 1.5\ntrain_hdr, train_skipped = build_header_frame(\n    train_use, series_by_study[\"train\"], TRAIN_DCM, deadline=deadline)\nprint(\"train headers: {} studies in {:.1f}s (skipped {})\".format(\n    len(train_hdr), time.time() - t0, train_skipped))\n\nfor name, df in ((\"train\", train_hdr), (\"test\", test_hdr)):\n    if len(df) == 0:\n        print(name, \"headers: none read\")\n        continue\n    print(\"{:5s} header reads available on {:.1f}% of studies | failed reads {} | missing series dirs {}\".format(\n        name, 100.0 * df[\"hdr_available\"].mean(), int(df[\"hdr_read_fail\"].sum()),\n        int(df[\"hdr_series_missing\"].sum())))"},{"cell_type":"code","execution_count":null,"id":"2d0ccfd9a7d8","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"RAW_CAT_COLS = [\"raw_\" + k for k in CAT_TAGS]\n\n\ndef assemble(study_ids, series_feat, header_feat):\n    \"\"\"Join the series and header feature blocks onto one row per study.\n\n    Args:\n        study_ids: Studies to produce rows for, in order.\n        series_feat: Output of series_features, indexed by study id.\n        header_feat: Output of build_header_frame, indexed by study id.\n\n    Returns:\n        DataFrame indexed by study id with both blocks present, columns in a\n        fixed order, header columns NaN or empty where no header could be read.\n    \"\"\"\n    idx = pd.Index([str(s) for s in study_ids], name=\"StudyInstanceUID\")\n    sf = series_feat.reindex(idx)[SERIES_FEATURE_COLUMNS]\n    if len(header_feat):\n        hf = header_feat.reindex(idx)\n    else:\n        hf = pd.DataFrame(index=idx)\n    template = study_header_features(\"__none__\", [], None)\n    for k, v in template.items():\n        if k not in hf.columns:\n            hf[k] = v\n    for c in RAW_CAT_COLS + [\"patient_id\"]:\n        hf[c] = hf[c].fillna(\"\").astype(str)\n    # Pin the header block to the template's key order too. Both blocks now have a\n    # declared order, so train and test cannot drift apart no matter which branch\n    # each of them took.\n    hf = hf[list(template.keys())]\n    return pd.concat([sf, hf], axis=1)\n\n\n# Only studies the header walk actually reached go into training. If the deadline\n# truncated the walk, the remainder is dropped rather than fed in on fill values.\ntrain_featured = list(train_hdr.index) if len(train_hdr) else list(train_use)\nif len(train_featured) < len(train_use):\n    print(\"header walk was truncated: training on {} of {} planned studies\".format(\n        len(train_featured), len(train_use)))\n\ntrain_X_raw = assemble(train_featured, train_sf, train_hdr)\ntest_X_raw = assemble(test_ids, test_sf, test_hdr)\nassert list(train_X_raw.columns) == list(test_X_raw.columns), (\n    \"assembled train and test columns diverged: {}\".format(\n        [(a, b) for a, b in zip(train_X_raw.columns, test_X_raw.columns) if a != b][:5]))\n\n# Categorical encoding is fitted on train alone and applied unchanged to test.\n# Codes: -1 for a value never seen in train, -2 for a missing value.\ncat_maps = {}\nfreq_maps = {}\nfor c in RAW_CAT_COLS:\n    vals = train_X_raw[c]\n    seen = sorted(v for v in vals.unique() if v != \"\")\n    cat_maps[c] = {v: i for i, v in enumerate(seen)}\n    freq_maps[c] = vals[vals != \"\"].value_counts(normalize=True).to_dict()\n\n\ndef encode_categoricals(df):\n    \"\"\"Replace raw categorical columns with train-fitted codes and frequencies.\n\n    Args:\n        df: Assembled feature frame carrying the raw_* string columns.\n\n    Returns:\n        A copy with code_* and freq_* columns and the raw_* columns dropped.\n    \"\"\"\n    out = df.copy()\n    for c in RAW_CAT_COLS:\n        s = out[c].astype(str)\n        out[\"code_\" + c[4:]] = [\n            -2 if v == \"\" else cat_maps[c].get(v, -1) for v in s\n        ]\n        out[\"freq_\" + c[4:]] = [\n            0.0 if v == \"\" else float(freq_maps[c].get(v, 0.0)) for v in s\n        ]\n    return out.drop(columns=RAW_CAT_COLS)\n\n\ntrain_enc = encode_categoricals(train_X_raw)\ntest_enc = encode_categoricals(test_X_raw)\n\ngroups_raw = train_X_raw[\"patient_id\"].astype(str).values\ngroups = np.array([g if g else \"study::\" + sid\n                   for g, sid in zip(groups_raw, train_X_raw.index)])\nprint(\"grouping key : {} distinct patients across {} studies\".format(\n    len(set(groups)), len(groups)))\nprint(\"             : {} studies fell back to their own study id\".format(\n    int((groups_raw == \"\").sum())))\n\n\ndef merge_duplicate_reports(group_keys, study_ids, report_by_study, max_share=0.25):\n    \"\"\"Merge studies that share an identical report text into one group.\n\n    Forty-nine report texts in this corpus are shared by 183 studies, the largest\n    block being 37 studies carrying one identical normal read. The extractor gives\n    every study in such a block the same twelve values, so a block split across\n    folds lets the model score a memorised target rather than a generalised one.\n    A union-find merges each block with whatever patient groups it touches, so the\n    block and those patients land in the same fold.\n\n    Args:\n        group_keys: Current group key per study, in study_ids order.\n        study_ids: Study ids aligned with group_keys.\n        report_by_study: Map from study id to its report text, empty where absent.\n        max_share: Refuse the merge if it would put more than this share of the\n            training rows into a single group, which would make a balanced split\n            impossible.\n\n    Returns:\n        Tuple of (group keys after merging, number of blocks merged, size of the\n        largest resulting group, whether the merge was applied).\n    \"\"\"\n    parent = {}\n\n    def find(a):\n        \"\"\"Return the representative of a's set, with path compression.\"\"\"\n        parent.setdefault(a, a)\n        while parent[a] != a:\n            parent[a] = parent[parent[a]]\n            a = parent[a]\n        return a\n\n    def union(a, b):\n        \"\"\"Merge the sets containing a and b.\"\"\"\n        ra, rb = find(a), find(b)\n        if ra != rb:\n            parent[ra] = rb\n\n    by_report = {}\n    for g, sid in zip(group_keys, study_ids):\n        txt = report_by_study.get(str(sid), \"\")\n        if not txt:\n            continue\n        by_report.setdefault(txt, []).append(g)\n\n    n_blocks = 0\n    for gs in by_report.values():\n        if len(set(gs)) < 2:\n            continue\n        n_blocks += 1\n        for other in gs[1:]:\n            union(other, gs[0])\n\n    merged = np.array([find(g) for g in group_keys])\n    sizes = pd.Series(merged).value_counts()\n    largest = int(sizes.iloc[0]) if len(sizes) else 0\n    if largest > max_share * max(1, len(merged)):\n        return np.asarray(group_keys), n_blocks, largest, False\n    return merged, n_blocks, largest, True\n\n\nreport_by_study = {}\nif report_col:\n    for _sid, _txt in zip(train[\"StudyInstanceUID\"].astype(str), train[report_col]):\n        report_by_study[_sid] = _txt.strip() if isinstance(_txt, str) else \"\"\n\ngroups, n_dup_blocks, largest_group, merge_ok = merge_duplicate_reports(\n    groups, list(train_X_raw.index), report_by_study)\nprint(\"             : {} duplicate-report blocks merged into their patient groups, \"\n      \"leaving {} groups\".format(n_dup_blocks, len(set(groups))))\nprint(\"             : largest group after merging holds {} of {} studies{}\".format(\n    largest_group, len(groups),\n    \"\" if merge_ok else \"; merge refused, a balanced split comes first\"))\n\n# A feature block present on one side and absent on the other is worse than no\n# feature block at all: the model fits it on train and extrapolates on test, and\n# every prediction collapses toward one row. Measure the gap, and drop the whole\n# header block rather than trust it.\nHDR_PREFIXES = (\"hdr_\", \"code_\", \"freq_\", \"n_distinct_\", \"n_slices\",\n                \"slices_per_series\", \"pixel_spacing\", \"fov_mm\", \"n_series_on_disk\")\n\n\ndef header_coverage(df):\n    \"\"\"Fraction of studies where at least one DICOM header could be read.\n\n    Args:\n        df: An encoded feature frame.\n\n    Returns:\n        Float in [0, 1]; 0.0 when the column is absent.\n    \"\"\"\n    if \"hdr_available\" not in df.columns or len(df) == 0:\n        return 0.0\n    return float(pd.to_numeric(df[\"hdr_available\"], errors=\"coerce\").fillna(0).mean())\n\n\ncov_train = header_coverage(train_enc)\ncov_test = header_coverage(test_enc)\nHEADER_OK = abs(cov_train - cov_test) <= 0.25\nprint(\"header coverage: train {:.1%} | test {:.1%}\".format(cov_train, cov_test))\nif not HEADER_OK:\n    print(\"coverage gap above 25 points, so every header-derived column is dropped;\")\n    print(\"the model falls back to the series manifest, which both sides always have\")\n\ndrop_cols = [\"patient_id\"]\nFEATURE_COLUMNS = [c for c in train_enc.columns if c not in drop_cols]\nFEATURE_COLUMNS = [c for c in FEATURE_COLUMNS\n                   if pd.api.types.is_numeric_dtype(train_enc[c])]\nif not HEADER_OK:\n    FEATURE_COLUMNS = [c for c in FEATURE_COLUMNS\n                       if not c.startswith(HDR_PREFIXES)]\n\n# One fill policy, decided here, applied identically to train and test.\nFILL_VALUES = train_enc[FEATURE_COLUMNS].median(numeric_only=True)\nFILL_VALUES = FILL_VALUES.fillna(0.0)\n\nX_train = train_enc.reindex(columns=FEATURE_COLUMNS).astype(float).fillna(FILL_VALUES)\nX_test = test_enc.reindex(columns=FEATURE_COLUMNS).astype(float).fillna(FILL_VALUES)\nX_train = X_train.replace([np.inf, -np.inf], np.nan).fillna(FILL_VALUES).fillna(0.0)\nX_test = X_test.replace([np.inf, -np.inf], np.nan).fillna(FILL_VALUES).fillna(0.0)\n\nassert list(X_train.columns) == list(X_test.columns), \"train/test feature mismatch\"\nassert np.isfinite(X_train.values).all(), \"non-finite value in the train matrix\"\nassert np.isfinite(X_test.values).all(), \"non-finite value in the test matrix\"\n\n# Constant columns carry nothing and only slow the fit. Never drop everything.\nnunique = X_train.nunique()\nkeep = [c for c in X_train.columns if nunique[c] > 1]\nif not keep:\n    keep = list(X_train.columns[:1])\n    print(\"every feature is constant; keeping one column so the fit can still run\")\ndropped = [c for c in X_train.columns if c not in keep]\nX_train = X_train[keep]\nX_test = X_test[keep]\nFEATURE_COLUMNS = keep\n\nprint()\nprint(\"feature matrix: train {} x {} | test {} x {}\".format(\n    X_train.shape[0], X_train.shape[1], X_test.shape[0], X_test.shape[1]))\nprint(\"dropped {} constant columns\".format(len(dropped)))\nn_all_fill = int((X_test.eq(FILL_VALUES.reindex(X_test.columns)).all(axis=1)).sum())\nprint(\"test studies sitting entirely on fill values: {} of {}\".format(\n    n_all_fill, len(X_test)))\n\nrule_by_study = rule.copy()\nrule_by_study.index = train[\"StudyInstanceUID\"].astype(str).values\ny = rule_by_study.reindex(X_train.index).astype(int)\nassert list(y.index) == list(X_train.index), \"label rows drifted from feature rows\"\nprint(\"label matrix : {} x {}\".format(y.shape[0], y.shape[1]))"},{"cell_type":"markdown","id":"4b04454b6966","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">5</span><span style=\"vertical-align:middle;\">Model</span></h2></div>\n\nOne LightGBM binary classifier per label, trained on the rule-extracted labels, with\nLogisticRegression as the fallback if LightGBM is unavailable. Fixed rounds, no early\nstopping on the validation fold, because early stopping on the fold you then score is\nhow an out-of-fold estimate gets inflated.\n\nFolds are `GroupKFold` grouped by DICOM `PatientID`, not by study. A patient can\ncontribute more than one study, and grouping by study would leak a patient across the\nsplit. Where no header could be read the study id is used as its own group, which is\nthe conservative choice.\n\nSomething else has to share a fold. Forty-nine report texts in this corpus are shared\nby 183 studies, the largest block being 37 studies carrying one identical normal read.\nThe extractor hands every study in such a block the same twelve values, so a block\nsplit across folds lets a model score a target it has already seen rather than one it\ngeneralised to. Each block is merged with the patient groups it touches before the\nsplit. The grouping cell prints how many blocks that was.\n\nTwo out-of-fold numbers get reported and they measure different things."},{"cell_type":"code","execution_count":null,"id":"1d95a108a2ad","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"from sklearn.model_selection import GroupKFold\nfrom sklearn.metrics import roc_auc_score\n\ntry:\n    import lightgbm as lgb\n    HAVE_LGB = True\nexcept Exception as exc:\n    HAVE_LGB = False\n    print(\"lightgbm unavailable, falling back to logistic regression:\", exc)\n\nif not HAVE_LGB:\n    from sklearn.linear_model import LogisticRegression\n    from sklearn.pipeline import make_pipeline\n    from sklearn.preprocessing import StandardScaler\n\nN_JOBS = max(1, min(4, os.cpu_count() or 2))\nLGB_PARAMS = dict(\n    objective=\"binary\", learning_rate=0.05, n_estimators=300, num_leaves=31,\n    min_child_samples=40, subsample=0.8, subsample_freq=1, colsample_bytree=0.8,\n    reg_lambda=1.0, random_state=SEED, n_jobs=N_JOBS, verbose=-1,\n    importance_type=\"gain\",\n)\n\n\ndef make_model():\n    \"\"\"Return a fresh per-label classifier.\n\n    Fixed rounds, no early stopping: stopping on the fold you then score inflates\n    the out-of-fold estimate.\n\n    Returns:\n        An unfitted sklearn-compatible binary classifier.\n    \"\"\"\n    if HAVE_LGB:\n        return lgb.LGBMClassifier(**LGB_PARAMS)\n    return make_pipeline(\n        StandardScaler(),\n        LogisticRegression(max_iter=2000, C=0.5, random_state=SEED),\n    )\n\n\nXv = X_train.values\nXt = X_test.values\nn_train, n_test_rows = Xv.shape[0], Xt.shape[0]\nn_splits = max(2, min(N_FOLDS, len(set(groups))))\ngkf = GroupKFold(n_splits=n_splits)\n\noof = np.full((n_train, len(LABELS)), np.nan)\ntest_pred = np.zeros((n_test_rows, len(LABELS)))\nimportance = np.zeros(len(FEATURE_COLUMNS))\nmodel_notes = []\n\nt0 = time.time()\nfor j, lab in enumerate(LABELS):\n    yv = y[lab].values.astype(int)\n    pos = int(yv.sum())\n    if pos < MIN_POSITIVES or (len(yv) - pos) < MIN_POSITIVES:\n        prior = float(yv.mean()) if len(yv) else 0.5\n        oof[:, j] = prior\n        test_pred[:, j] = prior\n        model_notes.append((lab, pos, \"constant prior, too few positives\"))\n        continue\n    fold_test = np.zeros(n_test_rows)\n    n_used = 0\n    for tr_idx, va_idx in gkf.split(Xv, yv, groups=groups):\n        if len(set(yv[tr_idx])) < 2:\n            continue\n        model = make_model()\n        model.fit(Xv[tr_idx], yv[tr_idx])\n        oof[va_idx, j] = model.predict_proba(Xv[va_idx])[:, 1]\n        fold_test += model.predict_proba(Xt)[:, 1]\n        n_used += 1\n        if HAVE_LGB and hasattr(model, \"feature_importances_\"):\n            importance += model.feature_importances_.astype(float)\n    if n_used == 0:\n        prior = float(yv.mean())\n        oof[:, j] = prior\n        test_pred[:, j] = prior\n        model_notes.append((lab, pos, \"constant prior, no usable fold\"))\n        continue\n    test_pred[:, j] = fold_test / n_used\n    model_notes.append((lab, pos, \"{} folds\".format(n_used)))\n\n# Any label whose folds all collapsed leaves NaN behind; fall back to its prior.\nfor j, lab in enumerate(LABELS):\n    m = ~np.isfinite(oof[:, j])\n    if m.any():\n        oof[m, j] = float(y[lab].mean())\n\nprint(\"trained {} labels in {:.1f}s using {}\".format(\n    len(LABELS), time.time() - t0, \"LightGBM\" if HAVE_LGB else \"LogisticRegression\"))\nfor lab, pos, note in model_notes:\n    print(\"  {:18s} rule positives {:5d}  {}\".format(lab, pos, note))"},{"cell_type":"code","execution_count":null,"id":"de5db7975490","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"gold_ids_present = [s for s in gold_ids if s in set(X_train.index)]\ngold_pos_in_X = X_train.index.get_indexer(pd.Index(gold_ids_present))\ngold_truth = (train.set_index(train[\"StudyInstanceUID\"].astype(str))\n                   .loc[gold_ids_present, LABELS].astype(int))\n\nrows = []\nfor j, lab in enumerate(LABELS):\n    yv = y[lab].values.astype(int)\n    try:\n        a_rule = roc_auc_score(yv, oof[:, j]) if len(set(yv)) == 2 else np.nan\n    except Exception:\n        a_rule = np.nan\n    yg = gold_truth[lab].values\n    try:\n        a_gold = roc_auc_score(yg, oof[gold_pos_in_X, j]) if len(set(yg)) == 2 else np.nan\n    except Exception:\n        a_gold = np.nan\n    rows.append({\"label\": lab, \"rule_pos\": int(yv.sum()),\n                 \"auc_vs_rule\": a_rule, \"auc_vs_gold\": a_gold})\n\nscores = pd.DataFrame(rows)\nmacro_rule = float(np.nanmean(scores[\"auc_vs_rule\"]))\nmacro_gold = float(np.nanmean(scores[\"auc_vs_gold\"]))\n\nprint(scores.round(4).to_string(index=False))\nprint()\nprint(\"macro AUC vs rule labels   : {:.4f}   (self-consistency, not a leaderboard estimate)\".format(macro_rule))\nprint(\"macro AUC vs gold labels   : {:.4f}   (n = {} studies, far too few to trust)\".format(\n    macro_gold, len(gold_ids_present)))"},{"cell_type":"markdown","id":"6eea6c01f236","metadata":{},"source":"### Why the two AUC numbers are not comparable\n\nThe first number is out-of-fold macro AUC **against the rule-extracted labels**. It\nanswers: can metadata predict what the keyword extractor said? Both sides of that\ncomparison come from the same noisy source, so it is a self-consistency check. It\ntells you the pipeline runs and the features carry some signal. It does not estimate\nthe leaderboard.\n\nThe second number is out-of-fold macro AUC **against the gold labels**, computed on\nthe handful of studies that have them. That is the right target, and it is the number\nwhose direction matters. It is also computed on far too few studies to trust. A\nper-label AUC on a few dozen studies has a standard error near 0.07, and the macro\naverage of twelve such estimates is better but still nowhere near the resolution\nneeded to choose between two models.\n\nTreat the first as a smoke test and the second as a weak sanity check. Neither is a\nmodel-selection tool. Model ranking in this competition has to come from a much larger\nreport-derived validation set, which is the next thing to build."},{"cell_type":"code","execution_count":null,"id":"d2bbd5d85418","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"plot = scores.iloc[::-1].reset_index(drop=True)\nypos = np.arange(len(plot))\nh, gap = 0.34, 0.05  # the gap keeps a surface line between the two bars of a group\n\nfig, ax = plt.subplots(figsize=(14.0, 8.4))\nax.barh(ypos + (h + gap) / 2, np.nan_to_num(plot[\"auc_vs_rule\"].values, nan=0.0),\n        height=h, color=PRIMARY, edgecolor=\"none\", zorder=3,\n        label=\"vs rule labels, {:,} training studies\".format(len(y)))\nax.barh(ypos - (h + gap) / 2, np.nan_to_num(plot[\"auc_vs_gold\"].values, nan=0.0),\n        height=h, color=ACCENT, edgecolor=\"none\", zorder=3,\n        label=\"vs gold labels, n = {}\".format(len(gold_ids_present)))\nlabel_bars(ax, plot[\"auc_vs_rule\"].values, ypos + (h + gap) / 2, fmt=\"{:.3f}\", pad=0.008)\nlabel_bars(ax, plot[\"auc_vs_gold\"].values, ypos - (h + gap) / 2, fmt=\"{:.3f}\", pad=0.008)\nax.axvline(0.5, color=AXIS, linewidth=1.2, zorder=2)\nax.text(0.5, len(plot) - 0.35, \" random\", fontsize=11.25, color=MUTED, va=\"center\")\nax.set_yticks(ypos)\nax.set_yticklabels(plot[\"label\"].values, fontsize=12.5, color=INK)\nax.set_ylim(-0.9, len(plot) - 0.3)\nax.set_xlim(0, 1.10)\nax.set_xlabel(\"out-of-fold ROC AUC\")\nstyle(ax, \"x\")\nax.spines[\"left\"].set_visible(False)\n# Legend below the axis so it can never sit on top of a bar.\nax.legend(loc=\"upper center\", bbox_to_anchor=(0.5, -0.075), ncol=2, labelcolor=INK_2)\n\nfinish(fig, \"Two numbers, two different questions, and only one of them is the target\",\n       detail=\"Macro average {:.3f} against the rule labels, {:.3f} against the gold \"\n              \"labels.\".format(macro_rule, macro_gold),\n       footroom=0.30,\n       caption=\"Blue is self-consistency: can metadata predict what the keyword extractor \"\n               \"said. Orange is the real target, measured on {} studies, which is far too \"\n               \"few to rank anything. Folds are GroupKFold on DICOM PatientID, with \"\n               \"studies sharing one report text held in the same fold.\".format(\n                   len(gold_ids_present)))"},{"cell_type":"code","execution_count":null,"id":"b2f18ce9ff1f","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"if HAVE_LGB and importance.sum() > 0:\n    n_show = min(20, len(FEATURE_COLUMNS))\n    imp = (pd.Series(importance, index=FEATURE_COLUMNS)\n             .sort_values(ascending=False).head(n_show))\n    print(imp.round(1).to_string())\n    plot_imp = imp.iloc[::-1]\n    ypos = np.arange(len(plot_imp))\n    top = float(plot_imp.values.max()) or 1.0\n\n    fig, ax = plt.subplots(figsize=(14.0, 8.6))\n    ax.barh(ypos, plot_imp.values, height=0.68, color=PRIMARY, edgecolor=\"none\",\n            zorder=3)\n    label_bars(ax, plot_imp.values, ypos, fmt=\"{:,.0f}\", pad=top * 0.014)\n    ax.set_yticks(ypos)\n    ax.set_yticklabels(plot_imp.index, fontsize=12.5, color=INK)\n    ax.set_ylim(-0.7, len(plot_imp) - 0.3)\n    ax.set_xlim(0, top * 1.20)\n    ax.set_xlabel(\"total split gain, summed over the twelve labels and every fold\")\n    ax.xaxis.set_major_formatter(\n        mpl.ticker.FuncFormatter(lambda v, _: \"{:,.0f}\".format(v)))\n    style(ax, \"x\")\n    ax.spines[\"left\"].set_visible(False)\n\n    finish(fig, \"Every feature it leans on describes the acquisition, not the knee\",\n           detail=(\"All {} surviving features by gain.\".format(n_show)\n                   if n_show == len(FEATURE_COLUMNS)\n                   else \"Top {} of {} features by gain.\".format(\n                       n_show, len(FEATURE_COLUMNS))),\n           caption=\"Protocol and site descriptors. Whatever this model knows, it learned \"\n                   \"from where and how the study was acquired, which is exactly the kind \"\n                   \"of shortcut the host's prevalence warning says will not transfer.\")\nelse:\n    print(\"no gain-based importance available (LightGBM not used)\")"},{"cell_type":"markdown","id":"d4804a18b684","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">6</span><span style=\"vertical-align:middle;\">Submission</span></h2></div>\n\nColumn order is taken from `sample_submission.csv` rather than assumed. Row identity\ncomes from `test.csv`, so the submission covers exactly the studies the grader expects,\nin the order they appear. Everything is asserted before the file is written: row\ncount, column set, finiteness, and range."},{"cell_type":"code","execution_count":null,"id":"81b1a71844c2","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"if list(sample_sub.columns[:1]) == [\"StudyInstanceUID\"] and set(LABELS).issubset(sample_sub.columns):\n    SUB_COLS = list(sample_sub.columns)\nelse:\n    SUB_COLS = [\"StudyInstanceUID\"] + LABELS\n    print(\"sample_submission columns unusable; falling back to the documented order\")\n\nsub = pd.DataFrame({\"StudyInstanceUID\": test[\"StudyInstanceUID\"].astype(str).values})\nfor j, lab in enumerate(LABELS):\n    sub[lab] = test_pred[:, j]\nsub = sub[SUB_COLS]\nsub[LABELS] = sub[LABELS].astype(float).clip(1e-6, 1 - 1e-6)\n\nvals = sub[LABELS].values\nassert len(sub) == len(test), \"row count {} != test.csv row count {}\".format(len(sub), len(test))\nassert list(sub.columns) == SUB_COLS, \"submission column order drifted\"\nassert sub[\"StudyInstanceUID\"].is_unique, \"duplicate StudyInstanceUID in the submission\"\nassert np.isfinite(vals).all(), \"non-finite prediction present\"\nassert (vals >= 0).all() and (vals <= 1).all(), \"prediction outside [0, 1]\"\n\nout_path = WORK / \"submission.csv\"\nsub.to_csv(out_path, index=False)\nprint(\"wrote\", out_path, \"|\", len(sub), \"rows x\", len(sub.columns), \"columns\")\nprint()\nprint(sub.head().to_string(index=False))\nprint()\nprint(\"per-label mean predicted probability\")\nprint(sub[LABELS].mean().round(4).to_string())"},{"cell_type":"markdown","id":"9d2fb86ed0ea","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">7</span><span style=\"vertical-align:middle;\">What this is not</span></h2></div>\n\nIt never opens a pixel. Every prediction comes from scanner and protocol metadata, so\nwhatever signal it finds is a correlation between imaging site and disease prevalence,\nnot a reading of the knee. The host warns that prevalence is not guaranteed to match\nacross the training, public and final evaluation sets, which is precisely the kind of\nshortcut that decays between splits. Whatever it scores, expect that number to move\nbetween the public and the private split.\n\nThe labels are keyword output with a negation guard. The audit above puts the honest\nceiling on what a model trained on them can learn.\n\nThe obvious next steps, in the order they pay: replace the keyword extractor with an\nLLM extraction audited against the same gold studies; build a report-derived\nvalidation set large enough to rank models; then train an imaging model and treat this\nmetadata model as a feature source rather than a predictor."},{"cell_type":"markdown","id":"d4cde36bf5bf","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"vertical-align:middle;\">How this notebook was made</span></h2></div>\n\nThis notebook was generated by an agentic system rather than typed by hand. The system\nis Claude Code running a custom multi-agent LLM harness: one agent drafted the pipeline,\nthen separate adversarial verifier agents tried to break it and sent each defect back to\nthe builder script instead of patching the notebook. The verifier agents are the part\nthat matters, because a notebook that has never been run is a draft.\n\nWhat the verification actually did. Two rounds of it, the second run by an agent told to\nassume nothing the first had reported was true. Between them the notebook was executed end\nto end more than a dozen times, always whole, never a fragment. Once against a local mirror\nof the Kaggle input mount. Repeatedly under `bubblewrap` against synthetic `/kaggle/input`\ntrees built to break it: forty study ids it had never seen, with pixels present for only a\nquarter of them, one study whose files were unreadable bytes and one whose manifest named a\nseries absent from disk; a tree carrying no test manifest and no readable pixel at all; and\na tree where 2,203 patients owned two or three studies each, placed far apart in the file,\nchecked against a ground-truth patient map the notebook was never shown. Every run emitted\none prediction row per test study, and no patient landed in two folds.\n\nThree real defects came out of it. The train and test branches of the series-feature\nbuilder emitted the same column *set* in a different *order*, which every set-based check\npasses and which only bites a model fed the frame positionally. The builder was not\nidempotent, so two consecutive builds of unchanged content produced different cell ids and\na meaningless diff. And the 49 report texts shared by 183 studies were being split across\nfolds, which hands the model a target it has already seen; those blocks are now merged into\ntheir patient groups before the split. All three were fixed in the builder.\n\nThe source of truth is a Python builder script, `notebooks/build_baseline_notebook.py`\nin the working repository, which emits this `.ipynb` through `nbformat`. The notebook is\nan artifact; review happens on plain Python, and the notebook is regenerated rather than\nedited. Publication uses `kaggle kernels push` and submission uses\n`kaggle competitions submit`, the same path the companion EDA notebook takes.\n\nWhat this is: a plumbing proof and a floor, built from metadata the test set actually\ncarries and labels a keyword extractor produced. It has no leaderboard score and claims\nnone. If you build on it, the label extractor is the part worth replacing first."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"mimetype":"text/x-python","name":"python"}},"nbformat":4,"nbformat_minor":5}