{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\n# RSNA Knee Abnormality Detection\n**Single Kaggle notebook.**\n\nThis version is built around the actual competition structure: 4,407 training studies, only 58 with expert labels, 12 target findings, 24,371 MRI series, and reports for the remaining studies. The competition provides `train_series` metadata with plane / fluid-sensitive / fat-suppression flags, and test reports are not provided.\n\n## Strategy\n\n1. Robustly discover the competition data.\n2. Build a study/series index without reading the 570 GB tree repeatedly.\n3. Parse reports into **probabilities with an explicit unknown state (0.5)**.\n4. If an attached LLM report-label dataset exists, use it as the preferred weak-label teacher; otherwise use the conservative parser.\n5. Pretrain an MRI model on all 4,407 studies using weak labels.\n6. Fine-tune / calibrate on the 58 gold studies with strict study-level CV.\n7. Use a **multi-series, 2.5D, attention-pooling model**:\n   - 6 anatomical/protocol slots\n   - adjacent-slice triplets\n   - pretrained ImageNet backbone when available\n   - slice attention\n   - series attention / masked slot fusion\n8. Generate OOF predictions.\n9. Blend weak-label prior + MRI predictions using OOF optimization.\n10. Retrain on all available training studies and infer test.\n11. Validate submission against `sample_submission.csv`.\n\n### Important\nThe notebook uses `DEV_MODE=True` by default. This is deliberate: first prove that the pipeline runs on Kaggle before spending hours on the full 570 GB DICOM tree.\n","metadata":{}},{"cell_type":"code","source":"\n# ============================================================\n# 0. CONFIG\n# ============================================================\nfrom pathlib import Path\nimport os, re, gc, json, math, random, warnings, time\nfrom collections import Counter, defaultdict\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\nwarnings.filterwarnings(\"ignore\")\nSEED = 42\nrandom.seed(SEED)\nnp.random.seed(SEED)\n\nLABELS = [\n    \"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\",\n    \"Medial OA\", \"Lateral OA\", \"PF OA\", \"Effusion\",\n    \"Synovitis\", \"Baker's\", \"Contusion\", \"Fracture\"\n]\n\nDEV_MODE = False\nN_FOLDS = 5\n\n# Development limits. Set to 0 for full run.\nDEV_STUDIES = 120\nDEV_SERIES_PER_STUDY = 6\nDEV_EPOCHS_WEAK = 1\nDEV_EPOCHS_GOLD = 3\n\n# Full-run defaults\nFULL_EPOCHS_WEAK = 3\nFULL_EPOCHS_GOLD = 8\n\nIMAGE_SIZE = 224\nN_SLICES = 12\nMAX_SERIES_PER_SLOT = 1\nBATCH_SIZE = 4\nLR = 1e-4\nWEIGHT_DECAY = 1e-4\nNUM_WORKERS = 2\nPSEUDO_LOSS_WEIGHT = 0.35\n\nOUT = Path(\"/kaggle/working/rsna_knee_top10\")\nOUT.mkdir(parents=True, exist_ok=True)\n\nDEVICE = \"cuda\" if __import__(\"torch\").cuda.is_available() else \"cpu\"\n\nprint(\"DEVICE:\", DEVICE)\nprint(\"DEV_MODE:\", DEV_MODE)\nprint(\"OUT:\", OUT)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n# Phase 0 — Data discovery\n\nThe official competition structure is:\n\n`train.csv`, `train_series.csv`, `train_series/<StudyInstanceUID>/<SeriesInstanceUID>/<SOPInstanceUID>.dcm`, plus corresponding test files and `sample_submission.csv`.\n\nThe discovery code below does not assume a particular Kaggle dataset folder name.\n","metadata":{}},{"cell_type":"code","source":"\n# ============================================================\n# 0A. FIND COMPETITION FILES\n# ============================================================\nINPUT_ROOT = Path(\"/kaggle/input\")\n\ndef find_named(name):\n    return list(INPUT_ROOT.rglob(name))\n\ndef choose(name, required=True):\n    xs = find_named(name)\n    print(name, \"=>\", len(xs), \"candidate(s)\")\n    for x in xs[:10]:\n        print(\"   \", x)\n    if required and not xs:\n        raise FileNotFoundError(\n            f\"Không tìm thấy {name}. Hãy Add Data -> RSNA Knee Abnormality Detection.\"\n        )\n    return xs[0] if xs else None\n\nTRAIN_CSV = choose(\"train.csv\")\nTRAIN_SERIES_CSV = choose(\"train_series.csv\")\nTEST_CSV = choose(\"test.csv\")\nTEST_SERIES_CSV = choose(\"test_series.csv\")\nSAMPLE_SUB = choose(\"sample_submission.csv\", required=False)\n\ntrain = pd.read_csv(TRAIN_CSV)\ntrain_series = pd.read_csv(TRAIN_SERIES_CSV)\ntest = pd.read_csv(TEST_CSV)\ntest_series = pd.read_csv(TEST_SERIES_CSV)\n\nprint(train.shape, train_series.shape, test.shape, test_series.shape)\nprint(\"train columns:\", train.columns.tolist())\nprint(\"train_series columns:\", train_series.columns.tolist())\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# ============================================================\n# 0B. DATA INTEGRITY\n# ============================================================\nassert train[\"StudyInstanceUID\"].is_unique\nassert train_series[\"SeriesInstanceUID\"].is_unique\n\ncomplete_gold = train[LABELS].notna().all(axis=1)\ngold = train.loc[complete_gold].copy()\nunlabelled = train.loc[~complete_gold].copy()\n\nprint(\"Total studies:\", len(train))\nprint(\"Gold studies:\", len(gold))\nprint(\"Report-only studies:\", len(unlabelled))\nprint(\"Total series:\", len(train_series))\n\ndisplay(gold[LABELS].mean().sort_values(ascending=False).to_frame(\"prevalence\"))\n\n# The competition has an all-or-nothing label pattern.\nassert train.loc[complete_gold, LABELS].notna().all().all()\nassert train.loc[~complete_gold, LABELS].isna().all().all()\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n# Phase 1 — Index every series once\n\nDo not repeatedly recurse through 819k DICOM files during training. Build an index from the known competition directory layout. We use DICOM metadata only when necessary for geometry and pixel loading.\n","metadata":{}},{"cell_type":"code","source":"\n# ============================================================\n# 1A. DICOM ROOT RESOLUTION\n# ============================================================\n# Prefer the parent containing train_series.\ncandidate_train_series_dirs = []\nfor p in INPUT_ROOT.rglob(\"train_series\"):\n    if p.is_dir():\n        candidate_train_series_dirs.append(p)\n\nif not candidate_train_series_dirs:\n    raise FileNotFoundError(\"Cannot locate train_series directory.\")\n\nTRAIN_SERIES_ROOT = candidate_train_series_dirs[0]\nTEST_SERIES_ROOT = TRAIN_SERIES_ROOT.parent / \"test_series\"\n\nprint(\"TRAIN_SERIES_ROOT:\", TRAIN_SERIES_ROOT)\nprint(\"TEST_SERIES_ROOT :\", TEST_SERIES_ROOT)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# ============================================================\n# 1B. SERIES INDEX\n# ============================================================\nseries_index = train_series.copy()\nseries_index[\"StudyInstanceUID\"] = series_index[\"StudyInstanceUID\"].astype(str)\nseries_index[\"SeriesInstanceUID\"] = series_index[\"SeriesInstanceUID\"].astype(str)\n\nseries_index[\"series_dir\"] = [\n    str(TRAIN_SERIES_ROOT / study / sid)\n    for study, sid in zip(\n        series_index[\"StudyInstanceUID\"],\n        series_index[\"SeriesInstanceUID\"]\n    )\n]\n\nseries_index[\"exists\"] = series_index[\"series_dir\"].map(lambda x: Path(x).is_dir())\n\nprint(\"Series directories found:\",\n      int(series_index[\"exists\"].sum()), \"/\", len(series_index))\n\ndisplay(series_index.head())\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n# Phase 2 — Report teacher\n\nA central competition constraint is that reports exist for all training studies but **not for test studies**. Therefore reports should be used to create labels / priors during training, not as a test-time feature. \n\nThe weak-label representation deliberately uses:\n\n- `1.0` = confident positive\n- `0.0` = confident negative\n- `0.5` = unknown / not addressed / uncertain\n\nThis avoids converting \"not mentioned\" into a false negative.\n","metadata":{}},{"cell_type":"code","source":"\n# ============================================================\n# 2A. OPTIONAL ATTACHED LLM LABEL DATASET\n# ============================================================\n# If you attach a public report-label dataset to Kaggle, this cell\n# searches for CSVs containing StudyInstanceUID + target columns.\n# The notebook never requires it.\n\nllm_candidates = []\nfor csv in INPUT_ROOT.rglob(\"*.csv\"):\n    try:\n        cols = pd.read_csv(csv, nrows=0).columns.tolist()\n        if \"StudyInstanceUID\" in cols and sum(c in cols for c in LABELS) >= 8:\n            if csv not in [TRAIN_CSV, TRAIN_SERIES_CSV, TEST_CSV, TEST_SERIES_CSV]:\n                llm_candidates.append(csv)\n    except Exception:\n        pass\n\nprint(\"Potential external label files:\")\nfor p in llm_candidates[:20]:\n    print(\" \", p)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# ============================================================\n# 2B. CONSERVATIVE REPORT PARSER\n# ============================================================\nTERMS = {\n    \"ACL\": [r\"\\bacl\\b\", r\"anterior cruciate\"],\n    \"MCL\": [r\"\\bmcl\\b\", r\"medial collateral\"],\n    \"Medial Meniscus\": [r\"medial meniscus\"],\n    \"Lateral Meniscus\": [r\"lateral meniscus\"],\n    \"Medial OA\": [r\"medial.*osteoarthritis\", r\"medial compartment.*arthr\"],\n    \"Lateral OA\": [r\"lateral.*osteoarthritis\", r\"lateral compartment.*arthr\"],\n    \"PF OA\": [r\"patellofemoral\", r\"patello-femoral\"],\n    \"Effusion\": [r\"effusion\", r\"joint.*fluid\"],\n    \"Synovitis\": [r\"synovitis\", r\"synovial\"],\n    \"Baker's\": [r\"baker\", r\"popliteal cyst\"],\n    \"Contusion\": [r\"contusion\", r\"bone bruise\", r\"marrow edema\"],\n    \"Fracture\": [r\"fracture\", r\"fractured\"],\n}\n\nNEG = r\"(no|not|without|negative for|absence of|free of|does not demonstrate|there is no)\"\nUNCERTAIN = r\"(possible|possibly|may represent|cannot exclude|suspicious for|question of|suggestive of)\"\n\ndef parse_report(text):\n    text = str(text)\n    lower = text.lower()\n    out = {}\n    for label, patterns in TERMS.items():\n        pos = neg = unc = False\n        for pat in patterns:\n            for m in re.finditer(pat, lower):\n                ctx = lower[max(0, m.start()-100): min(len(lower), m.end()+70)]\n                if re.search(UNCERTAIN, ctx):\n                    unc = True\n                elif re.search(NEG, ctx):\n                    neg = True\n                else:\n                    pos = True\n        if pos and not unc:\n            out[label] = 1.0\n        elif neg and not pos and not unc:\n            out[label] = 0.0\n        else:\n            out[label] = 0.5\n    return out\n\nweak = []\nfor _, r in train.iterrows():\n    parsed = parse_report(r.get(\"Report\", \"\"))\n    weak.append({\"StudyInstanceUID\": str(r[\"StudyInstanceUID\"]), **parsed})\n\nweak = pd.DataFrame(weak)\n\n# If an attached label file exists, prefer the first one with the full label set.\nllm_teacher = None\nfor p in llm_candidates:\n    try:\n        tmp = pd.read_csv(p)\n        if all(c in tmp.columns for c in [\"StudyInstanceUID\"] + LABELS):\n            llm_teacher = tmp[[\"StudyInstanceUID\"] + LABELS].copy()\n            llm_teacher[\"StudyInstanceUID\"] = llm_teacher[\"StudyInstanceUID\"].astype(str)\n            print(\"Using external weak-label teacher:\", p)\n            break\n    except Exception:\n        pass\n\nif llm_teacher is not None:\n    weak = llm_teacher.copy()\n    for c in LABELS:\n        weak[c] = pd.to_numeric(weak[c], errors=\"coerce\").clip(0, 1)\n\ndisplay(weak.head())\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# ============================================================\n# 2C. WEAK LABEL AUDIT AGAINST GOLD\n# ============================================================\nfrom sklearn.metrics import roc_auc_score\n\ngold_map = gold.set_index(gold[\"StudyInstanceUID\"].astype(str))\nweak_map = weak.set_index(\"StudyInstanceUID\")\n\nweak_auc = []\nfor label in LABELS:\n    ids = [sid for sid in gold_map.index if sid in weak_map.index]\n    y = gold_map.loc[ids, label].astype(float).values\n    p = weak_map.loc[ids, label].astype(float).values\n    weak_auc.append({\n        \"label\": label,\n        \"AUC_vs_gold\": roc_auc_score(y, p) if len(np.unique(y)) == 2 else np.nan,\n        \"mean_weak\": np.mean(p),\n    })\n\nweak_auc = pd.DataFrame(weak_auc)\ndisplay(weak_auc)\nprint(\"Weak-label macro AUC:\",\n      float(np.nanmean(weak_auc[\"AUC_vs_gold\"])))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n# Phase 3 — DICOM geometry and 2.5D preprocessing\n\nThe competition explicitly notes that orientations, intensities and resolutions vary, with multiple transfer syntaxes. \nWe therefore:\n- sort by `ImagePositionPatient` projected onto the slice normal,\n- use `PixelSpacing`,\n- normalize robustly,\n- resize to a fixed square,\n- form adjacent-slice triplets for a 2.5D CNN.\n\nThe cache is **not** a full 570 GB tensor dump. It stores compact preprocessed slices for only the selected series.\n","metadata":{}},{"cell_type":"code","source":"\n# ============================================================\n# 3A. DICOM FUNCTIONS\n# ============================================================\nimport pydicom\nfrom PIL import Image\nfrom tqdm.auto import tqdm\n\ndef get_dicom_paths(series_dir):\n    return sorted(Path(series_dir).glob(\"*.dcm\"))\n\ndef read_meta(path):\n    return pydicom.dcmread(str(path), stop_before_pixels=True, force=True)\n\ndef physical_sort(paths):\n    rows = []\n    for p in paths:\n        ds = read_meta(p)\n        iop = getattr(ds, \"ImageOrientationPatient\", None)\n        ipp = getattr(ds, \"ImagePositionPatient\", None)\n        key = None\n        if iop is not None and ipp is not None and len(iop) >= 6 and len(ipp) >= 3:\n            row = np.asarray(iop[:3], dtype=float)\n            col = np.asarray(iop[3:6], dtype=float)\n            normal = np.cross(row, col)\n            n = np.linalg.norm(normal)\n            if n > 1e-8:\n                normal /= n\n                key = float(np.dot(np.asarray(ipp[:3], dtype=float), normal))\n        if key is None:\n            key = float(getattr(ds, \"InstanceNumber\", 0))\n        rows.append((key, p))\n    rows.sort(key=lambda x: x[0])\n    return [p for _, p in rows]\n\ndef norm_u8(arr):\n    arr = arr.astype(np.float32)\n    finite = np.isfinite(arr)\n    if not finite.any():\n        return np.zeros(arr.shape, np.uint8)\n    x = arr[finite]\n    lo, hi = np.percentile(x, [1, 99])\n    if hi <= lo:\n        hi = lo + 1\n    arr = np.clip((arr-lo)/(hi-lo), 0, 1)\n    return (arr*255).astype(np.uint8)\n\ndef crop_resize(arr, spacing, size=IMAGE_SIZE, crop_mm=140):\n    sy, sx = float(spacing[0]), float(spacing[1])\n    h, w = arr.shape\n    ch = min(h, max(8, int(round(crop_mm/sy))))\n    cw = min(w, max(8, int(round(crop_mm/sx))))\n    y0, x0 = (h-ch)//2, (w-cw)//2\n    arr = arr[y0:y0+ch, x0:x0+cw]\n    return np.asarray(\n        Image.fromarray(norm_u8(arr)).resize((size, size), Image.Resampling.BILINEAR),\n        dtype=np.uint8\n    )\n\ndef load_series_volume(series_dir):\n    paths = physical_sort(get_dicom_paths(series_dir))\n    frames = []\n    for p in paths:\n        ds = pydicom.dcmread(str(p), force=True)\n        arr = ds.pixel_array.astype(np.float32)\n        spacing = getattr(ds, \"PixelSpacing\", [1.0, 1.0])\n        frames.append(crop_resize(arr, spacing))\n    return np.stack(frames)\n\ndef make_triplets(volume, n=N_SLICES):\n    # Uniform centers; each item is [prev, center, next].\n    L = len(volume)\n    centers = np.linspace(0, L-1, n).astype(int)\n    out = []\n    for c in centers:\n        a = volume[max(0,c-1)]\n        b = volume[c]\n        d = volume[min(L-1,c+1)]\n        out.append(np.stack([a,b,d], axis=0))\n    return np.stack(out)  # [N,3,H,W]\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n# Phase 4 — Multi-series slot selection\n\nRather than treating all series equally, we map each study to up to six slots:\n\n1. sagittal fluid-sensitive\n2. sagittal non-fluid\n3. coronal fluid-sensitive\n4. coronal non-fluid\n5. axial fluid-sensitive\n6. axial non-fluid\n\nThis preserves acquisition metadata and gives the model a consistent study-level structure.\n","metadata":{}},{"cell_type":"code","source":"\n# ============================================================\n# 4A. SLOT ASSIGNMENT\n# ============================================================\ndef slot_name(row):\n    plane = str(row[\"Anatomical_Plane\"]).strip().lower()\n    fluid = int(row[\"Fluid_Sensitive\"]) if pd.notna(row[\"Fluid_Sensitive\"]) else 0\n    if plane.startswith(\"sag\"):\n        p = \"sagittal\"\n    elif plane.startswith(\"cor\"):\n        p = \"coronal\"\n    elif plane.startswith(\"axi\"):\n        p = \"axial\"\n    else:\n        p = \"other\"\n    return f\"{p}_{'fluid' if fluid else 'nonfluid'}\"\n\nseries_index[\"slot\"] = series_index.apply(slot_name, axis=1)\n\nslot_priority = (\n    series_index\n    .groupby([\"StudyInstanceUID\",\"slot\"])\n    .size()\n    .rename(\"n_series\")\n    .reset_index()\n)\n\n# Keep one representative series per slot. If multiple exist,\n# choose fat-suppressed/fluid-sensitive and then the first stable ID.\nseries_index[\"_score\"] = (\n    series_index[\"Fluid_Sensitive\"].fillna(0).astype(float) * 2\n    + series_index[\"Fat_Suppression\"].fillna(0).astype(float)\n)\n\nseries_selected = (\n    series_index\n    .sort_values([\"StudyInstanceUID\",\"slot\",\"_score\",\"SeriesInstanceUID\"],\n                 ascending=[True,True,False,True])\n    .groupby([\"StudyInstanceUID\",\"slot\"], as_index=False)\n    .head(MAX_SERIES_PER_SLOT)\n    .copy()\n)\n\nprint(\"Selected series:\", len(series_selected))\ndisplay(series_selected[[\n    \"StudyInstanceUID\",\"SeriesInstanceUID\",\n    \"slot\",\"Fluid_Sensitive\",\"Fat_Suppression\",\"Anatomical_Plane\"\n]].head(20))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n# Phase 5 — Compact cache\n\nFor a development run, cache only a small set of studies. For the full run, cache the selected six slots per study. This is much smaller than caching every DICOM slice from all 24k series.\n","metadata":{}},{"cell_type":"code","source":"\n# ============================================================\n# 5A. SELECT DEVELOPMENT STUDIES\n# ============================================================\nall_studies = train[\"StudyInstanceUID\"].astype(str).tolist()\ngold_ids = set(gold[\"StudyInstanceUID\"].astype(str))\n\nif DEV_MODE:\n    # Include all gold studies plus a sample of report-only studies.\n    extra = [x for x in all_studies if x not in gold_ids][:max(0, DEV_STUDIES-len(gold_ids))]\n    selected_studies = list(gold_ids) + extra\nelse:\n    selected_studies = all_studies\n\nselected_studies = set(map(str, selected_studies))\nprint(\"Studies selected for cache:\", len(selected_studies))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# ============================================================\n# 5B. BUILD CACHE\n# ============================================================\nCACHE = OUT / \"cache\"\nCACHE.mkdir(exist_ok=True)\n\nmanifest_rows = []\n\nfor _, r in tqdm(\n    series_selected[series_selected[\"StudyInstanceUID\"].isin(selected_studies)].iterrows(),\n    total=int(series_selected[\"StudyInstanceUID\"].isin(selected_studies).sum()),\n    desc=\"Preprocessing selected series\"\n):\n    study = str(r[\"StudyInstanceUID\"])\n    sid = str(r[\"SeriesInstanceUID\"])\n    sdir = Path(r[\"series_dir\"])\n    if not sdir.exists():\n        continue\n\n    out = CACHE / f\"{study}__{sid}.npy\"\n    if not out.exists():\n        try:\n            vol = load_series_volume(sdir)\n            np.save(out, vol, allow_pickle=False)\n        except Exception as e:\n            manifest_rows.append({\n                \"StudyInstanceUID\": study,\n                \"SeriesInstanceUID\": sid,\n                \"slot\": r[\"slot\"],\n                \"error\": repr(e)\n            })\n            continue\n\n    manifest_rows.append({\n        \"StudyInstanceUID\": study,\n        \"SeriesInstanceUID\": sid,\n        \"slot\": r[\"slot\"],\n        \"path\": str(out)\n    })\n\ncache_manifest = pd.DataFrame(manifest_rows)\ncache_manifest.to_csv(OUT/\"cache_manifest.csv\", index=False)\n\nprint(\"Cached series:\", len(cache_manifest))\ndisplay(cache_manifest.head())\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n# Phase 6 — Strong 2.5D multi-series model\n\nThe model has three levels:\n\n`DICOM slices → slice encoder → attention over slices → slot embeddings → masked study fusion → 12 labels`\n\nThis is substantially closer to the competition structure than the earlier 12-channel CNN.\n","metadata":{}},{"cell_type":"code","source":"\n# ============================================================\n# 6A. PYTORCH + BACKBONE\n# ============================================================\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torch.utils.data import Dataset, DataLoader\n\ntry:\n    import timm\n    HAS_TIMM = True\nexcept Exception:\n    HAS_TIMM = False\n\nprint(\"timm:\", HAS_TIMM, \"device:\", DEVICE)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# ============================================================\n# 6B. DATASET\n# ============================================================\nSLOTS = [\n    \"sagittal_fluid\", \"sagittal_nonfluid\",\n    \"coronal_fluid\", \"coronal_nonfluid\",\n    \"axial_fluid\", \"axial_nonfluid\"\n]\n\nclass StudyDataset(Dataset):\n    def __init__(self, studies, labels_df=None, cache_manifest=None):\n        self.studies = [str(x) for x in studies]\n        self.labels_df = labels_df.set_index(labels_df[\"StudyInstanceUID\"].astype(str)) if labels_df is not None else None\n\n        cm = cache_manifest.copy()\n        cm[\"StudyInstanceUID\"] = cm[\"StudyInstanceUID\"].astype(str)\n        self.paths = {\n            (str(r[\"StudyInstanceUID\"]), str(r[\"slot\"])): r[\"path\"]\n            for _, r in cm.dropna(subset=[\"path\"]).iterrows()\n        }\n\n    def __len__(self):\n        return len(self.studies)\n\n    def __getitem__(self, idx):\n        study = self.studies[idx]\n        slots = []\n        mask = []\n\n        for slot in SLOTS:\n            p = self.paths.get((study, slot))\n            if p is None:\n                slots.append(np.zeros((N_SLICES,3,IMAGE_SIZE,IMAGE_SIZE), np.uint8))\n                mask.append(0.0)\n            else:\n                vol = np.load(p, mmap_mode=\"r\")\n                trip = make_triplets(vol, N_SLICES)\n                slots.append(trip)\n                mask.append(1.0)\n\n        x = np.stack(slots).astype(np.float32) / 255.0\n        x = torch.from_numpy(x)\n        mask = torch.tensor(mask, dtype=torch.float32)\n\n        if self.labels_df is None:\n            return x, mask, study\n\n        y = torch.tensor(\n            self.labels_df.loc[study, LABELS].astype(float).values,\n            dtype=torch.float32\n        )\n        return x, mask, y, study\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# ============================================================\n# 6C. MODEL\n# ============================================================\nclass SliceBackbone(nn.Module):\n    def __init__(self, emb=256):\n        super().__init__()\n        if HAS_TIMM:\n            self.backbone = timm.create_model(\n                \"resnet18\",\n                pretrained=True,\n                num_classes=0,\n                in_chans=3\n            )\n            d = self.backbone.num_features\n        else:\n            from torchvision.models import resnet18\n            self.backbone = resnet18(weights=None)\n            d = self.backbone.fc.in_features\n            self.backbone.fc = nn.Identity()\n        self.proj = nn.Linear(d, emb)\n        self.norm = nn.LayerNorm(emb)\n\n    def forward(self, x):\n        z = self.backbone(x)\n        return self.norm(self.proj(z))\n\nclass AttentionPool(nn.Module):\n    def __init__(self, dim):\n        super().__init__()\n        self.score = nn.Sequential(\n            nn.Linear(dim, dim//2),\n            nn.Tanh(),\n            nn.Linear(dim//2, 1)\n        )\n\n    def forward(self, x, mask=None):\n        # x [B,N,D]\n        s = self.score(x).squeeze(-1)\n        if mask is not None:\n            s = s.masked_fill(~mask.bool(), -1e4)\n        a = torch.softmax(s, dim=1)\n        return (x * a.unsqueeze(-1)).sum(1), a\n\nclass MultiSeriesKneeModel(nn.Module):\n    def __init__(self, n_classes=12, emb=256):\n        super().__init__()\n\n        self.slice_encoder = SliceBackbone(emb)\n        self.slice_pool = AttentionPool(emb)\n\n        self.slot_embedding = nn.Parameter(\n            torch.zeros(len(SLOTS), emb)\n        )\n\n        self.slot_pool = AttentionPool(emb)\n\n        self.head = nn.Sequential(\n            nn.LayerNorm(emb),\n            nn.Linear(emb, n_classes)\n        )\n\n    def forward(self, x, slot_mask):\n        \"\"\"\n        x:\n            [B, S, N, C, H, W]\n\n            B = batch size\n            S = number of series slots = 6\n            N = number of slices\n            C = 3 for 2.5D triplet\n            H,W = image size\n\n        slot_mask:\n            [B, S]\n        \"\"\"\n\n        # --------------------------------------------------\n        # 1. Read input shape\n        # --------------------------------------------------\n        B, S, N, C, H, W = x.shape\n\n        if S != len(SLOTS):\n            raise RuntimeError(\n                f\"Expected {len(SLOTS)} slots, got S={S}\"\n            )\n\n        # --------------------------------------------------\n        # 2. Encode every slice\n        # --------------------------------------------------\n        # [B,S,N,C,H,W]\n        #       ↓\n        # [B*S*N,C,H,W]\n        x = x.reshape(B * S * N, C, H, W)\n\n        z = self.slice_encoder(x)\n\n        # --------------------------------------------------\n        # 3. Restore study / series / slice dimensions\n        # --------------------------------------------------\n        # [B*S*N,D]\n        #       ↓\n        # [B,S,N,D]\n        D = z.shape[-1]\n\n        z = z.reshape(B, S, N, D)\n\n        # --------------------------------------------------\n        # 4. Pool slices inside each MRI series\n        # --------------------------------------------------\n        series_z = []\n\n        for s in range(S):\n\n            # z[:, s]\n            # [B,N,D]\n\n            zz, _ = self.slice_pool(\n                z[:, s, :, :]\n            )\n\n            # [B,D]\n            series_z.append(zz)\n\n        # [B,S,D]\n        series_z = torch.stack(\n            series_z,\n            dim=1\n        )\n\n        # --------------------------------------------------\n        # 5. Add series/plane identity embedding\n        # --------------------------------------------------\n        series_z = (\n            series_z\n            + self.slot_embedding.unsqueeze(0)\n        )\n\n        # --------------------------------------------------\n        # 6. Validate slot mask\n        # --------------------------------------------------\n        slot_mask = slot_mask.to(\n            device=series_z.device,\n            dtype=torch.float32\n        )\n\n        expected_shape = (B, S)\n\n        if slot_mask.ndim != 2:\n            raise RuntimeError(\n                f\"slot_mask must be 2D [B,S], \"\n                f\"got shape={tuple(slot_mask.shape)}\"\n            )\n\n        if tuple(slot_mask.shape) != expected_shape:\n            raise RuntimeError(\n                f\"slot_mask shape={tuple(slot_mask.shape)}, \"\n                f\"expected={expected_shape}\"\n            )\n\n        # --------------------------------------------------\n        # 7. Pool the six MRI series\n        # --------------------------------------------------\n        # [B,S,D] + [B,S]\n        #       ↓\n        # [B,D]\n        study_z, _ = self.slot_pool(\n            series_z,\n            slot_mask\n        )\n\n        # --------------------------------------------------\n        # 8. Multi-label prediction\n        # --------------------------------------------------\n        # [B,D] → [B,12]\n        logits = self.head(study_z)\n\n        return logits\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n# Phase 7 — Gold / weak-label training\n\nBecause only 58 studies have expert labels, the main leverage is to use the report-derived labels to pretrain the MRI representation, then fine-tune on gold.\n\nWe do **not** treat weak labels as equal to expert labels.\n\nFor weak training:\n- `0.5` means unknown and contributes no loss.\n- confident weak cells are weighted by `PSEUDO_LOSS_WEIGHT`.\n\nFor gold training:\n- all 12 expert targets receive full weight.\n","metadata":{}},{"cell_type":"code","source":"\n# ============================================================\n# 7A. MASKED BCE\n# ============================================================\ndef masked_bce(logits, target, weight=None):\n    known = (target != 0.5).float()\n    target01 = target.clamp(0,1)\n    loss = F.binary_cross_entropy_with_logits(logits, target01, reduction=\"none\")\n    if weight is not None:\n        loss = loss * weight\n    loss = loss * known\n    denom = known.sum().clamp_min(1.0)\n    return loss.sum() / denom\n\ndef auc_score_matrix(y, p):\n    vals = []\n    for j in range(y.shape[1]):\n        if len(np.unique(y[:,j])) >= 2:\n            vals.append(roc_auc_score(y[:,j], p[:,j]))\n        else:\n            vals.append(np.nan)\n    return np.asarray(vals), float(np.nanmean(vals))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# ============================================================\n# 7B. GENERIC TRAIN / PREDICT\n# ============================================================\nfrom sklearn.metrics import roc_auc_score\n\ndef train_model(train_studies, train_labels, val_studies, val_labels,\n                epochs, weak=False):\n    tr_ds = StudyDataset(train_studies, train_labels, cache_manifest)\n    va_ds = StudyDataset(val_studies, val_labels, cache_manifest)\n\n    tr_loader = DataLoader(tr_ds, batch_size=BATCH_SIZE, shuffle=True,\n                           num_workers=NUM_WORKERS, pin_memory=True)\n    va_loader = DataLoader(va_ds, batch_size=BATCH_SIZE, shuffle=False,\n                           num_workers=NUM_WORKERS, pin_memory=True)\n\n    model = MultiSeriesKneeModel().to(DEVICE)\n    opt = torch.optim.AdamW(model.parameters(), lr=LR, weight_decay=WEIGHT_DECAY)\n    scaler = torch.cuda.amp.GradScaler(enabled=(DEVICE==\"cuda\"))\n\n    best = -np.inf\n    best_state = None\n\n    for epoch in range(epochs):\n        model.train()\n        losses = []\n        for batch in tr_loader:\n            if weak:\n                x, sm, y, ids = batch\n                y = y.to(DEVICE)\n            else:\n                x, sm, y, ids = batch\n                y = y.to(DEVICE)\n\n            x = x.to(DEVICE, non_blocking=True)\n            sm = sm.to(DEVICE, non_blocking=True)\n            opt.zero_grad(set_to_none=True)\n\n            with torch.cuda.amp.autocast(enabled=(DEVICE==\"cuda\")):\n                logits = model(x, sm)\n                loss = masked_bce(\n                    logits, y,\n                    None if not weak else torch.full_like(y, PSEUDO_LOSS_WEIGHT)\n                )\n\n            scaler.scale(loss).backward()\n            scaler.step(opt)\n            scaler.update()\n            losses.append(float(loss.detach().cpu()))\n\n        # Validation\n        model.eval()\n        ys, ps = [], []\n        with torch.no_grad():\n            for x, sm, y, ids in va_loader:\n                x = x.to(DEVICE)\n                sm = sm.to(DEVICE)\n                p = torch.sigmoid(model(x, sm)).cpu().numpy()\n                ys.append(y.numpy())\n                ps.append(p)\n\n        yv = np.concatenate(ys)\n        pv = np.concatenate(ps)\n        aucs, macro = auc_score_matrix(yv, pv)\n\n        print(f\"epoch {epoch+1}/{epochs} loss={np.mean(losses):.4f} macro_auc={macro:.4f}\")\n        if macro > best:\n            best = macro\n            best_state = {k:v.detach().cpu().clone() for k,v in model.state_dict().items()}\n\n    model.load_state_dict(best_state)\n    return model, best\n\ndef predict_model(model, studies, labels_df=None):\n    ds = StudyDataset(studies, labels_df, cache_manifest)\n    loader = DataLoader(ds, batch_size=BATCH_SIZE, shuffle=False,\n                        num_workers=NUM_WORKERS)\n    model.eval()\n    preds, ids = [], []\n    with torch.no_grad():\n        for batch in loader:\n            x, sm = batch[0].to(DEVICE), batch[1].to(DEVICE)\n            p = torch.sigmoid(model(x, sm)).cpu().numpy()\n            preds.append(p)\n            ids.extend(batch[-1])\n    return np.concatenate(preds), [str(x) for x in ids]\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n# Phase 8 — 5-fold OOF gold evaluation\n\nThe official metric is per-study probability for each of the twelve findings; macro-AUC is therefore the correct model-selection target. We use study-level folds so no series from the same study crosses train/validation.\n","metadata":{}},{"cell_type":"code","source":"\n# ============================================================\n# 8A. BUILD FOLDS\n# ============================================================\nfrom sklearn.model_selection import KFold\n\ngold2 = gold.copy()\ngold2[\"StudyInstanceUID\"] = gold2[\"StudyInstanceUID\"].astype(str)\n\nkf = KFold(N_FOLDS, shuffle=True, random_state=SEED)\nfold_map = {}\nfor f, (_, va) in enumerate(kf.split(gold2)):\n    for sid in gold2.iloc[va][\"StudyInstanceUID\"]:\n        fold_map[str(sid)] = f\ngold2[\"fold\"] = gold2[\"StudyInstanceUID\"].map(fold_map)\n\ndisplay(gold2[\"fold\"].value_counts().sort_index())\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# ============================================================\n# 8B. GOLD OOF\n# ============================================================\ngold_ids_available = set(cache_manifest[\"StudyInstanceUID\"].astype(str))\ngold2_run = gold2[gold2[\"StudyInstanceUID\"].isin(gold_ids_available)].copy()\n\noof_gold = pd.DataFrame(\n    np.nan,\n    index=gold2[\"StudyInstanceUID\"],\n    columns=LABELS\n)\n\nfold_results = []\n\nfor fold in range(N_FOLDS):\n    tr = gold2_run[gold2_run[\"fold\"] != fold]\n    va = gold2_run[gold2_run[\"fold\"] == fold]\n    if len(tr) == 0 or len(va) == 0:\n        continue\n\n    print(\"\\n========== FOLD\", fold, \"==========\")\n    model, score = train_model(\n        tr[\"StudyInstanceUID\"].tolist(), tr,\n        va[\"StudyInstanceUID\"].tolist(), va,\n        epochs=DEV_EPOCHS_GOLD if DEV_MODE else FULL_EPOCHS_GOLD,\n        weak=False\n    )\n\n    pred, ids = predict_model(model, va, va)\n    for sid, p in zip(ids, pred):\n        oof_gold.loc[sid, LABELS] = p\n\n    fold_results.append({\"fold\": fold, \"macro_auc\": score})\n    del model\n    gc.collect()\n    if torch.cuda.is_available():\n        torch.cuda.empty_cache()\n\ndisplay(pd.DataFrame(fold_results))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# ============================================================\n# 8C. GOLD OOF METRICS\n# ============================================================\nmask = ~oof_gold.isna().all(axis=1)\nif mask.any():\n    y = gold2.set_index(\"StudyInstanceUID\").loc[oof_gold.index[mask], LABELS].values.astype(float)\n    p = oof_gold.loc[mask, LABELS].values.astype(float)\n    aucs, macro = auc_score_matrix(y,p)\n    gold_metrics = pd.DataFrame({\"label\":LABELS, \"auc\":aucs})\n    print(\"GOLD OOF MACRO-AUC:\", macro)\n    display(gold_metrics.sort_values(\"auc\"))\nelse:\n    gold_metrics = pd.DataFrame({\"label\":LABELS, \"auc\":np.nan})\n    print(\"No OOF predictions.\")\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n# Phase 9 — Weak-label MRI pretraining\n\nThis phase is optional in `DEV_MODE`. In the full run it is the key semi-supervised step:\n\n`4,407 reports → weak targets → MRI pretraining → 58-study gold fine-tuning`.\n\nThe model learns from the much larger training population without pretending those labels are expert annotations.\n","metadata":{}},{"cell_type":"code","source":"\n# ============================================================\n# 9A. WEAK TRAINING TABLE\n# ============================================================\nweak_train = weak.copy()\nweak_train[\"StudyInstanceUID\"] = weak_train[\"StudyInstanceUID\"].astype(str)\n\nweak_studies = [s for s in selected_studies if s in set(weak_train[\"StudyInstanceUID\"])]\nweak_df = weak_train.set_index(\"StudyInstanceUID\").loc[weak_studies].reset_index()\n\nif DEV_MODE:\n    weak_df = weak_df.head(min(DEV_STUDIES, len(weak_df)))\n\nprint(\"Weak studies:\", len(weak_df))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# ============================================================\n# 9B. WEAK PRETRAINING\n# ============================================================\nweak_model = None\n\nif len(weak_df) >= 10:\n    # Hold out a small portion only to detect catastrophic failure.\n    cut = max(1, int(len(weak_df)*0.1))\n    weak_val = weak_df.iloc[:cut].copy()\n    weak_tr = weak_df.iloc[cut:].copy()\n\n    weak_model, _ = train_model(\n        weak_tr[\"StudyInstanceUID\"].tolist(), weak_tr,\n        weak_val[\"StudyInstanceUID\"].tolist(), weak_val,\n        epochs=DEV_EPOCHS_WEAK if DEV_MODE else FULL_EPOCHS_WEAK,\n        weak=True\n    )\n\n    torch.save(weak_model.state_dict(), OUT/\"weak_pretrained.pt\")\nelse:\n    print(\"Not enough cached weak-label studies for pretraining.\")\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n# Phase 10 — OOF blend / prior fusion\n\nA report teacher is available at training time but not test time. Its role is therefore to **shape the MRI model**, and optionally provide a calibrated prior for OOF analysis.\n\nWe optimize the blend only on gold OOF data.\n\nFor AUC, monotonic transformations do not matter, but relative ranking does. We therefore search a small grid of blend weights rather than using an arbitrary 50/50 blend.\n","metadata":{}},{"cell_type":"code","source":"\n# ============================================================\n# 10A. OOF BLEND\n# ============================================================\nif mask.any():\n    ids = list(oof_gold.index[mask])\n    y = gold2.set_index(\"StudyInstanceUID\").loc[ids, LABELS].values.astype(float)\n    p_mri = oof_gold.loc[ids, LABELS].values.astype(float)\n    p_report = weak_map.loc[ids, LABELS].values.astype(float)\n\n    candidates = []\n    for w in np.linspace(0,1,21):\n        p = w*p_mri + (1-w)*p_report\n        _, macro = auc_score_matrix(y,p)\n        candidates.append((w,macro))\n\n    best_w, best_macro = max(candidates, key=lambda x:x[1])\n    print(\"Best OOF MRI weight:\", best_w)\n    print(\"Blended OOF macro-AUC:\", best_macro)\n    print(\"MRI-only OOF macro-AUC:\",\n          auc_score_matrix(y,p_mri)[1])\n    print(\"Report-only OOF macro-AUC:\",\n          auc_score_matrix(y,p_report)[1])\n\n    blend_curve = pd.DataFrame(candidates, columns=[\"mri_weight\",\"macro_auc\"])\n    display(blend_curve.sort_values(\"macro_auc\", ascending=False).head(10))\nelse:\n    best_w = 1.0\n    print(\"No OOF blend available.\")\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n# Phase 11 — Model ablations \n\nRun these as separate experiments and keep every OOF result:\n\n1. **Gold-only MRI**\n2. **Weak-pretrained → gold fine-tune**\n3. **Different series-slot policies**\n4. **Different slice counts**\n5. **Different image sizes**\n6. **Different backbones**\n7. **Different random seeds**\n8. **OOF ensemble**\n\nDo not pick a model from a single public leaderboard submission. Pick it from OOF, then use the leaderboard only as a secondary diagnostic.\n","metadata":{}},{"cell_type":"code","source":"\n# ============================================================\n# 11A. EXPERIMENT REGISTRY\n# ============================================================\nexperiment_registry = pd.DataFrame([\n    {\"experiment\":\"gold_only_2p5d_multiseries\",\"status\":\"implemented\"},\n    {\"experiment\":\"weak_pretrain_then_gold\",\"status\":\"implemented\"},\n    {\"experiment\":\"report_prior_blend\",\"status\":\"implemented\"},\n    {\"experiment\":\"multi_seed_ensemble\",\"status\":\"planned\"},\n    {\"experiment\":\"backbone_sweep\",\"status\":\"planned\"},\n    {\"experiment\":\"slice_count_sweep\",\"status\":\"planned\"},\n    {\"experiment\":\"slot_policy_sweep\",\"status\":\"planned\"},\n])\nexperiment_registry.to_csv(OUT/\"experiment_registry.csv\", index=False)\ndisplay(experiment_registry)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n# Phase 12 — Full-data training and test inference\n\nThis section is deliberately protected by `RUN_FINAL = False`.\n\nOnce the OOF-selected configuration is frozen:\n- cache the selected test series,\n- train the chosen MRI model on all 4,407 studies using the chosen weak→gold schedule,\n- infer all test studies,\n- validate the exact submission schema,\n- save `submission.csv`.\n\nTest reports are not available, so no report feature is used at inference time. \n","metadata":{}},{"cell_type":"code","source":"\n# ============================================================\n# 12A. TEST CACHE\n# ============================================================\nRUN_FINAL = False\n\ndef cache_test_series():\n    rows = []\n    for _, r in tqdm(test_series.iterrows(), total=len(test_series), desc=\"Caching test\"):\n        study = str(r[\"StudyInstanceUID\"])\n        sid = str(r[\"SeriesInstanceUID\"])\n        slot = slot_name(r)\n        sdir = TEST_SERIES_ROOT / study / sid\n        if not sdir.exists():\n            continue\n        out = OUT/\"test_cache\"/f\"{study}__{sid}.npy\"\n        out.parent.mkdir(exist_ok=True)\n        if not out.exists():\n            try:\n                np.save(out, load_series_volume(sdir), allow_pickle=False)\n            except Exception:\n                continue\n        rows.append({\"StudyInstanceUID\":study,\"SeriesInstanceUID\":sid,\n                     \"slot\":slot,\"path\":str(out)})\n    return pd.DataFrame(rows)\n\nif RUN_FINAL:\n    test_cache_manifest = cache_test_series()\n    test_cache_manifest.to_csv(OUT/\"test_cache_manifest.csv\", index=False)\n    print(\"Test cached:\", len(test_cache_manifest))\nelse:\n    print(\"RUN_FINAL=False — test cache skipped.\")\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# ============================================================\n# 12B. SUBMISSION VALIDATION\n# ============================================================\ndef validate_submission(sub):\n    required = [\"StudyInstanceUID\"] + LABELS\n    missing = [c for c in required if c not in sub.columns]\n    if missing:\n        raise ValueError(\"Missing columns: \" + str(missing))\n\n    if len(sub) != len(test):\n        raise ValueError(f\"Row count {len(sub)} != test {len(test)}\")\n\n    if not sub[\"StudyInstanceUID\"].astype(str).equals(\n        test[\"StudyInstanceUID\"].astype(str).reset_index(drop=True)\n    ):\n        raise ValueError(\"StudyInstanceUID order does not match test.csv\")\n\n    vals = sub[LABELS].to_numpy(dtype=float)\n    if not np.isfinite(vals).all():\n        raise ValueError(\"NaN/Inf predictions found\")\n\n    if ((vals < 0) | (vals > 1)).any():\n        raise ValueError(\"Predictions outside [0,1]\")\n\n    return True\n\nif RUN_FINAL:\n    # Replace this placeholder with final model predictions.\n    final_pred = np.full((len(test), len(LABELS)), 0.5, dtype=np.float32)\n    submission = pd.DataFrame(final_pred, columns=LABELS)\n    submission.insert(0, \"StudyInstanceUID\", test[\"StudyInstanceUID\"].values)\n    validate_submission(submission)\n    submission.to_csv(OUT/\"submission.csv\", index=False)\n    print(\"Saved:\", OUT/\"submission.csv\")\nelse:\n    print(\"Final submission disabled. Set RUN_FINAL=True after freezing OOF-selected configuration.\")\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n# Phase 13 — Final report\n\nArtifacts are written to `/kaggle/working/rsna_knee_top10`.\n\n## Required decision gates\n\n### Gate A — Data\n- all series directories resolve\n- DICOM geometry is valid\n- visual QC passes\n\n### Gate B — Weak labels\n- report teacher AUC vs 58 gold studies is measured\n- uncertain / unmentioned cells are not forced negative\n- external LLM labels are preferred when attached and validated\n\n### Gate C — MRI\n- 5-fold study-level OOF\n- per-label AUC\n- macro-AUC\n- no leakage\n\n### Gate D — Final\n- freeze architecture\n- freeze preprocessing\n- freeze ensemble weights\n- train full data\n- infer test\n- validate submission\n\n## What this notebook changes versus the previous version\n\n- Uses the actual 58/4,407 semi-supervised structure.\n- Adds report weak-label teacher and optional attached LLM-label detection.\n- Adds multi-series six-slot representation.\n- Adds 2.5D adjacent-slice triplets.\n- Adds slice attention.\n- Adds study-level fusion.\n- Adds weak pretraining → gold fine-tuning path.\n- Adds OOF report/MRI blend search.\n- Keeps AUC at fold, per-label, and macro levels.\n- Protects final submission from accidental development runs.\n\n","metadata":{}}]}