{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# RSNA Knee — Notebook 01: Baseline / Supervision / Fold Audit\n\nThis notebook is the **first preparation stage** for continual fine-tuning of the existing ~0.89 DINOv2 ensemble.\n\nIt intentionally **does not fine-tune yet**. The uploaded inference notebook has a 20-member manifest checkpoint package, but it does not contain the original training loop or a verified `StudyInstanceUID -> fold` mapping. Re-creating folds arbitrarily and calling them the original folds could cause leakage.\n\n### What this notebook does\n\n1. Finds the RSNA Knee training data and the existing DINOv2 weight package.\n2. Audits `train.csv` labels while treating missing labels as **unknown**, not negative.\n3. Reads the existing `manifest.json` and builds a checkpoint/member registry.\n4. Inspects checkpoint structure without loading all checkpoints into memory at once.\n5. Searches for a **real original fold-map artifact**; it never invents one silently.\n6. Writes masked-label tables and expert subsets for later fine-tuning.\n7. Creates a compact `stage1_bundle.zip` that becomes an input to Notebook 02.\n\n### Important rule\n\nIf no original study-to-fold mapping is found, Notebook 02 must **not** pretend the checkpoint fold IDs define validation studies. We can still fine-tune later, but validation must be designed explicitly.","metadata":{}},{"cell_type":"code","source":"from __future__ import annotations\n\nimport os\nimport gc\nimport json\nimport shutil\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\nimport torch\nfrom IPython.display import display\n\nSEED = 2026\nnp.random.seed(SEED)\ntorch.manual_seed(SEED)\n\nTARGETS = [\n    \"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\", \"Medial OA\",\n    \"Lateral OA\", \"PF OA\", \"Effusion\", \"Synovitis\", \"Baker's\",\n    \"Contusion\", \"Fracture\",\n]\n\nWORK = Path(\"/kaggle/working\")\nOUT = WORK / \"rsna_stage1_bundle\"\nOUT.mkdir(parents=True, exist_ok=True)\n\nprint(\"torch:\", torch.__version__)\nprint(\"cuda available:\", torch.cuda.is_available())\nprint(\"output:\", OUT)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-11T19:37:10.439184Z","iopub.execute_input":"2026-08-11T19:37:10.439526Z","iopub.status.idle":"2026-08-11T19:37:16.448751Z","shell.execute_reply.started":"2026-08-11T19:37:10.439491Z","shell.execute_reply":"2026-08-11T19:37:16.447906Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 1. Locate competition data and the existing 20-member weight package","metadata":{}},{"cell_type":"code","source":"def find_train_root() -> Path:\n    explicit = os.environ.get(\"KNEE_INPUT_DIR\", \"\").strip()\n    candidates = []\n    if explicit:\n        candidates.append(Path(explicit))\n\n    candidates += [\n        Path(\"/kaggle/input/competitions/rsna-knee-abnormality-detection\"),\n        Path(\"/kaggle/input/rsna-knee-abnormality-detection\"),\n        Path(\"data\"),\n        Path(\".\"),\n    ]\n\n    for c in candidates:\n        if (c / \"train.csv\").is_file() and (c / \"train_series.csv\").is_file():\n            return c\n\n    base = Path(\"/kaggle/input\")\n    if base.is_dir():\n        for root, dirs, files in os.walk(base):\n            dirs[:] = [d for d in dirs if d not in (\"train_series\", \"test_series\")]\n            root = Path(root)\n            if (root / \"train.csv\").is_file() and (root / \"train_series.csv\").is_file():\n                return root\n\n    raise FileNotFoundError(\n        \"RSNA Knee training data not found. Attach the competition dataset or set KNEE_INPUT_DIR.\"\n    )\n\n\ndef valid_weight_package(root: Path) -> bool:\n    root = Path(root)\n    manifest_path = root / \"manifest.json\"\n    if not manifest_path.is_file():\n        return False\n\n    try:\n        manifest = json.loads(manifest_path.read_text())\n    except Exception:\n        return False\n\n    members = manifest.get(\"members\")\n    if not isinstance(members, list) or len(members) == 0:\n        return False\n\n    missing = [\n        str(m.get(\"file\"))\n        for m in members\n        if not (root / str(m.get(\"file\"))).is_file()\n    ]\n    if missing:\n        raise FileNotFoundError(\n            f\"Weight package exists but is incomplete. First missing checkpoint: {missing[0]}\"\n        )\n    return True\n\n\ndef find_weights() -> Path:\n    explicit = os.environ.get(\"KNEE_WEIGHTS_DIR\", \"\").strip()\n    candidates = []\n    if explicit:\n        candidates.append(Path(explicit))\n\n    candidates += [Path(\"/kaggle/input/rsna-knee-weights\")]\n\n    seen = set()\n    for p in candidates:\n        key = str(p)\n        if key in seen:\n            continue\n        seen.add(key)\n        if valid_weight_package(p):\n            return p\n\n    base = Path(\"/kaggle/input\")\n    if base.is_dir():\n        for root, dirs, files in os.walk(base):\n            dirs[:] = [d for d in dirs if d not in (\"train_series\", \"test_series\")]\n            if \"manifest.json\" not in files:\n                continue\n            root = Path(root)\n            if valid_weight_package(root):\n                return root\n\n    raise FileNotFoundError(\n        \"Compatible DINO manifest package not found. Attach rsna-knee-weights or set KNEE_WEIGHTS_DIR.\"\n    )\n\n\nROOT = find_train_root()\nWEIGHTS = find_weights()\n\nprint(\"competition root:\", ROOT)\nprint(\"weights root:\", WEIGHTS)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-11T19:37:16.45039Z","iopub.execute_input":"2026-08-11T19:37:16.450785Z","iopub.status.idle":"2026-08-11T19:37:16.504482Z","shell.execute_reply.started":"2026-08-11T19:37:16.450759Z","shell.execute_reply":"2026-08-11T19:37:16.503708Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2. Load training metadata and audit label availability","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(ROOT / \"train.csv\", dtype={\"StudyInstanceUID\": str})\ntrain_series = pd.read_csv(\n    ROOT / \"train_series.csv\",\n    dtype={\"StudyInstanceUID\": str, \"SeriesInstanceUID\": str},\n)\n\nmissing_target_columns = [t for t in TARGETS if t not in train_df.columns]\nif missing_target_columns:\n    raise KeyError(f\"train.csv is missing target columns: {missing_target_columns}\")\n\nif \"StudyInstanceUID\" not in train_df.columns:\n    raise KeyError(\"train.csv must contain StudyInstanceUID\")\n\nprint(\"studies:\", len(train_df))\nprint(\"series:\", len(train_series))\nprint(\"unique train UIDs:\", train_df[\"StudyInstanceUID\"].nunique())\nprint(\"duplicate train UIDs:\", int(train_df[\"StudyInstanceUID\"].duplicated().sum()))\n\nannotated = train_df[TARGETS].notna().sum()\npositive = train_df[TARGETS].sum(skipna=True)\nnegative = annotated - positive\n\nlabel_stats = pd.DataFrame({\n    \"target\": TARGETS,\n    \"annotated\": [int(annotated[t]) for t in TARGETS],\n    \"positive\": [float(positive[t]) for t in TARGETS],\n    \"negative\": [float(negative[t]) for t in TARGETS],\n})\nlabel_stats[\"positive_rate_on_annotated\"] = (\n    label_stats[\"positive\"] / label_stats[\"annotated\"].replace(0, np.nan)\n)\nlabel_stats.to_csv(OUT / \"label_stats.csv\", index=False)\n\ndisplay(label_stats.style.format({\"positive_rate_on_annotated\": \"{:.2%}\"}))\n\nn_complete = int(train_df[TARGETS].notna().all(axis=1).sum())\nn_any = int(train_df[TARGETS].notna().any(axis=1).sum())\nprint(f\"any image annotation: {n_any} / {len(train_df)}\")\nprint(f\"all 12 image annotations: {n_complete} / {len(train_df)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-11T19:37:20.330226Z","iopub.execute_input":"2026-08-11T19:37:20.330509Z","iopub.status.idle":"2026-08-11T19:37:20.641862Z","shell.execute_reply.started":"2026-08-11T19:37:20.330487Z","shell.execute_reply":"2026-08-11T19:37:20.641182Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Missing labels remain unknown\n\nFor every target we create both:\n\n- the numeric label column;\n- a `__mask` column indicating whether that target is actually annotated.\n\nNotebook 02 can therefore use a **masked loss** and avoid treating `NaN` as class `0`.","metadata":{}},{"cell_type":"code","source":"masked = train_df[[\"StudyInstanceUID\", *TARGETS]].copy()\n\nfor t in TARGETS:\n    masked[f\"{t}__mask\"] = masked[t].notna().astype(np.uint8)\n\n# Keep NaN in the label itself. The mask decides whether a loss term is valid.\nmasked.to_csv(OUT / \"train_labels_masked.csv\", index=False)\n\nexpert_complete = train_df.loc[\n    train_df[TARGETS].notna().all(axis=1),\n    [\"StudyInstanceUID\", *TARGETS]\n].copy()\nexpert_complete.to_csv(OUT / \"expert_complete_12.csv\", index=False)\n\npartially_labeled = train_df.loc[\n    train_df[TARGETS].notna().any(axis=1)\n    & ~train_df[TARGETS].notna().all(axis=1),\n    [\"StudyInstanceUID\", *TARGETS]\n].copy()\npartially_labeled.to_csv(OUT / \"partial_or_incomplete_labels.csv\", index=False)\n\nprint(\"masked label table:\", masked.shape)\nprint(\"complete 12-target expert subset:\", expert_complete.shape)\nprint(\"partial/incomplete subset:\", partially_labeled.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-11T19:37:22.595988Z","iopub.execute_input":"2026-08-11T19:37:22.596939Z","iopub.status.idle":"2026-08-11T19:37:22.661636Z","shell.execute_reply.started":"2026-08-11T19:37:22.596909Z","shell.execute_reply":"2026-08-11T19:37:22.660746Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3. Read the original DINO manifest and build a member registry","metadata":{}},{"cell_type":"code","source":"manifest_path = WEIGHTS / \"manifest.json\"\nmanifest = json.loads(manifest_path.read_text())\nmembers = manifest[\"members\"]\n\nprint(\"manifest top-level keys:\", sorted(manifest.keys()))\nprint(\"member count:\", len(members))\n\nrows = []\nfor i, m in enumerate(members):\n    cfg = m.get(\"config\") or {}\n    ck_path = WEIGHTS / str(m.get(\"file\"))\n    rows.append({\n        \"member_index\": i,\n        \"member_id\": m.get(\"id\"),\n        \"fold\": m.get(\"fold\"),\n        \"holdout\": m.get(\"holdout\"),\n        \"checkpoint_file\": m.get(\"file\"),\n        \"pixel_group\": m.get(\"pixel_group\"),\n        \"variant\": cfg.get(\"variant\"),\n        \"pool\": cfg.get(\"pool\", \"cls_mean\"),\n        \"prior\": cfg.get(\"prior\", False),\n        \"unfreeze_last\": cfg.get(\"unfreeze_last\"),\n        \"img\": cfg.get(\"img\"),\n        \"group\": cfg.get(\"group\"),\n        \"slices\": cfg.get(\"slices\"),\n        \"crop_mm\": cfg.get(\"crop_mm\"),\n        \"band\": json.dumps(cfg.get(\"band\")),\n        \"slots\": json.dumps(cfg.get(\"slots\")),\n        \"checkpoint_exists\": ck_path.is_file(),\n        \"checkpoint_bytes\": ck_path.stat().st_size if ck_path.is_file() else np.nan,\n    })\n\nregistry = pd.DataFrame(rows)\nregistry.to_csv(OUT / \"member_registry.csv\", index=False)\ndisplay(registry.sort_values([\"fold\", \"holdout\"], ascending=[True, False]))\n\nprint(\"\\nMembers per fold:\")\ndisplay(registry.groupby(\"fold\", dropna=False).size().rename(\"n_members\").to_frame())\n\nprint(\"\\nPixel groups:\")\ndisplay(registry.groupby(\"pixel_group\", dropna=False).size().rename(\"n_members\").to_frame())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-11T19:37:24.318563Z","iopub.execute_input":"2026-08-11T19:37:24.319356Z","iopub.status.idle":"2026-08-11T19:37:24.384362Z","shell.execute_reply.started":"2026-08-11T19:37:24.319309Z","shell.execute_reply":"2026-08-11T19:37:24.383462Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4. Inspect checkpoint structure safely","metadata":{}},{"cell_type":"code","source":"def summarize_checkpoint(path: Path) -> dict:\n    ck = torch.load(path, map_location=\"cpu\", weights_only=False)\n\n    info = {\n        \"file\": path.name,\n        \"container_type\": type(ck).__name__,\n    }\n\n    if isinstance(ck, dict):\n        info[\"keys\"] = json.dumps(sorted(map(str, ck.keys())))\n        state = ck.get(\"model\")\n\n        if isinstance(state, dict):\n            info[\"model_tensor_count\"] = len(state)\n            info[\"model_parameter_values\"] = int(\n                sum(v.numel() for v in state.values() if torch.is_tensor(v))\n            )\n            info[\"first_state_keys\"] = json.dumps(list(state.keys())[:12])\n        else:\n            info[\"model_tensor_count\"] = np.nan\n            info[\"model_parameter_values\"] = np.nan\n            info[\"first_state_keys\"] = None\n\n        info[\"has_fingerprint\"] = \"fingerprint\" in ck\n        info[\"has_optimizer\"] = any(\n            k in ck for k in (\"optimizer\", \"optimizer_state_dict\", \"optim\")\n        )\n        info[\"has_scheduler\"] = any(\n            k in ck for k in (\"scheduler\", \"scheduler_state_dict\")\n        )\n        info[\"has_epoch\"] = \"epoch\" in ck\n    else:\n        info[\"keys\"] = None\n        info[\"model_tensor_count\"] = np.nan\n        info[\"model_parameter_values\"] = np.nan\n        info[\"first_state_keys\"] = None\n        info[\"has_fingerprint\"] = False\n        info[\"has_optimizer\"] = False\n        info[\"has_scheduler\"] = False\n        info[\"has_epoch\"] = False\n\n    del ck\n    gc.collect()\n    return info\n\n\n# One representative checkpoint per (fold, pixel_group) is enough for structure audit.\naudit_indices = (\n    registry.sort_values([\"fold\", \"holdout\"], ascending=[True, False])\n    .drop_duplicates(subset=[\"fold\", \"pixel_group\"])\n    [\"member_index\"]\n    .astype(int)\n    .tolist()\n)\n\ncheckpoint_audit = []\nfor idx in audit_indices:\n    m = members[idx]\n    p = WEIGHTS / str(m[\"file\"])\n    print(\"auditing:\", p.name)\n    row = summarize_checkpoint(p)\n    row[\"member_id\"] = m.get(\"id\")\n    row[\"fold\"] = m.get(\"fold\")\n    checkpoint_audit.append(row)\n\ncheckpoint_audit = pd.DataFrame(checkpoint_audit)\ncheckpoint_audit.to_csv(OUT / \"checkpoint_structure_audit.csv\", index=False)\ndisplay(checkpoint_audit)\n\nif not checkpoint_audit.empty and not bool(checkpoint_audit[\"has_optimizer\"].all()):\n    print(\n        \"\\nNOTE: at least one audited checkpoint has no saved optimizer state. \"\n        \"That is fine: Notebook 02 can load model weights and create a NEW optimizer.\"\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-11T19:37:27.524504Z","iopub.execute_input":"2026-08-11T19:37:27.524836Z","iopub.status.idle":"2026-08-11T19:37:30.296862Z","shell.execute_reply.started":"2026-08-11T19:37:27.524808Z","shell.execute_reply":"2026-08-11T19:37:30.295962Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5. Search for a real original study-to-fold map\n\nA manifest member's `fold=2` tells us the checkpoint identity, **not which StudyInstanceUIDs belonged to fold 2**.\n\nThe code below only accepts a CSV as a fold map if it contains:\n\n- `StudyInstanceUID`\n- a column named `fold`, `Fold`, `fold_id`, or `cv_fold`\n\nand each study maps to exactly one fold.","metadata":{}},{"cell_type":"code","source":"FOLD_COLUMN_CANDIDATES = [\"fold\", \"Fold\", \"fold_id\", \"cv_fold\"]\n\ndef candidate_csvs(search_roots):\n    seen = set()\n    out = []\n\n    for root in search_roots:\n        root = Path(root)\n        if not root.exists():\n            continue\n\n        # Avoid recursively traversing huge DICOM directories.\n        for pattern in (\"*.csv\", \"*/*.csv\", \"*/*/*.csv\"):\n            for p in root.glob(pattern):\n                if any(part in {\"train_series\", \"test_series\"} for part in p.parts):\n                    continue\n                key = str(p.resolve())\n                if key in seen:\n                    continue\n                seen.add(key)\n                out.append(p)\n    return out\n\n\ndef inspect_fold_csv(path: Path):\n    try:\n        x = pd.read_csv(path, nrows=20000, dtype={\"StudyInstanceUID\": str})\n    except Exception:\n        return None\n\n    if \"StudyInstanceUID\" not in x.columns:\n        return None\n\n    fold_col = next((c for c in FOLD_COLUMN_CANDIDATES if c in x.columns), None)\n    if fold_col is None:\n        return None\n\n    y = x[[\"StudyInstanceUID\", fold_col]].dropna().copy()\n    if y.empty:\n        return None\n\n    unique_per_uid = y.groupby(\"StudyInstanceUID\")[fold_col].nunique()\n    valid_unique = bool((unique_per_uid <= 1).all())\n\n    return {\n        \"path\": str(path),\n        \"fold_column\": fold_col,\n        \"rows_read\": len(x),\n        \"mapped_rows\": len(y),\n        \"unique_uids\": y[\"StudyInstanceUID\"].nunique(),\n        \"one_fold_per_uid\": valid_unique,\n        \"fold_values\": json.dumps(sorted(map(str, y[fold_col].unique().tolist()))),\n    }\n\n\nfold_candidates = []\nfor p in candidate_csvs([WEIGHTS, ROOT]):\n    info = inspect_fold_csv(p)\n    if info is not None:\n        fold_candidates.append(info)\n\nfold_candidates_df = pd.DataFrame(fold_candidates)\nfold_candidates_df.to_csv(OUT / \"fold_artifact_candidates.csv\", index=False)\n\nif len(fold_candidates_df):\n    display(fold_candidates_df)\nelse:\n    print(\"No CSV study-to-fold artifact was found.\")\n\naccepted_fold_map = None\naccepted_source = None\n\n# Conservative automatic acceptance:\n# require one-fold-per-UID and at least 90% coverage of the training UID set.\ntrain_uids = set(train_df[\"StudyInstanceUID\"].astype(str))\n\nfor info in fold_candidates:\n    if not info[\"one_fold_per_uid\"]:\n        continue\n\n    p = Path(info[\"path\"])\n    x = pd.read_csv(p, dtype={\"StudyInstanceUID\": str})\n    fc = info[\"fold_column\"]\n    x = x[[\"StudyInstanceUID\", fc]].dropna().drop_duplicates().copy()\n\n    coverage = (\n        len(train_uids.intersection(set(x[\"StudyInstanceUID\"])))\n        / max(len(train_uids), 1)\n    )\n\n    if coverage >= 0.90:\n        x = x.rename(columns={fc: \"fold\"})\n        accepted_fold_map = x\n        accepted_source = str(p)\n        break\n\nHAS_SAFE_ORIGINAL_FOLD_MAP = accepted_fold_map is not None\n\nif HAS_SAFE_ORIGINAL_FOLD_MAP:\n    accepted_fold_map.to_csv(OUT / \"original_fold_map.csv\", index=False)\n    print(\"Accepted fold map:\", accepted_source)\n    print(\"mapped studies:\", accepted_fold_map[\"StudyInstanceUID\"].nunique())\n    print(\"fold counts:\")\n    display(accepted_fold_map.groupby(\"fold\").size().rename(\"n_studies\").to_frame())\nelse:\n    print(\n        \"SAFE RESULT: no high-coverage original fold map was automatically verified.\\n\"\n        \"Notebook 02 must not invent a historical fold assignment.\"\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-11T19:37:30.298241Z","iopub.execute_input":"2026-08-11T19:37:30.298614Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6. Build a resume manifest for Notebook 02","metadata":{}},{"cell_type":"code","source":"resume_members = []\nfor m in members:\n    cfg = m.get(\"config\") or {}\n    resume_members.append({\n        \"id\": m.get(\"id\"),\n        \"fold\": m.get(\"fold\"),\n        \"holdout\": m.get(\"holdout\"),\n        \"base_checkpoint_file\": m.get(\"file\"),\n        \"pixel_group\": m.get(\"pixel_group\"),\n        \"config\": cfg,\n        \"finetuned_checkpoint_file\": None,\n        \"finetuned_best_metric\": None,\n        \"finetuned_epoch\": None,\n    })\n\nresume_manifest = {\n    \"stage\": \"M0_BASELINE_AUDIT\",\n    \"source_weight_package\": str(WEIGHTS),\n    \"source_manifest\": str(manifest_path),\n    \"seed\": SEED,\n    \"targets\": TARGETS,\n    \"n_train_studies\": int(len(train_df)),\n    \"n_fully_annotated_12\": int(len(expert_complete)),\n    \"has_verified_original_fold_map\": bool(HAS_SAFE_ORIGINAL_FOLD_MAP),\n    \"verified_fold_map_source\": accepted_source,\n    \"members\": resume_members,\n}\n\n(OUT / \"resume_manifest.json\").write_text(json.dumps(resume_manifest, indent=2))\n(OUT / \"manifest_snapshot.json\").write_text(json.dumps(manifest, indent=2))\n\nprint(json.dumps({\n    \"stage\": resume_manifest[\"stage\"],\n    \"members\": len(resume_members),\n    \"n_train_studies\": resume_manifest[\"n_train_studies\"],\n    \"n_fully_annotated_12\": resume_manifest[\"n_fully_annotated_12\"],\n    \"has_verified_original_fold_map\": resume_manifest[\"has_verified_original_fold_map\"],\n}, indent=2))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 7. Write the Stage 1 safety contract","metadata":{}},{"cell_type":"code","source":"status = {\n    \"stage1_ok\": True,\n    \"competition_root\": str(ROOT),\n    \"weights_root\": str(WEIGHTS),\n    \"member_count\": int(len(members)),\n    \"fold_ids_in_manifest\": sorted(\n        [\n            x.item() if hasattr(x, \"item\") else x\n            for x in registry[\"fold\"].dropna().unique().tolist()\n        ],\n        key=lambda z: str(z),\n    ),\n    \"targets\": TARGETS,\n    \"masked_label_file\": \"train_labels_masked.csv\",\n    \"expert_complete_file\": \"expert_complete_12.csv\",\n    \"member_registry_file\": \"member_registry.csv\",\n    \"resume_manifest_file\": \"resume_manifest.json\",\n    \"original_fold_map_file\": (\n        \"original_fold_map.csv\" if HAS_SAFE_ORIGINAL_FOLD_MAP else None\n    ),\n    \"has_verified_original_fold_map\": bool(HAS_SAFE_ORIGINAL_FOLD_MAP),\n    \"next_stage_rule\": (\n        \"Use verified historical fold mapping for fold-safe continuation.\"\n        if HAS_SAFE_ORIGINAL_FOLD_MAP\n        else\n        \"Do not claim fold-safe OOF from the base checkpoints until an original \"\n        \"study-to-fold map or an explicitly independent validation design is supplied.\"\n    ),\n}\n\n(OUT / \"stage1_status.json\").write_text(json.dumps(status, indent=2))\n\nreadme_lines = [\n    \"RSNA Knee continual fine-tuning - Stage 1 bundle\",\n    \"\",\n    f\"Competition root: {ROOT}\",\n    f\"Weights root: {WEIGHTS}\",\n    f\"Members: {len(members)}\",\n    f\"Train studies: {len(train_df)}\",\n    f\"Complete 12-target expert rows: {len(expert_complete)}\",\n    f\"Verified original fold map: {HAS_SAFE_ORIGINAL_FOLD_MAP}\",\n    \"\",\n    \"Files:\",\n    \"- member_registry.csv\",\n    \"- checkpoint_structure_audit.csv\",\n    \"- label_stats.csv\",\n    \"- train_labels_masked.csv\",\n    \"- expert_complete_12.csv\",\n    \"- partial_or_incomplete_labels.csv\",\n    \"- fold_artifact_candidates.csv\",\n    \"- original_fold_map.csv (only when verified)\",\n    \"- manifest_snapshot.json\",\n    \"- resume_manifest.json\",\n    \"- stage1_status.json\",\n    \"\",\n    \"Safety rule:\",\n    status[\"next_stage_rule\"],\n]\nreadme = \"\\n\".join(readme_lines)\n(OUT / \"README.txt\").write_text(readme)\n\nprint(readme)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 8. Create the transfer bundle for Notebook 02","metadata":{}},{"cell_type":"code","source":"bundle_path = WORK / \"rsna_stage1_bundle.zip\"\nif bundle_path.exists():\n    bundle_path.unlink()\n\nshutil.make_archive(\n    str(bundle_path.with_suffix(\"\")),\n    \"zip\",\n    root_dir=OUT,\n)\n\nprint(\"Created:\", bundle_path)\nprint(\"Size MB:\", round(bundle_path.stat().st_size / (1024**2), 3))\n\nprint(\"\\nBundle contents:\")\nfor p in sorted(OUT.iterdir()):\n    print(\" -\", p.name)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Notebook 01 output\n\nThe file to carry into the next stage is:\n\n`/kaggle/working/rsna_stage1_bundle.zip`\n\nThe **original checkpoint files are not duplicated** into this ZIP because they are large and are already available through the existing `rsna-knee-weights` Kaggle input. Notebook 02 should mount both:\n\n1. the original weight package; and\n2. the Stage 1 bundle.\n\nNotebook 02 will then implement **conservative fine-tuning**, after checking `stage1_status.json` and the validation/fold safety condition.","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}