{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"acaf29a1-6270-4282-8c15-d4427a674eb8","cell_type":"markdown","source":"# M1 — build the cache and gate it, one CPU run\n\nDecodes every study of the split into the compact cache the training notebook reads, then\nruns the acceptance gate on its own output and prints **PASS/FAIL**.\n\nNothing here asks you to make a decision mid-run. The builder measures itself before it\ncommits, stops itself before the session is killed, and resumes from the manifest if it did\nnot finish — so \"will 4,407 studies fit in 12 h\" is a thing you observe, not a thing you have\nto guess right in advance.\n\n| need | how |\n|---|---|\n| **CPU** accelerator | 12 h limit, and it does not touch the 30 h/week GPU quota |\n| competition data | `rsna-knee-abnormality-detection` → `/kaggle/input/competitions/rsna-knee-abnormality-detection` |\n| `knee-tools` dataset | this repo — `geometry.py`, `cache_gate.py`, `knee_cache.py` → `/kaggle/input/datasets/tamerlanomralinov/knee-tools` |\n| *(optional)* an earlier run's output | attach it; the build continues where it stopped |\n\nWhy the build and the gate are the same run: our first submission scored **0.643** because\nthe cache covered one shard of four and nobody re-checked, and because `resolve_laterality`\nread an image *corner* as the knee's position and mirrored an arbitrary subset of coronal and\naxial series. Both failures are silent. A separate verification step is a step you can skip.","metadata":{}},{"id":"803a678b-d9c9-4eb5-a29c-cbc4eaad7415","cell_type":"code","source":"# ----------------------------------------------------------------------------------\n# Config. Everything else is derived — study counts come from the CSV, cache directories from\n# knee_cache.PRIOR_DIRS, so there is no second copy of anything to keep in sync.\n# ----------------------------------------------------------------------------------\nCFG = dict(\n    split=\"train\",                 # \"train\" | \"test\"\n    out_dir=\"/kaggle/working/cache\",\n    tools_path=\"/kaggle/input/datasets/tamerlanomralinov/knee-tools/tools\",\n    notebooks_path=\"/kaggle/input/datasets/tamerlanomralinov/knee-tools/notebooks\",\n\n    # Wall clock for the WHOLE notebook, measured from the cell below. Kaggle kills a CPU\n    # session at 12 h; the gate needs a few minutes and \"Save Version\" re-runs everything, so\n    # the build gets 10.8 h and the rest is margin. Lower it to test the resume path on purpose.\n    budget_h=10.8,\n\n    smoke_studies=3,               # decoded end to end before committing the session\n    want_slots=None,               # e.g. (\"sag_fs\", \"cor_fs\") to halve the time and the disk\n    retry_failed=False,            # True only after changing the decode path\n    skip_build=False,              # True = gate an already-built cache and nothing else\n)\n\nimport os                                                                  # noqa: E402\nimport sys                                                                 # noqa: E402\nimport glob                                                                # noqa: E402\nimport time                                                                # noqa: E402\nimport shutil                                                              # noqa: E402\nimport pickle                                                              # noqa: E402\nimport traceback                                                           # noqa: E402\n\nNB_T0 = time.time()\n\n\ndef find_repo(hint_tools, hint_notebooks):\n    \"\"\"Locate the knee-tools payload wherever Kaggle mounted it, or say exactly how to fix it.\n\n    The notebooks import geometry / cache_gate / knee_cache instead of pasting them, so that\n    training and inference provably preprocess identically -- which means the repo has to\n    exist on Kaggle as a dataset. Hardcoding one mount path turns \"the dataset is attached\n    under a slightly different slug\" into an ImportError with no instructions attached.\n    \"\"\"\n    def has(d, *names):\n        return d and all(os.path.isfile(os.path.join(d, n)) for n in names)\n\n    def where(hint, leaf, *needs):\n        # depth 3 is the layout this account gets: /kaggle/input/datasets/<owner>/<slug>/<leaf>\n        cands = [hint, *sorted(glob.glob(f\"/kaggle/input/*/{leaf}\")),\n                 *sorted(glob.glob(f\"/kaggle/input/*/*/{leaf}\")),\n                 *sorted(glob.glob(f\"/kaggle/input/*/*/*/{leaf}\")), leaf, f\"../{leaf}\"]\n        return next((d for d in cands if has(d, *needs)), None)\n\n    tools = where(hint_tools, \"tools\", \"geometry.py\", \"cache_gate.py\")\n    nbs = where(hint_notebooks, \"notebooks\", \"knee_cache.py\")\n    if not (tools and nbs):\n        # list two extra depths: with the datasets/<owner>/<slug> layout, /kaggle/input/*\n        # prints ['competitions', 'datasets'] and tells you nothing about what is attached.\n        mounted = sorted(glob.glob(\"/kaggle/input/*\") + glob.glob(\"/kaggle/input/*/*/*\"))\n        raise SystemExit(\n            \"knee-tools not found (tools=%r notebooks=%r).\\n\"\n            \"/kaggle/input holds: %s\\n\\n\"\n            \"Add Input -> Datasets -> knee-tools. If it is missing or out of date, push it \"\n            \"from the repo:\\n    python tools/kaggle_push.py -m \\\"what changed\\\"\\n\"\n            \"and attach the new version -- an attached-but-stale dataset imports fine and \"\n            \"runs last week's code.\" % (tools, nbs, mounted))\n    return tools, nbs\n\n\nTOOLS, NOTEBOOKS = find_repo(CFG[\"tools_path\"], CFG[\"notebooks_path\"])\nsys.path.insert(0, TOOLS)\nsys.path.insert(0, NOTEBOOKS)\n\n# Which version of the code this session is actually running. kaggle_push.py stamps VERSION.txt\n# with a push time and a sha1 per file; if this line is older than your last edit, the dataset\n# is stale and nothing below is testing what you think it is.\n_v = os.path.join(os.path.dirname(TOOLS.rstrip(\"/\")), \"VERSION.txt\")\nprint(open(_v).readline().strip() if os.path.isfile(_v) else\n      f\"tools {TOOLS} (no VERSION.txt -- pushed before tools/kaggle_push.py existed)\")\n\nimport knee_cache as kc                                                    # noqa: E402\nimport cache_gate                                                          # noqa: E402\n\n# knee_cache.py is imported, not pasted: the submission notebook runs this same file on the\n# hidden test set, so train and inference cannot drift apart in how they decode, window, order\n# or mirror a slice. A mismatch there is invisible in CV and fatal on LB.\nkc.SPLIT = CFG[\"split\"]\nkc.OUT_DIR = CFG[\"out_dir\"]\nkc.SHARD, kc.N_SHARDS = 0, 1\nkc.WANT_SLOTS = CFG[\"want_slots\"]\nkc.RETRY_FAILED = CFG[\"retry_failed\"]\n\nCACHE_DIRS = [CFG[\"out_dir\"], *kc.PRIOR_DIRS]      # the gate reads the union: a build resumed\n                                                   # across sessions is legitimately split\nDEADLINE = NB_T0 + CFG[\"budget_h\"] * 3600\nprint(f\"cache params: {kc.cache_params()}\")\nprint(f\"workers {kc.N_WORKERS} | budget {CFG['budget_h']} h | cache dirs {CACHE_DIRS}\")","metadata":{},"outputs":[],"execution_count":null},{"id":"9dd47035-2025-45d6-a6f2-983b5808493b","cell_type":"markdown","source":"## Preflight\n\nTen seconds of checks against the ten hours they protect: is the competition actually\nmounted, is there room on disk, and — the one that matters — does a real study decode end to\nend with this code, on this pydicom, from these DICOMs. A missing `geometry.py`, an\nunreadable transfer syntax or a wrong dataset should cost you a cell, not a session.","metadata":{}},{"id":"b3c60022-9fe0-4b86-92ac-61b7ffb94589","cell_type":"code","source":"root = kc.resolve_root(CFG[\"split\"])\nby_study = kc.load_series_index(CFG[\"split\"])\nEXPECTED = len(by_study)                       # derived, never hardcoded: 4,407 is the train\n                                               # count and gating test against it is nonsense\ndone, attached, failed_before = kc.already_built()\ntodo = [u for u in kc.work_order(by_study) if u not in done]\n\nfree_gb = shutil.disk_usage(os.path.dirname(CFG[\"out_dir\"]) or \"/\").free / 1e9\nprint(f\"root      : {root}\")\nprint(f\"studies   : {EXPECTED} unique in {CFG['split']}_series.csv\")\nprint(f\"cached    : {len(done)} ({attached} from attached datasets, {failed_before} failed \"\n      f\"earlier and not being retried)\")\nprint(f\"to build  : {len(todo)}\")\nprint(f\"disk free : {free_gb:.1f} GB\")\n\n# One study end to end, in this process, with the traceback visible. imap_unordered swallows\n# exceptions into a manifest column by design -- fine for one bad study out of 4,407, useless\n# when the cause is systematic and you find out after 40 minutes of writing error rows.\nos.makedirs(CFG[\"out_dir\"], exist_ok=True)\nsmoke, s_per_study, mb_per_study = [], None, None\nn_smoke = 0 if CFG[\"skip_build\"] else CFG[\"smoke_studies\"]\nfor uid in (todo or list(by_study))[:n_smoke]:\n    t = time.time()\n    try:\n        data, meta = kc.build_study(uid, by_study[uid])\n    except Exception:\n        traceback.print_exc()\n        raise SystemExit(\"preflight failed to decode a study -- fix this before the build\")\n    nbytes = len(pickle.dumps(data, protocol=4))\n    smoke.append((time.time() - t, nbytes, meta))\n    print(f\"  {uid[-12:]}  {time.time() - t:5.1f} s  {nbytes / 1e6:5.2f} MB  \"\n          f\"slots={meta['slots'] or '-'}  lat={meta['laterality'] or '?'}({meta['lat_route']})\"\n          f\"  decode_failed={meta['decode_failed']}\")\n\nif smoke:\n    # /N_WORKERS because the real run is a process pool; Kaggle's 4 vCPU scale close to linearly\n    # on DICOM decode, which is CPU-bound and independent per study.\n    s_per_study = sum(s for s, _, _ in smoke) / len(smoke) / kc.N_WORKERS\n    mb_per_study = sum(b for _, b, _ in smoke) / len(smoke) / 1e6\n    hours = s_per_study * len(todo) / 3600\n    gb = mb_per_study * EXPECTED / 1000\n    print(f\"\\nprojection: {s_per_study:.2f} s/study on {kc.N_WORKERS} workers -> \"\n          f\"{hours:.1f} h for the {len(todo)} left, {gb:.1f} GB for all {EXPECTED}\")\n    if gb > free_gb:\n        print(f\"!! projected {gb:.1f} GB does not fit in {free_gb:.1f} GB free. Cut \"\n              f\"CFG['want_slots'] to the slots you train on, or lower kc.SIZE.\")\n    if hours > CFG[\"budget_h\"]:\n        print(f\"-> more than one session ({hours:.1f} h > {CFG['budget_h']} h). Nothing to \"\n              f\"change: it will build what it can, stop cleanly, and tell you to attach this\\n\"\n              f\"   output to a second run. Set want_slots if you would rather have it in one.\")\n    else:\n        print(\"-> fits this session.\")","metadata":{},"outputs":[],"execution_count":null},{"id":"3bdc8d12-75af-4726-abd4-bad79dd6a35e","cell_type":"markdown","source":"## Build\n\nDeterministic shuffled order, so a run that hits the deadline leaves an unbiased sample of\nthe corpus rather than an alphabetical prefix. Progress prints s/study and an ETA every 50.","metadata":{}},{"id":"606583f7-d757-4803-bad3-f09305f9b60f","cell_type":"code","source":"if CFG[\"skip_build\"]:\n    build = None\n    print(\"skip_build -- gating the existing cache only\")\nelse:\n    build = kc.main(deadline=DEADLINE)","metadata":{},"outputs":[],"execution_count":null},{"id":"4f3ac514-8ff5-467d-b1f1-c573a3b6b55e","cell_type":"markdown","source":"## The gate\n\n`cache_gate.run` is a module, not cells, so the check that runs here and the check you run\nlater against an attached cache are literally the same code. Two copies would drift, and the\none that drifted would be the one that passed.\n\n| it decides | it cannot decide |\n|---|---|\n| coverage against the split's own study count | **the L/R grids — you have to look** |\n| laterality routes, plus the positional route audited against the DICOM tag | |\n| site keys (`GroupKFold` needs them; unrecoverable after the build) | |\n| decode failures — a failed slice becomes zeros, indistinguishable from a black knee | |\n| slot coverage, size against the 20 GB cap | |","metadata":{}},{"id":"14a19c44-6938-4003-b44d-c799e8d8b618","cell_type":"code","source":"result = cache_gate.run(CACHE_DIRS, split=CFG[\"split\"], expected=EXPECTED)","metadata":{},"outputs":[],"execution_count":null},{"id":"66b5bf9a-68f5-4664-847d-ca9ab01ce9a1","cell_type":"markdown","source":"## Read the grids before believing the PASS\n\nOn coronal the **fibula sits on the lateral side**. It must appear on the same side of the\nframe in every panel of both rows. If the top row looks like a mirror of the bottom row, the\nlaterality pass is wrong however clean the manifest was — that is exactly the failure that\ncost four columns last time, and no number above can see it.\n\nThe slice-progression strip should change smoothly. A spike in the difference plot means a\nmisordered slice, which also breaks the 2.5D triples that hand the encoder three adjacent\nslices as three channels.","metadata":{}},{"id":"15d5845d-61bf-45c4-a5f6-6255a95e168f","cell_type":"code","source":"# ----------------------------------------------------------------------------------\n# What this run says about the submission notebook. The hidden test set has to be cached\n# inside the same limit as inference, and the cache build is the expensive half -- this number\n# sizes the ensemble in CLIMB.md Rung 5, so record it now rather than rediscovering it later.\n# ----------------------------------------------------------------------------------\ntry:                                    # the hidden test set is not mounted during training,\n    n_test = len(kc.load_series_index(\"test\"))     # but its series CSV usually is\nexcept OSError:\n    n_test = 0\n# The PUBLIC test_series.csv is a 3-study placeholder -- it reads fine, so the except above\n# never fires, and the first real build cheerfully reported \"3 test studies take 0.00 h,\n# leaving 9.00 h for inference\". A runtime budget that says you have all of it is worse than\n# no runtime budget, so anything this small is treated as the placeholder it is.\nn_test_src = \"test_series.csv\"\nif n_test < 50:\n    n_test, n_test_src = 1300, f\"PLAN.md estimate; test_series.csv held only {n_test}\"\nmeasured = build[\"s_per_study\"] if build and build[\"built\"] else s_per_study\nprint(f\"elapsed this notebook: {(time.time() - NB_T0) / 3600:.2f} h of the {CFG['budget_h']} h \"\n      f\"budget\")\nif measured:\n    test_h = measured * n_test / 3600\n    src = \"measured over the build\" if build and build[\"built\"] else \"extrapolated from preflight\"\n    print(f\"{measured:.2f} s/study ({src}) -> {n_test} test studies ({n_test_src}) take \"\n          f\"{test_h:.2f} h, leaving {9 - test_h:.2f} h of a 9 h submission run for inference\")\n    if test_h > 4:\n        print(\"!! More than 4 h just to decode the test set. Before Rung 5: build only the \"\n              \"slots the\\n   ensemble uses (CFG['want_slots']) and keep the arrays in RAM \"\n              \"instead of the JPEG round trip.\")\n\nif build and build[\"stopped_at_deadline\"]:\n    print(f\"\\nNEXT: save this output as a dataset named knee-cache*, attach it to a fresh run \"\n          f\"of this\\nnotebook, and press go. {build['left']} studies left; nothing is \"\n          f\"recomputed and nothing is copied.\")\nelif result[\"passed\"]:\n    print(\"\\nNEXT: save this output as a Kaggle dataset named knee-cache (training and \"\n          \"submission\\nglob knee-cache*), and write down the s/study above.\")","metadata":{},"outputs":[],"execution_count":null},{"id":"8c24c4a7-313e-42a1-aee9-858a28eb788f","cell_type":"code","source":"import numpy as np, cache_gate as cg\n\nman, ok, bad, parts = cg.load_manifest(CACHE_DIRS, \"train\")\n\ndef mean_mid(side, slot=\"cor_fs\", n=60):\n    sub = ok[(ok[\"laterality\"] == side)\n             & ok[\"slots\"].fillna(\"\").str.contains(slot, regex=False)].head(n)\n    acc = [cg.load_slot(CACHE_DIRS, u, slot) for u in sub[\"StudyInstanceUID\"]]\n    return np.mean([f[len(f) // 2].astype(np.float32) for f in acc if f], 0)\n\nL, R = mean_mid(\"L\"), mean_mid(\"R\")\nsame, mirror = np.abs(L - R).mean(), np.abs(L - R[:, ::-1]).mean()\nself_asym = np.abs(L - L[:, ::-1]).mean()\n\nprint(f\"|meanL - meanR|         = {same:.2f}   <- должно быть МЕНЬШЕ\")\nprint(f\"|meanL - mirror(meanR)| = {mirror:.2f}\")\nprint(f\"|meanL - mirror(meanL)| = {self_asym:.2f}   <- запас мощности теста\")\nprint(\"OK: канонизация держится\" if same < mirror else \"!! ЗЕРКАЛО: ряды не согласованы\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"e9e3883f-da59-4f98-84ea-dcba10b6f3fb","cell_type":"markdown","source":"## If it failed\n\n**Coverage short after a clean stop** — save the output, attach it, run again. `PRIOR_DIRS`\nglobs `knee-cache*`, so the builder skips everything the attached manifests report as done.\n\n**Centre-vs-tag agreement** — fix the geometry, not the cache. Reproduce it as a failing case\nin `tools/test_cache_laterality.py`; that suite is stdlib-only and runs in a second on any\nmachine, so there is no reason to debug it against 569 GB of DICOM.\n\n**Grids look mirrored** — check `geometry.compose`, which XORs the scanner's stored row\ndirection against the laterality flip. Applying those two independently double-flips and\nsilently undoes the correction, which is the whole reason that function exists.\n\n**Many decode errors** — read the top error strings the gate printed. If they are systematic\n(a transfer syntax, a missing tag), fix the decode path and rerun with\n`CFG[\"retry_failed\"] = True`, which is the only thing that re-attempts studies already\nrecorded as failed.","metadata":{}}]}