{"cells":[{"cell_type":"markdown","metadata":{},"source":"# Lateral OA / Effusion specialist -- scaling nekkon's report-vs-image experiment\n\n`nekkon/the-scan-knows-two-things-the-report-does-not` measured, on 929 of the\n4,407 studies (disk-limited), that a weak-label-trained image model beats the\nreport-derived label outright on two findings: **Lateral compartment OA**\n(image AUC 0.77 vs report-derived label 0.56) and **Effusion** (0.72 vs\n0.67). Everywhere else the report wins by a lot, because the community's\nexisting ensembles (DINOsaur/H&S, CoAtNet, RadImageNet arms) are all trained\non the same report-derived weak labels for every finding uniformly -- none of\nthem specifically re-weights toward images where images are known to add\ninformation.\n\nThis notebook reruns nekkon's exact recipe (EfficientNet-B3, 6-slot MIL,\nper-finding attention pooling, soft BCE against `weak_labels_v2`) with the\nonly change being **data volume**: instead of a local, disk-limited copy of\n929 studies, it reads directly from Kaggle's live competition mount and\ncaches as much of the full 4,407-study corpus as fits in the session's time\nbudget. Section 5 of the source notebook found radiologist-AUC still rising\nat 929 studies (`0.593 -> 0.609 -> 0.646`, ~5 points per doubling, no sign of\nflattening) -- this is the natural next data point on that curve.\n\nThis is an experiment, not a submission: no `submission.csv` is written. The\ndeliverable is an honest, gold58 (radiologist-annotated, never-in-training)\nAUC reading for all twelve findings, to decide whether Lateral OA / Effusion\nspecifically are worth gold58-gating into the production blend as a new,\nindependent signal source -- per this project's own finding that only\n*genuinely new signal*, not fixed-weight reblends of existing components, has\never reproduced on the public leaderboard.\n\nCredit: recipe, weak-label pipeline and the report-vs-image measurement are\n[nekkon](https://www.kaggle.com/nekkon)'s\n(`rsna-knee-image-oof-gold`, `weak-labels-for-all-12-knee-mri-findings`,\nCC0-1.0)."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"\"\"\"Cell 0: run-size configuration. FULL RUN (smoke test passed).\"\"\"\nimport os\nos.environ['CACHE_TIME_BUDGET_S'] = '3600'\nos.environ['EPOCHS'] = '10'\nos.environ['BS'] = '16'\nos.environ['N_FOLDS'] = '5'\n"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"\"\"\"Cell 1: environment setup, path resolution, weak-label loading.\n\nAdapted from nekkon/rsna-knee-image-oof-gold (rsna_cache.py + rsna_train.py),\nscaled from the author's disk-limited 929 studies to as much of the full\n4,407-study corpus as fits in this session's time budget, using Kaggle's\nlive competition mount instead of a local DICOM copy.\n\"\"\"\nimport os, sys, glob, json, time, math\nfrom pathlib import Path\nimport numpy as np\nimport pandas as pd\nimport pydicom\nimport torch\n\nWORK = Path('/kaggle/working')\nCACHE = WORK / 'cache'\nCACHE.mkdir(exist_ok=True)\n\n_COMP_ROOTS = [\n    Path('/kaggle/input/rsna-knee-abnormality-detection'),\n    Path('/kaggle/input/competitions/rsna-knee-abnormality-detection'),\n]\nCOMP = next((p for p in _COMP_ROOTS if (p / 'train.csv').is_file()), _COMP_ROOTS[0])\nDATA = COMP / 'train_series'\n\n_wl_hits = sorted(Path('/kaggle/input').rglob('weak_labels_v2.csv'))\nif not _wl_hits:\n    print('weak_labels_v2.csv not found by rglob; /kaggle/input tree:')\n    for p in sorted(Path('/kaggle/input').rglob('*')):\n        print(' ', p)\n    raise FileNotFoundError('weak_labels_v2.csv not found anywhere under /kaggle/input')\nWL_PATH = _wl_hits[0]\n\nprint('competition root:', COMP)\nprint('dicom root      :', DATA)\nprint('weak label file :', WL_PATH, '| all hits:', _wl_hits)\nprint('cuda available  :', torch.cuda.is_available())\nfor i in range(torch.cuda.device_count()):\n    print(' gpu', i, torch.cuda.get_device_name(i))\nprint('cpu count       :', os.cpu_count())\n\ntrain = pd.read_csv(COMP / 'train.csv')\nseries = pd.read_csv(COMP / 'train_series.csv')\nW = pd.read_csv(WL_PATH).set_index('StudyInstanceUID')\n\nLAB = ['ACL', 'MCL', 'Medial Meniscus', 'Lateral Meniscus', 'Medial OA', 'Lateral OA',\n       'PF OA', 'Effusion', 'Synovitis', \"Baker's\", 'Contusion', 'Fracture']\nassert LAB == [c for c in train.columns if c not in ('StudyInstanceUID', 'Report')]\n\nY = W[LAB].copy()\nfor l in ('Medial OA', 'Lateral OA'):\n    if f'{l}_contrast' in W.columns:\n        Y[l] = (W[l] + W[f'{l}_contrast'].rank(pct=True)) / 2\n\ngold_uid = set(train.StudyInstanceUID[train[LAB].notna().all(axis=1)])\nG = train.set_index('StudyInstanceUID')[LAB]\nfold = (pd.util.hash_pandas_object(train.Report, index=False) % 5).values\nfold_of = dict(zip(train.StudyInstanceUID, fold))\n\nprint(f'studies total={len(train)} gold={len(gold_uid)}')\nprint('label prevalence in gold58:')\nprint(G.loc[list(gold_uid)].mean().round(3))\n"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"\"\"\"Cell 2: DICOM -> uint8 cache functions (adapted from nekkon/rsna_cache.py).\n\nPer study -> per (plane, fluid) slot -> S slices over the middle band, cropped\nto a fixed physical extent (mm) and resampled to IMG px, percentile-normalised.\nUnchanged from the source recipe except for path resolution (Kaggle mount,\nnot a local DICOM copy) and an added wall-clock budget so a partial cache is\nusable if the full corpus does not fit in one session.\n\"\"\"\nfrom scipy.ndimage import zoom\nfrom concurrent.futures import ProcessPoolExecutor\n\nIMG = 256\nS = 12\nCROP_MM = 140.0\nBAND = (0.15, 0.85)\nSLOTS = [(p, f) for p in ('Sagittal', 'Coronal', 'Axial') for f in (1, 0)]\nN_WORKERS = min(os.cpu_count() or 4, 4)\n\ndef order_files(files):\n    recs = []\n    for f in files:\n        try:\n            d = pydicom.dcmread(f, stop_before_pixels=True)\n            iop = np.array(d.ImageOrientationPatient, float)\n            ipp = np.array(d.ImagePositionPatient, float)\n            n = np.cross(iop[:3], iop[3:])\n            recs.append((float(ipp @ n), f, ipp, d))\n        except Exception:\n            continue\n    recs.sort(key=lambda r: r[0])\n    return recs\n\ndef read_slice(f):\n    d = pydicom.dcmread(f)\n    a = d.pixel_array.astype(np.float32)\n    if getattr(d, 'PhotometricInterpretation', '') == 'MONOCHROME1':\n        a = a.max() - a\n    ps = [float(x) for x in getattr(d, 'PixelSpacing', [1, 1])]\n    return a, ps\n\ndef crop_resize(a, ps):\n    h, w = a.shape\n    nh, nw = int(round(CROP_MM / ps[0])), int(round(CROP_MM / ps[1]))\n    cy, cx = h // 2, w // 2\n    y0, x0 = max(0, cy - nh // 2), max(0, cx - nw // 2)\n    c = a[y0:y0 + nh, x0:x0 + nw]\n    if c.shape[0] < nh or c.shape[1] < nw:\n        pad = np.zeros((nh, nw), np.float32)\n        pad[:c.shape[0], :c.shape[1]] = c\n        c = pad\n    return zoom(c, (IMG / nh, IMG / nw), order=1)\n\ndef do_study(uid):\n    out = CACHE / f'{uid}.npz'\n    if out.exists():\n        return uid, 'skip'\n    sd = DATA / uid\n    if not sd.exists():\n        return uid, 'nodata'\n    rows = series[series.StudyInstanceUID == uid]\n    arrays = {}\n    centres = []\n    for (plane, fl) in SLOTS:\n        cand = rows[(rows.Anatomical_Plane == plane) & (rows.Fluid_Sensitive == fl)]\n        best = None\n        for _, r in cand.iterrows():\n            fs = sorted(glob.glob(str(sd / r.SeriesInstanceUID / '*.dcm')))\n            if len(fs) >= (len(best) if best else 0):\n                best = fs\n        if not best:\n            continue\n        recs = order_files(best)\n        if len(recs) < 3:\n            continue\n        n = len(recs)\n        lo, hi = int(BAND[0] * (n - 1)), int(BAND[1] * (n - 1))\n        pick = np.unique(np.linspace(lo, hi, S).astype(int))\n        sl = []\n        for k in pick:\n            a, ps = read_slice(recs[k][1])\n            sl.append(crop_resize(a, ps))\n        st = np.stack(sl)\n        p1, p99 = np.percentile(st, [1, 99])\n        st = np.clip((st - p1) / max(p99 - p1, 1e-3), 0, 1)\n        arrays[f'{plane}_{fl}'] = (st * 255).astype(np.uint8)\n        centres.append(np.mean([r[2][0] for r in recs]))\n    if not arrays:\n        return uid, 'empty'\n    lat = 'R' if (centres and np.median(centres) < -20) else ('L' if (centres and np.median(centres) > 20) else 'U')\n    np.savez_compressed(out, **arrays, lat=np.array(lat), n_slots=np.array(len(arrays)))\n    return uid, f'ok:{len(arrays)}'\n\nprint(f'cache functions ready | IMG={IMG} S={S} CROP={CROP_MM}mm workers={N_WORKERS}')\n"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"\"\"\"Cell 3: run the caching stage with a wall-clock budget.\n\nGold studies are cached first (guarantees a complete, honest evaluation set\nregardless of how much of the corpus we get through); the remaining studies\nare shuffled (seed fixed) so a time-limited partial cache is still an\nunbiased random subset of the corpus, not just \"whatever sorts first\".\n\"\"\"\nCACHE_TIME_BUDGET_S = float(os.environ.get('CACHE_TIME_BUDGET_S', 5.0 * 3600))\n\nall_uids = train.StudyInstanceUID.tolist()\ngold_list = [u for u in all_uids if u in gold_uid]\nrest = [u for u in all_uids if u not in gold_uid]\nrng = np.random.default_rng(2026)\nrest = list(rng.permutation(rest))\nordered_uids = gold_list + rest\ntodo = [u for u in ordered_uids if not (CACHE / f'{u}.npz').exists()]\nprint(f'{len(todo)} studies to cache (of {len(ordered_uids)}), time budget {CACHE_TIME_BUDGET_S:.0f}s')\n\nt0 = time.time()\nstat = {}\ndone_count = 0\nwith ProcessPoolExecutor(N_WORKERS) as ex:\n    futures = {ex.submit(do_study, u): u for u in todo}\n    for fut in futures:\n        pass  # submit-order iteration below via as_completed instead\n    from concurrent.futures import as_completed\n    for fut in as_completed(futures):\n        u, s = fut.result()\n        stat[s.split(':')[0]] = stat.get(s.split(':')[0], 0) + 1\n        done_count += 1\n        elapsed = time.time() - t0\n        if done_count % 50 == 0 or done_count == len(todo):\n            rate = done_count / elapsed\n            print(f'{done_count}/{len(todo)} {elapsed:.0f}s rate={rate:.2f}/s eta_full={(len(todo)-done_count)/max(rate,1e-6):.0f}s {stat}', flush=True)\n        if elapsed > CACHE_TIME_BUDGET_S:\n            print(f'TIME BUDGET REACHED at {done_count}/{len(todo)}, cancelling remaining', flush=True)\n            for f2 in futures:\n                f2.cancel()\n            break\n\ncached = sorted(p.stem for p in CACHE.glob('*.npz'))\ncached_gold = [u for u in cached if u in gold_uid]\nprint(f'DONE caching | total cached={len(cached)} | gold cached={len(cached_gold)}/{len(gold_uid)} | elapsed={time.time()-t0:.0f}s')\nassert len(cached_gold) == len(gold_uid), 'not all gold studies were cached -- evaluation would be incomplete'\n"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"\"\"\"Cell 4: multi-slot MIL model + dataset (adapted from nekkon/rsna_train.py).\n\nEncoder: timm EfficientNet-B3 (ImageNet) applied per slot to 3 adjacent\nslices as a 2.5D input; a per-finding attention head pools the six\nplane/fluid slots (masked softmax over slots present in a given study).\nTargets: weak_labels_v2 soft blend (soft BCE). Unchanged from the source\nrecipe apart from ARCH (B3, matching the published README) and reading the\ncache built in cell 3 instead of a local disk cache.\n\"\"\"\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport timm\nfrom sklearn.metrics import roc_auc_score\n\nARCH = os.environ.get('ARCH', 'efficientnet_b3')\nSLOT_NAMES = [f'{p}_{f}' for p in ('Sagittal', 'Coronal', 'Axial') for f in (1, 0)]\ndev = 'cuda' if torch.cuda.is_available() else 'cpu'\n\nclass DS(torch.utils.data.Dataset):\n    def __init__(s, uids, aug):\n        s.u = uids\n        s.aug = aug\n\n    def __len__(s):\n        return len(s.u)\n\n    def __getitem__(s, i):\n        u = s.u[i]\n        z = np.load(CACHE / f'{u}.npz')\n        x = np.zeros((len(SLOT_NAMES), 3, IMG, IMG), np.float32)\n        m = np.zeros(len(SLOT_NAMES), np.float32)\n        lat = str(z['lat'])\n        for k, sl in enumerate(SLOT_NAMES):\n            if sl not in z:\n                continue\n            a = z[sl].astype(np.float32) / 255.\n            n = a.shape[0]\n            c = n // 2 if not s.aug else np.random.randint(1, max(2, n - 1))\n            c = min(max(c, 1), n - 2)\n            tri = a[c - 1:c + 2]\n            if lat == 'R':\n                tri = tri[:, :, ::-1] if sl.split('_')[0] != 'Sagittal' else tri[::-1]\n            if s.aug:\n                sh = np.random.randint(-8, 9, 2)\n                tri = np.roll(tri, tuple(sh), axis=(1, 2))\n                tri = tri * np.random.uniform(.9, 1.1)\n            x[k] = tri\n            m[k] = 1\n        y = Y.loc[u].values.astype(np.float32)\n        return torch.from_numpy(x), torch.from_numpy(m), torch.from_numpy(y), u\n\nclass Net(nn.Module):\n    def __init__(s):\n        super().__init__()\n        s.enc = timm.create_model(ARCH, pretrained=True, num_classes=0)\n        d = s.enc.num_features\n        H = 256\n        s.proj = nn.Linear(d, H)\n        s.slot_emb = nn.Parameter(torch.zeros(len(SLOT_NAMES), H))\n        s.q = nn.Parameter(torch.randn(len(LAB), H) * .02)\n        s.out = nn.ModuleList([nn.Linear(H, 1) for _ in LAB])\n\n    def forward(s, x, m):\n        B, K = x.shape[:2]\n        h = s.enc(x.flatten(0, 1)).view(B, K, -1)\n        h = s.proj(h) + s.slot_emb\n        att = torch.einsum('bkh,oh->bok', h, s.q) / math.sqrt(h.shape[-1])\n        att = att.masked_fill(m[:, None, :] == 0, -1e4).softmax(-1)\n        c = torch.einsum('bok,bkh->boh', att, h)\n        return torch.cat([s.out[o](c[:, o]) for o in range(len(LAB))], 1)\n\ndef evaluate(model, uids, bs, n_workers, gold=False):\n    model.eval()\n    P = []\n    U = []\n    dl = torch.utils.data.DataLoader(DS(uids, False), batch_size=bs, num_workers=n_workers)\n    with torch.no_grad(), torch.autocast('cuda', dtype=torch.float16):\n        for x, m, y, u in dl:\n            out = torch.sigmoid(model(x.to(dev), m.to(dev))).float().cpu().numpy()\n            P.append(np.nan_to_num(out, nan=0.5))\n            U += list(u)\n    P = np.concatenate(P)\n    T = (G.loc[U].values if gold else Y.loc[U].values)\n    aucs = []\n    for j in range(len(LAB)):\n        t = T[:, j]\n        tb = (t >= 0.5).astype(int) if not gold else t.astype(int)\n        aucs.append(roc_auc_score(tb, P[:, j]) if 0 < tb.sum() < len(tb) else np.nan)\n    return float(np.nanmean(aucs)), aucs, P, U\n\nprint(f'train functions ready | ARCH={ARCH} device={dev}')\n"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"\"\"\"Cell 5: train all 5 folds on whatever got cached, evaluate on gold58 each,\nthen rank-average the 5 folds' gold predictions (nekkon's own methodology:\n\"the metric is AUC, so ranks are the only thing it reads\").\n\"\"\"\nN_FOLDS = int(os.environ.get('N_FOLDS', 5))\nEPOCHS = int(os.environ.get('EPOCHS', 10))\nBS = int(os.environ.get('BS', 16))\nLR_HEAD = 1e-3\nLR_BB = float(os.environ.get('LR_BB', 1e-4))\nN_DL_WORKERS = min(os.cpu_count() or 4, 4)\n\nhave = [p.stem for p in CACHE.glob('*.npz')]\nhave = [u for u in have if u in Y.index]\ngo = [u for u in have if u in gold_uid]\nprint(f'cached studies {len(have)} | gold among them {len(go)}', flush=True)\nassert len(go) == 58, f'expected all 58 gold studies cached, got {len(go)}'\n\ndef train_one_fold(FOLD):\n    tr = [u for u in have if u not in gold_uid and fold_of[u] != FOLD]\n    va = [u for u in have if u not in gold_uid and fold_of[u] == FOLD]\n    print(f'[fold {FOLD}] train {len(tr)}  weak-holdout {len(va)}  gold {len(go)}', flush=True)\n\n    model = Net().to(dev)\n    opt = torch.optim.AdamW(\n        [{'params': model.enc.parameters(), 'lr': LR_BB},\n         {'params': [p for n, p in model.named_parameters() if not n.startswith('enc.')], 'lr': LR_HEAD}],\n        weight_decay=1e-4,\n    )\n    steps = EPOCHS * math.ceil(len(tr) / BS)\n    sched = torch.optim.lr_scheduler.OneCycleLR(opt, max_lr=[LR_BB, LR_HEAD], total_steps=steps, pct_start=.15)\n    scaler = torch.amp.GradScaler()\n    dl = torch.utils.data.DataLoader(\n        DS(tr, True), batch_size=BS, shuffle=True, num_workers=N_DL_WORKERS,\n        drop_last=True, persistent_workers=True,\n    )\n\n    best = -1\n    best_go_each = None\n    best_goP = None\n    log = []\n    t_fold0 = time.time()\n    for ep in range(EPOCHS):\n        model.train()\n        t0 = time.time()\n        tl = 0\n        for x, m, y, _ in dl:\n            x, m, y = x.to(dev, non_blocking=True), m.to(dev), y.to(dev)\n            with torch.autocast('cuda', dtype=torch.float16):\n                loss = F.binary_cross_entropy_with_logits(model(x, m), y)\n            if not torch.isfinite(loss):\n                opt.zero_grad(set_to_none=True)\n                sched.step()\n                continue\n            opt.zero_grad(set_to_none=True)\n            scaler.scale(loss).backward()\n            scaler.unscale_(opt)\n            torch.nn.utils.clip_grad_norm_(model.parameters(), 2.0)\n            scaler.step(opt)\n            scaler.update()\n            sched.step()\n            tl += loss.item()\n        va_auc, va_each, _, _ = evaluate(model, va, BS, N_DL_WORKERS) if va else (float('nan'), [float('nan')] * len(LAB), None, None)\n        go_auc, go_each, goP, goU = evaluate(model, go, BS, N_DL_WORKERS, gold=True)\n        log.append(dict(ep=ep, loss=tl / len(dl), weak_holdout_auc=va_auc, gold_auc=go_auc, sec=time.time() - t0))\n        print(f'[fold {FOLD}] ep{ep} loss {tl/len(dl):.4f}  weak-holdout AUC {va_auc:.4f}  GOLD-58 AUC {go_auc:.4f}  {time.time()-t0:.0f}s', flush=True)\n        save_now = (va_auc > best) if va else (ep == EPOCHS - 1)\n        if save_now:\n            best = va_auc if va else best\n            best_go_each = go_each\n            best_goP = goP[[goU.index(u) for u in go]]\n            torch.save(model.state_dict(), WORK / f'model_effspec_f{FOLD}.pt')\n    json.dump(\n        dict(log=log, gold_per_finding=dict(zip(LAB, best_go_each)), n_train=len(tr),\n             n_weak_holdout=len(va), n_gold=len(go), fold=FOLD),\n        open(WORK / f'model_effspec_f{FOLD}.json', 'w'), indent=1,\n    )\n    np.save(WORK / f'pred_effspec_f{FOLD}_gold.npy', best_goP)\n    print(f'[fold {FOLD}] done in {time.time()-t_fold0:.0f}s, best weak-holdout {best:.4f}', flush=True)\n    return best_goP\n\nfold_preds = []\nt_all0 = time.time()\nfor FOLD in range(N_FOLDS):\n    fold_preds.append(train_one_fold(FOLD))\nprint(f'all {N_FOLDS} folds done in {time.time()-t_all0:.0f}s total', flush=True)\n\nfold_preds = np.stack(fold_preds)  # (N_FOLDS, 58, 12)\ngold_targets = G.loc[go].values.astype(int)\n\ndef rank_avg(preds_stack):\n    # preds_stack: (n_folds, n_studies, n_labels) -> percentile-rank each fold's column, then average\n    out = np.zeros(preds_stack.shape[1:])\n    for f in range(preds_stack.shape[0]):\n        for j in range(preds_stack.shape[2]):\n            out[:, j] += pd.Series(preds_stack[f, :, j]).rank(pct=True).to_numpy()\n    return out / preds_stack.shape[0]\n\nfinal_pred = rank_avg(fold_preds)\nfinal_auc = {}\nfor j, lab in enumerate(LAB):\n    t = gold_targets[:, j]\n    final_auc[lab] = roc_auc_score(t, final_pred[:, j]) if 0 < t.sum() < len(t) else float('nan')\n\nprint('=== FINAL rank-averaged 5-fold gold58 AUC ===')\nprint(json.dumps(final_auc, indent=2))\nprint(f'macro AUC: {np.nanmean(list(final_auc.values())):.4f}')\nprint('--- comparison to nekkon n=929 (1-fold-equivalent, README/notebook) ---')\nprint('Lateral OA  nekkon(n=929): image 0.77 vs report 0.56')\nprint('Effusion    nekkon(n=929): image 0.72 vs report 0.67')\nprint('this run    n_train per fold ~', len(have) - 58 - len(have) // 5)\nprint('Lateral OA  this run (5-fold rank-avg):', final_auc.get('Lateral OA'))\nprint('Effusion    this run (5-fold rank-avg):', final_auc.get('Effusion'))\n\njson.dump(\n    dict(final_auc_5fold_rankavg=final_auc, n_cached=len(have), n_gold=len(go),\n         epochs=EPOCHS, n_folds=N_FOLDS, bs=BS, arch=ARCH),\n    open(WORK / 'effspec_summary.json', 'w'), indent=2,\n)\nnp.save(WORK / 'effspec_gold_uids.npy', np.array(go))\nnp.save(WORK / 'effspec_final_pred_rankavg.npy', final_pred)\n"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"\"\"\"Cell 6: cleanup. Remove the large DICOM cache so kernel output stays small\n-- everything needed downstream (5 fold model weights, per-fold gold\npredictions/JSON, the rank-averaged summary) has already been written above\nand survives this.\n\"\"\"\nimport shutil\ncache_size_mb = sum(f.stat().st_size for f in CACHE.glob('*.npz')) / 1e6\nprint(f'removing cache ({cache_size_mb:.0f} MB, {len(list(CACHE.glob(\"*.npz\")))} files)')\nshutil.rmtree(CACHE, ignore_errors=True)\n\nkept = sorted(p.name for p in WORK.glob('*') if p.is_file())\nprint('kernel output files kept:', kept)\ntotal_kept_mb = sum((WORK / f).stat().st_size for f in kept) / 1e6\nprint(f'total kept output size: {total_kept_mb:.1f} MB')\n"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12"}},"nbformat":4,"nbformat_minor":5}