{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# M2 -- baseline, end to end (train)\n#\nOne model, one fold, the simplest thing that can score. The point of this notebook is **not\nits AUC** -- it is to produce the numbers PLAN.md currently has to guess: real GPU-hours per\nepoch, real cache read throughput, a real per-column AUC baseline, and a checkpoint that the\nsubmission notebook can actually run.\n#\n**Architecture** (reasoning in `tools/knee_model.py`): 32 cached slices of one sagittal\nfluid-sensitive series -> 2.5D triples (i-1, i, i+1) as RGB channels -> ImageNet-pretrained 2D\nbackbone shared across slices -> gated-attention pooling over the 32 embeddings -> 12-way\nmulti-label head. **No horizontal flip**: it mirrors the knee, which swaps medial and lateral,\nwhich is 4 of the 12 scored columns.\n#\n## Before running\n#\n| need | how |\n|---|---|\n| GPU (T4 x2) | Settings -> Accelerator. Spends the 30 h/week quota. |\n| Internet **ON** | `timm` downloads the ImageNet weights. Off only in the submission notebook. |\n| `knee-tools` dataset | this repo, so `tools/knee_model.py` and `results/labels_v1.csv` are importable — mounts at `/kaggle/input/datasets/tamerlanomralinov/knee-tools` |\n| `knee-cache` dataset(s) | the M1 output. Add every shard; `cache_globs` picks them all up. |\n#\n## Why this run is cheaper than the first attempt\n#\nThe first attempt died at epoch 4 of 8 with a bare `Kernel died` and no Python traceback --\nwhich is not an exception, it is the Linux OOM killer, so nothing in the log says why. It was\n**host** RAM, not VRAM: the DataLoader shipped the finished network input (32 slices x 3\nchannels x 288 px x float32 = 255 MB per batch, 510 MB for a validation batch), three loaders\neach kept `persistent_workers` alive with `prefetch_factor=4`, and `pin_memory` holds a second\ncopy of everything queued. Four changes, none of which touch what the model sees:\n#\n| change | effect |\n|---|---|\n| loader yields **uint8**; the 2.5D triple and the ImageNet normalisation move to the GPU | 255 MB -> 21 MB per batch, and 1/12 the PCIe traffic |\n| `prefetch_factor` 4 -> 2, `persistent_workers` on the train loader only | validation workers exit and give their queues back between epochs |\n| augmentation composes rotation + crop + aspect into **one** `grid_sample` | was two full bilinear passes over the 32-slice stack per study |\n| `torch.set_num_threads(1)` per worker, cv2 for JPEG decode | 3 workers x 4 OMP threads on 4 vCPUs was mostly contention |\n#\nThe epoch loop now prints host RAM next to the loss and writes `m2_last.pt` after every epoch,\nso the next crash costs one epoch instead of the whole run. `CFG['resume']` picks it back up.\n#\n## How to read the output\n#\nModel selection happens on a 10% holdout of the **noisy** labels. The 58 gold studies are\nnever trained on and never selected on -- they are a read-out only, and at n=58 each column\ncarries roughly +/-0.06. Tune on gold and you will overfit it within three decisions.","metadata":{"_uuid":"a791859a-1acc-4fd9-b9d7-e2e5aa3bcb73","_cell_guid":"15801c7a-00dd-4926-9f42-045ea1b25931","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# ----------------------------------------------------------------------------------\n# Config. Everything tweakable is here; nothing below hardcodes a choice.\n# ----------------------------------------------------------------------------------\nCFG = dict(\n    # --- data -------------------------------------------------------------------\n    tools_path=\"/kaggle/input/datasets/tamerlanomralinov/knee-tools/tools\",\n    labels_csv=\"/kaggle/input/datasets/tamerlanomralinov/knee-tools/results/labels_v1.csv\",\n    # both mount layouts: this account gets datasets/<owner>/<slug>, some sessions get the\n    # flat /kaggle/input/<slug>. Globbing both costs nothing and a glob that matches nothing\n    # is silent -- it just trains on however much of the cache it happened to find.\n    cache_globs=[\"/kaggle/input/notebooks/tamerlanomralinov/knee-cache/cache\",\n                 \"/kaggle/input/datasets/*/knee-cache*\",\n                 \"/kaggle/input/knee-cache*/cache\", \"/kaggle/input/knee-cache*\"],\n    slot_priority=(\"sag_fs\", \"sag_nf\"),   # PLAN M2: sagittal fluid-sensitive, nf as fallback\n    size=288,                             # the cache is 288 (CLIMB Rung 2, the 0.903 recipe's\n                                          # resolution): train random-crops, val resizes. This\n                                          # read 224 while a comment claimed the cache was 256\n                                          # -- a silent downscale of every training image.\n    group_by_site=True,                   # whole sites stay on one side of the split. False\n                                          # falls back to the i.i.d. md5 buckets, which read\n                                          # optimistic; see ANALYSIS.md §5 and tools/folds.py\n    val_bucket=9,                         # 1 of 10 buckets -> ~430 studies for selection\n    mask_unresolved_lateral=True,         # drop the 4 medial/lateral columns on studies whose\n                                          # laterality could not be resolved (1.5% of the\n                                          # corpus); their image may mirror their label\n    source_weight={\"manual\": 1.0, \"rules\": 1.0},   # manual is 0.889 vs 0.774 -- try 2.0 later\n\n    # --- model ------------------------------------------------------------------\n    backbone=\"resnet34\",                  # fast, robust, 2.5D-friendly. A baseline, not a claim.\n    pretrained=True,\n    attn_hidden=128,\n    dropout=0.1,\n    per_column_attn=False,                # 12 attention heads instead of 1 -- likely next step\n\n    # --- optimisation -----------------------------------------------------------\n    epochs=8,\n    batch_size=8,                         # x32 slices = 256 images per forward pass\n    lr=3e-4,\n    weight_decay=1e-4,\n    warmup_frac=0.05,\n    grad_clip=1.0,\n    amp=True,\n    multi_gpu=True,                       # DataParallel over 2xT4: ~1.6x, halves quota burn\n    channels_last=True,                   # NHWC: what cudnn's fp16 kernels want on a T4\n    seed=42,\n\n    # --- the loader, which is what actually killed the first run -----------------\n    # A batch is batch_size x 32 slices x 288 px. Shipped as the finished network input\n    # (3 channels, float32) that is 255 MB; shipped as uint8 with the 2.5D triple and the\n    # ImageNet normalisation done on the GPU it is 21 MB. Multiply by workers x prefetch x\n    # (loader queue + pinned staging copy) and the first spelling does not fit in the 29 GB\n    # a Kaggle T4x2 session has. See knee_model.to_input / to_network.\n    num_workers=4,                        # 4 vCPU; the loader is the bottleneck, not the GPU\n    prefetch=2,                           # batches in flight per worker (was 4)\n    val_workers=2,                        # validation loaders are NOT persistent: their\n                                          # workers exit between epochs and give the RAM back\n    # --- switches ---------------------------------------------------------------\n    smoke_test=True,                      # overfit 32 studies first; catches pipeline bugs\n    max_studies=None,                     # set e.g. 400 for a fast dry run\n    resume=True,                          # pick up m2_last.pt if the session already wrote one\n    time_budget_h=None,                   # stop cleanly after N hours and still write artefacts\n    out_dir=\"/kaggle/working\",\n)\nCFG","metadata":{"_uuid":"f7ae11e3-515a-49f0-bf94-083ff4ea8645","_cell_guid":"8c0ff420-fbce-4d07-8fab-48f477a367f3","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport sys\nimport glob\nimport math\nimport time\nimport random\n\nimport numpy as np\nimport pandas as pd\nimport torch\nimport torch.nn as nn\nimport matplotlib.pyplot as plt\nfrom torch.utils.data import DataLoader\n\nsys.path.insert(0, CFG[\"tools_path\"])\nfrom knee_model import (LABELS, LATERAL_COLUMNS, IMAGENET_MEAN, IMAGENET_STD, KneeStudies,\n                        KneeNet, masked_bce, binarise, bucket, grouped_buckets, group_leak,\n                        index_cache, load_study, slots_of, per_column_auc, auc_table,\n                        to_network, worker_init)\n\ntry:\n    import psutil\n    _VM = psutil.virtual_memory\n\n    def ram():\n        \"\"\"Host RAM used/total. The number the OOM killer reads -- and the one that ended the\n        first attempt of this notebook at epoch 4 with a bare 'Kernel died': no Python\n        traceback, because the process was SIGKILLed rather than raised in.\"\"\"\n        v = _VM()\n        return (v.total - v.available) / 1e9, v.total / 1e9\nexcept ImportError:                       # psutil ships with the Kaggle image; be safe anyway\n    def ram():\n        return float(\"nan\"), float(\"nan\")\n\n\ndef ram_str():\n    u, t = ram()\n    return f\"RAM {u:.1f}/{t:.1f} GB\"\n\n\ndef seed_all(s):\n    random.seed(s)\n    np.random.seed(s)\n    torch.manual_seed(s)\n    torch.cuda.manual_seed_all(s)\n\n\ndef make_scaler(enabled):\n    \"\"\"torch moved GradScaler between 2.x minors; support both rather than pin a version.\"\"\"\n    try:\n        return torch.amp.GradScaler(\"cuda\", enabled=enabled)\n    except (AttributeError, TypeError):\n        return torch.cuda.amp.GradScaler(enabled=enabled)\n\n\nseed_all(CFG[\"seed\"])\ntorch.backends.cudnn.benchmark = True\ndevice = \"cuda\" if torch.cuda.is_available() else \"cpu\"\nprint(f\"torch {torch.__version__} | {torch.cuda.device_count()} GPU(s) \"\n      f\"{[torch.cuda.get_device_name(i) for i in range(torch.cuda.device_count())]}\")\nprint(f\"{os.cpu_count()} vCPU | {ram_str()} at import\")\nassert device == \"cuda\", \"turn the GPU accelerator on -- this is unusable on CPU\"","metadata":{"_uuid":"c31eb77d-cd7d-404d-a92e-5f8d0fd76665","_cell_guid":"e5d0516b-ee25-46f4-afca-732e292d506b","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ----------------------------------------------------------------------------------\n# Index the cache, join the labels, and be explicit about what gets dropped. Slot coverage\n# comes from the M1 manifest (instant) rather than from unpickling 10 GB.\n# ----------------------------------------------------------------------------------\npaths = index_cache(CFG[\"cache_globs\"])\nprint(f\"cache: {len(paths)} study pickles\")\nassert paths, f\"nothing under {CFG['cache_globs']} -- add the M1 cache dataset(s)\"\n\nlabels = pd.read_csv(CFG[\"labels_csv\"])\nassert list(labels.columns) == [\"StudyInstanceUID\", \"source\"] + LABELS, labels.columns.tolist()\nprint(f\"labels: {len(labels)} rows, sources {labels['source'].value_counts().to_dict()}\")\n\nman_files = [p for pat in CFG[\"cache_globs\"]\n             for p in glob.glob(os.path.join(pat, \"manifest_train_*.csv\"))]\nif man_files:\n    man = (pd.concat([pd.read_csv(p) for p in man_files], ignore_index=True)\n             .drop_duplicates(\"StudyInstanceUID\"))\n    ok = man[man[\"error\"].isna() | (man[\"error\"] == \"\")]\n    slots = {r.StudyInstanceUID: set(str(r.slots).split(\"|\")) for r in ok.itertuples()}\n    # site_id and lat_route are recorded at build time because neither can be recovered from\n    # the pickles afterwards -- the header tags are gone. Read them here or not at all.\n    site_of = dict(zip(ok[\"StudyInstanceUID\"], ok[\"site_id\"].fillna(\"\").astype(str))) \\\n        if \"site_id\" in ok.columns else {}\n    unresolved = set(ok.loc[ok[\"lat_route\"] == \"unresolved\", \"StudyInstanceUID\"])\n    print(f\"manifest: {len(man_files)} shards, {len(ok)} ok / {len(man)} rows\")\n    print(f\"laterality routes: {ok['lat_route'].value_counts().to_dict()}\")\n    print(f\"-> unresolved laterality {len(unresolved)} \"\n          f\"({100*len(unresolved)/max(len(ok),1):.1f}%): those studies keep their original\"\n          f\"\\n   orientation, so for them the image may be the mirror of what the label says.\")\n    n_sites = len({s for s in site_of.values() if s})\n    print(f\"sites: {n_sites} distinct keys over {len(site_of)} studies\")\nelse:\n    print(\"no M1 manifest found -- scanning pickle keys instead (slower)\")\n    slots = {uid: set(slots_of(p)) for uid, p in paths.items()}\n    site_of, unresolved, n_sites = {}, set(), 0\n    print(\"!! no site_id and no lat_route without the manifest: the split falls back to i.i.d.\"\n          \"\\n   and unresolved-laterality studies cannot be masked. Attach the cache dataset\"\n          \"\\n   including its manifest_train_*.csv rather than only the pickles.\")\n\nwant = set(CFG[\"slot_priority\"])\nhave_slot = [u for u in labels[\"StudyInstanceUID\"] if u in paths and (slots.get(u, set()) & want)]\nin_cache = sum(u in paths for u in labels[\"StudyInstanceUID\"])\nprint(f\"\\nof {len(labels)} labelled studies: {in_cache} cached, \"\n      f\"{len(have_slot)} with {'/'.join(CFG['slot_priority'])} \"\n      f\"({100*len(have_slot)/len(labels):.1f}%)\")\nprint(f\"   missing from cache      : {len(labels) - in_cache}\")\nprint(f\"   cached, no sagittal slot: {in_cache - len(have_slot)}\")\nprint(\"-> a study without the slot cannot train this baseline at all. If that share is large,\"\n      \"\\n   fix M1 before spending GPU quota here.\")\n\nassert have_slot, (\"no cached study has the sagittal slot -- check the M1 output with \"\n                   \"tools/cache_gate.py before going further\")\n\n# the manifest is a claim about the pickles; check it against the pickles on a sample, and\n# measure the real cache read rate while doing so (the dataloader budget depends on it)\nt0 = time.time()\nsample = have_slot[:20]\nfor uid in sample:\n    vol, slot = load_study(paths[uid], CFG[\"slot_priority\"])\n    assert vol is not None and slot in want, f\"{uid}: manifest says {slots.get(uid)}, got {slot}\"\n    assert vol.dtype == np.uint8 and vol.ndim == 3, (uid, vol.shape, vol.dtype)\ndt = (time.time() - t0) / len(sample)\ncache_shape = tuple(int(v) for v in vol.shape)   # goes into the checkpoint: the submission\n                                                 # notebook must rebuild a cache of this shape\nprint(f\"\\ncache read+decode: {dt*1000:.0f} ms/study ({vol.shape} uint8), so one epoch of \"\n      f\"{len(have_slot)} studies needs {dt*len(have_slot)/60:.1f} CPU-min of decode; \"\n      f\"{CFG['num_workers']} workers -> {dt*len(have_slot)/60/CFG['num_workers']:.1f} min. \"\n      f\"If that exceeds the GPU time per epoch, the loader is the bottleneck.\")\n\ndf = labels[labels[\"StudyInstanceUID\"].isin(have_slot)].reset_index(drop=True)\nif CFG[\"max_studies\"]:                    # keep every gold row: it is the whole read-out\n    gold_rows = df[df[\"source\"] == \"gold\"]\n    rest = df[df[\"source\"] != \"gold\"].sample(\n        min(CFG[\"max_studies\"], (df[\"source\"] != \"gold\").sum()), random_state=CFG[\"seed\"])\n    df = pd.concat([gold_rows, rest]).reset_index(drop=True)\n    print(f\"\\nDRY RUN: cut to {len(df)} studies ({len(gold_rows)} gold kept)\")","metadata":{"_uuid":"0b1ce07b-5e83-4442-8ac2-fb5fc17e27e0","_cell_guid":"ca957c99-8304-4f78-8c68-b747727591a3","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ----------------------------------------------------------------------------------\n# Split. gold is held out entirely (ground truth, read-out only); one bucket of the noisy rows\n# is the selection set.\n#\n# GROUPED BY SITE where the manifest allows it. A site is a scanner, a population and a\n# protocol at once, so an i.i.d. split lets a study leak into its own validation fold through\n# any of the three and reads optimistic by an amount nobody can estimate in advance. That is\n# also the split whose optimism would be mistaken for a labels problem when the LB comes in\n# lower. tools/folds.py holds the assignment; tools/test_folds.py asserts it does not leak.\n# ----------------------------------------------------------------------------------\nuse_groups = bool(CFG[\"group_by_site\"] and n_sites > 1)\nif use_groups:\n    buckets = grouped_buckets(df[\"StudyInstanceUID\"], site_of, n=10)\n    df[\"bucket\"] = [buckets[u] for u in df[\"StudyInstanceUID\"]]\nelse:\n    df[\"bucket\"] = [bucket(u) for u in df[\"StudyInstanceUID\"]]\n    print(\"!! i.i.d. split: \"\n          + (\"group_by_site is off\" if not CFG[\"group_by_site\"] else \"no usable site keys\")\n          + \". val_noisy will read optimistic against the LB.\")\n\nis_gold = df[\"source\"] == \"gold\"\nval_gold = df[is_gold].reset_index(drop=True)\nval_noisy = df[~is_gold & (df[\"bucket\"] == CFG[\"val_bucket\"])].reset_index(drop=True)\ntrain = df[~is_gold & (df[\"bucket\"] != CFG[\"val_bucket\"])].reset_index(drop=True)\n\nfor a, b in ((train, val_gold), (train, val_noisy), (val_gold, val_noisy)):\n    assert not set(a[\"StudyInstanceUID\"]) & set(b[\"StudyInstanceUID\"])\nassert len(val_gold) and len(val_noisy) and len(train)\nprint(f\"train {len(train)} | val_noisy {len(val_noisy)} | val_gold {len(val_gold)}\"\n      f\" | split {'grouped by site' if use_groups else 'i.i.d.'}\")\n\nif use_groups:\n    # The property, asserted rather than described. A grouped split that quietly leaks looks\n    # exactly like one that does not, right up until the leaderboard disagrees with CV.\n    leak = group_leak(train[\"StudyInstanceUID\"], val_noisy[\"StudyInstanceUID\"], site_of)\n    assert not leak, f\"{len(leak)} site(s) on both sides of the split: {sorted(leak)[:5]}\"\n    v_sites = {site_of.get(u, \"\") for u in val_noisy[\"StudyInstanceUID\"]} - {\"\"}\n    print(f\"   val_noisy holds {len(v_sites)} whole site(s), no site shared with train\")\n    if len(v_sites) < 3:\n        print(\"   !! selection set spans fewer than 3 sites -- it is now measuring those \"\n              \"scanners\\n      as much as the model. Read val_noisy deltas under ~0.01 as noise.\")\n    # gold is a read-out, never selected on, but if it shares sites with train say so plainly\n    g_leak = group_leak(train[\"StudyInstanceUID\"], val_gold[\"StudyInstanceUID\"], site_of)\n    if g_leak:\n        print(f\"   note: gold shares {len(g_leak)} site(s) with train. gold is ground truth and\"\n              f\"\\n   is never selected on, so this is a read-out caveat, not a leak in selection.\")\n\n\nLATERAL_IDX = [LABELS.index(k) for k in LATERAL_COLUMNS]\n\n\ndef targets(frame):\n    \"\"\"-> (binary targets, mask, soft targets). Soft is what the loss sees; binary is what AUC\n    is measured against; mask is 0 where the label source had no evidence at all.\n\n    Also masks the four medial/lateral columns on studies whose laterality was never resolved.\n    Their image keeps the scanner's original orientation, so it may be the mirror of what the\n    label describes -- and a mirrored medial meniscus is a lateral one. Training on those\n    entries teaches a coin flip; scoring them measures one. The other eight columns name a\n    structure rather than a side of the frame and stay supervised.\n    \"\"\"\n    y_soft = frame[LABELS].to_numpy(np.float32)\n    y_bin, mask = binarise(y_soft)\n    if CFG[\"mask_unresolved_lateral\"] and unresolved:\n        hit = frame[\"StudyInstanceUID\"].isin(unresolved).to_numpy()\n        dropped = int(mask[np.ix_(hit, LATERAL_IDX)].sum())\n        mask[np.ix_(hit, LATERAL_IDX)] = 0.0\n        if hit.any():\n            print(f\"   masked {dropped} labelled cells on {int(hit.sum())} unresolved-laterality\"\n                  f\" studies ({'/'.join(LATERAL_COLUMNS)})\")\n    return y_bin, mask, y_soft\n\n\nytr_bin, mtr, ytr_soft = targets(train)\nyvn_bin, mvn, _ = targets(val_noisy)\nyvg_bin, mvg, _ = targets(val_gold)\n\nprint(f\"\\nsupervision actually available per column (train, of {len(train)} studies):\")\nprint(f\"{'label':<18}{'n_labelled':>11}{'n_pos':>7}{'prev':>7}  |{'gold n_pos':>11}\")\nfor j, k in enumerate(LABELS):\n    n = int(mtr[:, j].sum())\n    pos = int(ytr_bin[mtr[:, j] > 0, j].sum())\n    print(f\"{k:<18}{n:>11}{pos:>7}{pos/max(n,1):>7.3f}  |{int(yvg_bin[:, j].sum()):>11}\")\nprint(\"\\n-> n_labelled < n_train because 0.500 means 'the report said nothing', which is not\"\n      \"\\n   evidence of absence; those entries are masked out of the loss rather than taught\"\n      \"\\n   as 0.5. Set NEUTRAL handling in tools/knee_model.py if you want to test the other way.\"\n      \"\\n-> gold n_pos is what every gold AUC below rests on. Single digits means +/-0.10, easily.\")","metadata":{"_uuid":"4eda63ef-3cce-4708-8d58-efe1cd6f6754","_cell_guid":"5e9fceb2-616e-4172-ac87-9af9577a6eef","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ----------------------------------------------------------------------------------\n# Loaders, then LOOK at a batch. A mirrored or misordered cache produces perfectly plausible\n# images and a quietly worse score; this is the cheapest place left to catch it.\n# ----------------------------------------------------------------------------------\nAUG = dict(max_rot=10.0, scale=(0.8, 1.0), bc=0.15)   # no hflip -- see knee_model's docstring\n\n\ndef make_loader(frame, y=None, mask=None, train_mode=False, bs=None, workers=None):\n    \"\"\"Workers yield uint8 (S, H, W) per study; KneeNet.to_network builds the 2.5D triple and\n    normalises on the GPU. Bytes per batch, batch_size=8: 21 MB instead of 255 MB.\n\n    persistent_workers is deliberately train-only. The validation loaders are touched twice per\n    epoch and holding their workers -- and their prefetch queues, and the pinned copy of every\n    queued batch -- alive for the other 12 minutes is exactly the RAM that killed the run at the\n    epoch-4 train->validate boundary.\"\"\"\n    w = frame[\"source\"].map(CFG[\"source_weight\"]).fillna(1.0).to_numpy(np.float32) \\\n        if train_mode else None\n    ds = KneeStudies(frame[\"StudyInstanceUID\"], paths, labels=y, masks=mask, weights=w,\n                     size=CFG[\"size\"], train=train_mode,\n                     slot_priority=CFG[\"slot_priority\"], aug=AUG if train_mode else None,\n                     as_uint8=True)\n    nw = CFG[\"num_workers\"] if workers is None else workers\n    extra = dict(persistent_workers=bool(train_mode), prefetch_factor=CFG[\"prefetch\"],\n                 worker_init_fn=worker_init) if nw else {}\n    return DataLoader(ds, batch_size=bs or CFG[\"batch_size\"], shuffle=train_mode,\n                      num_workers=nw, pin_memory=True, drop_last=train_mode, **extra)\n\n\ntrain_loader = make_loader(train, ytr_soft, mtr, train_mode=True)\nvn_loader = make_loader(val_noisy, yvn_bin, mvn, bs=2 * CFG[\"batch_size\"],\n                        workers=CFG[\"val_workers\"])\nvg_loader = make_loader(val_gold, yvg_bin, mvg, bs=2 * CFG[\"batch_size\"],\n                        workers=CFG[\"val_workers\"])\n\nmb = lambda t: t.numel() * t.element_size() / 1e6                                # noqa: E731\nbatch = next(iter(train_loader))\nxb = to_network(batch[\"x\"])               # exactly what the GPU builds, built here on the CPU\nprint(f\"loader x {tuple(batch['x'].shape)} {batch['x'].dtype} {mb(batch['x']):.0f} MB/batch \"\n      f\"-> network x {tuple(xb.shape)} {xb.dtype} {mb(xb):.0f} MB\")\nprint(f\"x mean {xb.mean():.3f} std {xb.std():.3f} | y {tuple(batch['y'].shape)} \"\n      f\"| mask kept {batch['mask'].mean():.2f} | {ram_str()}\")\nprint(f\"in flight at most {CFG['num_workers']}x{CFG['prefetch']} train + \"\n      f\"{CFG['val_workers']}x{CFG['prefetch']} val batches, x2 for the pinned staging copy \"\n      f\"-> ~{2*mb(batch['x'])*(CFG['num_workers']*CFG['prefetch'] + 2*CFG['val_workers']*CFG['prefetch'])/1000:.1f} GB \"\n      f\"of loader queue at the worst moment.\")\n# cache_shape[0], not a literal 32: knee_cache.SLICES is per-plane and has changed once\n# already. A hardcoded slice count turns \"the cache was rebuilt with a different depth\" into\n# a confusing shape error three cells later instead of here.\nassert tuple(batch[\"x\"].shape[1:]) == (cache_shape[0], CFG[\"size\"], CFG[\"size\"]), (\n    batch[\"x\"].shape, cache_shape)\nassert tuple(xb.shape[1:]) == (cache_shape[0], 3, CFG[\"size\"], CFG[\"size\"]), xb.shape\n\n\ndef show(x, title):\n    \"\"\"x: (S,3,H,W) normalised. Row 1: every 4th slice. Row 2: the 2.5D triple of the mid slice.\"\"\"\n    d = (x * IMAGENET_STD + IMAGENET_MEAN).clamp(0, 1)\n    mid = x.shape[0] // 2\n    panels = [(f\"slice {i}\", d[i, 1], \"gray\") for i in range(0, x.shape[0], 4)]\n    panels += [(f\"ch{c} = slice {mid-1+c}\", d[mid, c], \"gray\") for c in range(3)]\n    panels += [(\"|ch2 - ch0|\", (d[mid, 2] - d[mid, 0]).abs(), \"magma\")]\n    fig, ax = plt.subplots(2, 8, figsize=(16, 4.6))\n    for a, p in zip(ax.ravel(), panels):\n        a.imshow(p[1], cmap=p[2])\n        a.set_title(p[0], fontsize=8)\n    for a in ax.ravel():\n        a.axis(\"off\")\n    fig.suptitle(f\"{title}\\nrow 1: slice order (lateral -> medial after canonicalisation).  \"\n                 \"row 2: the 2.5D triple -- ch0/ch1/ch2 must be adjacent, not identical.\",\n                 fontsize=9)\n    plt.tight_layout()\n    plt.show()\n\n\nshow(xb[0], f\"train sample ...{batch['uid'][0][-12:]} (augmented)\")\nprint(\"Check by eye: these are knees; consecutive slices move smoothly; |ch2-ch0| is small but\"\n      \"\\nnot zero; handedness is consistent across studies. cache_gate.py's L-vs-R grid is\"\n      \"\\nthe authoritative laterality check -- this is a second pair of eyes on the tensor itself.\")","metadata":{"_uuid":"1471a6c4-7f4f-4c06-bb6d-6a8c5268e803","_cell_guid":"b48c5f01-fe78-4e4c-988e-858ddfc5d3b1","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ----------------------------------------------------------------------------------\n# Model, optimiser, schedule\n# ----------------------------------------------------------------------------------\ndef build():\n    m = KneeNet(backbone=CFG[\"backbone\"], pretrained=CFG[\"pretrained\"], n_out=len(LABELS),\n                hidden=CFG[\"attn_hidden\"], drop=CFG[\"dropout\"],\n                per_column_attn=CFG[\"per_column_attn\"],\n                channels_last=CFG[\"channels_last\"]).to(device)\n    if CFG[\"multi_gpu\"] and torch.cuda.device_count() > 1:\n        m = nn.DataParallel(m)\n    return m\n\n\ndef unwrap(m):\n    return m.module if isinstance(m, nn.DataParallel) else m\n\n\ndef make_opt(m, total_steps):\n    opt = torch.optim.AdamW(m.parameters(), lr=CFG[\"lr\"], weight_decay=CFG[\"weight_decay\"])\n    warm = max(1, int(CFG[\"warmup_frac\"] * total_steps))\n\n    def lr_at(step):\n        if step < warm:\n            return (step + 1) / warm\n        p = (step - warm) / max(1, total_steps - warm)\n        return 0.5 * (1 + math.cos(math.pi * min(p, 1.0)))\n\n    return opt, torch.optim.lr_scheduler.LambdaLR(opt, lr_at)\n\n\ndef train_step(m, bt, opt, sch, scaler):\n    with torch.autocast(\"cuda\", dtype=torch.float16, enabled=CFG[\"amp\"]):\n        logits, _ = m(bt[\"x\"].to(device, non_blocking=True))\n    loss = masked_bce(logits.float(), bt[\"y\"].to(device), bt[\"mask\"].to(device),\n                      bt[\"w\"].to(device))\n    opt.zero_grad(set_to_none=True)\n    scaler.scale(loss).backward()\n    scaler.unscale_(opt)\n    nn.utils.clip_grad_norm_(m.parameters(), CFG[\"grad_clip\"])\n    scaler.step(opt)\n    scaler.update()\n    sch.step()\n    return loss.item()\n\n\nprobe = build()\nprint(f\"{CFG['backbone']}: {sum(p.numel() for p in probe.parameters())/1e6:.1f}M params, \"\n      f\"{unwrap(probe).encoder.num_features} features/slice, \"\n      f\"per_column_attn={CFG['per_column_attn']}\")\ndel probe\ntorch.cuda.empty_cache()","metadata":{"_uuid":"6ed7fefc-ee8c-4884-9a3c-5996cb244e53","_cell_guid":"f7937441-332b-4c69-85e0-913f2130e86f","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ----------------------------------------------------------------------------------\n# Smoke test: can this pipeline memorise 32 studies? The floor is the entropy of their own soft\n# labels -- a perfect memoriser reaches it and no lower. Getting nowhere near it means the data,\n# the loss or the pooling is broken, and no amount of GPU quota will fix that.\n# ----------------------------------------------------------------------------------\nif CFG[\"smoke_test\"]:\n    seed_all(CFG[\"seed\"])\n    tiny = train.head(32).reset_index(drop=True)\n    t_bin, t_mask, t_soft = targets(tiny)\n    p = np.clip(t_soft[t_mask > 0], 1e-6, 1 - 1e-6)\n    floor = float(-(p * np.log(p) + (1 - p) * np.log(1 - p)).mean())\n\n    tiny_loader = make_loader(tiny, t_soft, t_mask, train_mode=True, workers=2)\n    sm = build()\n    n_ep = 30\n    opt, sch = make_opt(sm, n_ep * len(tiny_loader))\n    scaler = make_scaler(CFG[\"amp\"])\n    sm.train()\n    losses, t0 = [], time.time()\n    for ep in range(n_ep):\n        ep_loss = [train_step(sm, bt, opt, sch, scaler) for bt in tiny_loader]\n        losses.append(float(np.mean(ep_loss)))\n        if ep % 6 == 0 or ep == n_ep - 1:\n            print(f\"  ep {ep:>2} loss {losses[-1]:.4f}  (entropy floor {floor:.4f})\")\n    got = min(losses[-3:])\n    print(f\"\\nsmoke: {losses[0]:.4f} -> {got:.4f}, floor {floor:.4f}, \"\n          f\"{time.time()-t0:.0f}s for {n_ep} passes over 32 studies\")\n    assert got < floor + 0.5 * (losses[0] - floor), (\n        \"the model cannot get halfway to memorising 32 studies -- fix the pipeline, not the \"\n        \"hyperparameters. Suspect: the loss mask, the bag pooling, or the cache itself.\")\n    # Peak VRAM, measured on the real batch shape rather than reasoned about. batch_size x\n    # SLICES images go through the encoder in one forward pass -- 8 x 32 = 256 at 288 px --\n    # so this is the number that decides whether the full run OOMs eight epochs from now,\n    # and it costs nothing to read here instead of finding out then.\n    peak = max(torch.cuda.max_memory_allocated(i) for i in range(torch.cuda.device_count()))\n    total = min(torch.cuda.get_device_properties(i).total_memory\n                for i in range(torch.cuda.device_count()))\n    print(f\"peak VRAM {peak/1e9:.1f} GB of {total/1e9:.1f} GB per GPU \"\n          f\"({CFG['batch_size']}x{cache_shape[0]} = \"\n          f\"{CFG['batch_size']*cache_shape[0]} images at {CFG['size']} px) | {ram_str()}\")\n    if peak > 0.85 * total:\n        print(\"!! within 15% of the card. Halve CFG['batch_size'] now rather than losing the \"\n              \"run\\n   to an OOM mid-epoch; the schedule is step-based, so nothing else changes.\")\n    elif peak < 0.45 * total:\n        print(f\"-> less than half the card is in use. CFG['batch_size'] can go to about \"\n              f\"{int(CFG['batch_size'] * 0.8 * total / max(peak, 1)) // 2 * 2}, which cuts the \"\n              f\"DataParallel scatter overhead per\\n   study. It only helps if the GPU is the \"\n              f\"bottleneck -- compare studies/s below against the\\n   cache read rate printed \"\n              f\"further up before spending the change.\")\n    for i in range(torch.cuda.device_count()):\n        torch.cuda.reset_peak_memory_stats(i)\n\n    del sm, opt, sch, scaler, tiny_loader\n    torch.cuda.empty_cache()\n    print(\"OK: gradients flow and the bag pooling can fit signal.\")","metadata":{"_uuid":"d3cdbf48-eea1-47f6-81fa-35443f574118","_cell_guid":"70151dbf-bd43-4570-a6b9-031099aa4789","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ----------------------------------------------------------------------------------\n# Train. Selection on val_noisy mean AUC; gold printed every epoch but never selected on.\n# ----------------------------------------------------------------------------------\n@torch.no_grad()\ndef predict(m, loader, want_attn=False):\n    m.eval()\n    P, A, U = [], [], []\n    for bt in loader:\n        with torch.autocast(\"cuda\", dtype=torch.float16, enabled=CFG[\"amp\"]):\n            logits, a = m(bt[\"x\"].to(device, non_blocking=True))\n        P.append(torch.sigmoid(logits.float()).cpu().numpy())\n        if want_attn:\n            A.append(a.float().cpu().numpy())\n        U += list(bt[\"uid\"])\n    return np.concatenate(P), (np.concatenate(A) if want_attn else None), U\n\n\nseed_all(CFG[\"seed\"])\nmodel = build()\nopt, sch = make_opt(model, CFG[\"epochs\"] * len(train_loader))\nscaler = make_scaler(CFG[\"amp\"])\n\nos.makedirs(CFG[\"out_dir\"], exist_ok=True)\nBEST_PT = os.path.join(CFG[\"out_dir\"], \"m2_best_state.pt\")   # weights only, rewritten on\nLAST_PT = os.path.join(CFG[\"out_dir\"], \"m2_last.pt\")         # improvement; full resume state\n\n# The best weights live on disk rather than in a dict held in this process. It costs one 85 MB\n# write per improvement and buys two things: the host RAM is not carrying a second copy of the\n# model for the whole run, and an epoch that finished is an epoch that survives -- the first\n# attempt at this notebook died at epoch 4 of 8 and left nothing at all behind.\nhistory, best, start_ep = [], (-1.0, -1), 0\nif CFG[\"resume\"] and os.path.isfile(LAST_PT):\n    st = torch.load(LAST_PT, map_location=\"cpu\", weights_only=False)\n    unwrap(model).load_state_dict(st[\"model\"])\n    opt.load_state_dict(st[\"opt\"])\n    sch.load_state_dict(st[\"sch\"])\n    scaler.load_state_dict(st[\"scaler\"])\n    history, best, start_ep = st[\"history\"], tuple(st[\"best\"]), st[\"epoch\"] + 1\n    print(f\"resumed from {LAST_PT}: epoch {st['epoch']} done, best val_noisy {best[0]:.4f}\")\n\ndeadline = time.time() + CFG[\"time_budget_h\"] * 3600 if CFG[\"time_budget_h\"] else None\nt_start = time.time()\nstopped_early = False\nfor ep in range(start_ep, CFG[\"epochs\"]):\n    model.train()\n    t0, run, seen = time.time(), 0.0, 0\n    for i, bt in enumerate(train_loader, 1):\n        run += train_step(model, bt, opt, sch, scaler)\n        seen += 1\n        if i % 50 == 0:\n            used, tot = ram()\n            print(f\"  ep{ep} [{i}/{len(train_loader)}] loss {run/seen:.4f} \"\n                  f\"| {i*CFG['batch_size']/(time.time()-t0):.1f} studies/s \"\n                  f\"| lr {sch.get_last_lr()[0]:.2e} | RAM {used:.1f}/{tot:.1f} GB\")\n            # A bare \"Kernel died\" with no traceback is the OOM killer, and by then the log\n            # says nothing about why. Say it while there is still a process to say it.\n            if used > 0.90 * tot:\n                print(\"  !! host RAM over 90%. The kernel will be SIGKILLed, not raised in. \"\n                      \"Cut CFG['prefetch']\\n     or CFG['num_workers'] and restart -- this \"\n                      \"run is not going to finish.\")\n    t_train = time.time() - t0\n\n    pn, _, un = predict(model, vn_loader)\n    pg, _, ug = predict(model, vg_loader)\n    assert un == val_noisy[\"StudyInstanceUID\"].tolist()      # order must match the targets\n    assert ug == val_gold[\"StudyInstanceUID\"].tolist()\n    mean_n = per_column_auc(pn, yvn_bin, mvn)[1]\n    mean_g = per_column_auc(pg, yvg_bin, mvg)[1]\n    history.append(dict(epoch=ep, loss=run/seen, val_noisy=mean_n, gold=mean_g,\n                        train_min=t_train/60, epoch_min=(time.time()-t0)/60))\n    print(f\"epoch {ep}: loss {run/seen:.4f} | val_noisy {mean_n:.4f} | GOLD {mean_g:.4f} \"\n          f\"| {t_train/60:.1f} min train, {(time.time()-t0)/60:.1f} min with validation \"\n          f\"| {ram_str()}\")\n    if mean_n > best[0]:\n        best = (mean_n, ep)\n        torch.save(unwrap(model).state_dict(), BEST_PT)\n        print(\"  ^ best so far (selected on val_noisy, never on gold)\")\n\n    torch.save(dict(model=unwrap(model).state_dict(), opt=opt.state_dict(),\n                    sch=sch.state_dict(), scaler=scaler.state_dict(),\n                    epoch=ep, history=history, best=list(best)), LAST_PT)\n    pd.DataFrame(history).to_csv(os.path.join(CFG[\"out_dir\"], \"m2_history.csv\"), index=False)\n\n    if deadline and time.time() + (time.time() - t0) > deadline:\n        stopped_early = True\n        print(f\"\\n!! stopping after epoch {ep}: another epoch would pass the \"\n              f\"{CFG['time_budget_h']} h budget.\\n   Everything below runs on the best \"\n              f\"checkpoint so far, which is a real result, not a lost run.\")\n        break\n\nn_done = len(history)                       # epochs ever run, resumed ones included\nn_here = max(n_done - start_ep, 1)           # epochs this session, which is what el covers\nel = (time.time() - t_start) / 3600\nprint(f\"\\n=== {n_done - start_ep} epochs this session ({n_done} total) in {el:.2f} h ===\")\nprint(f\"GPU-HOURS: {el:.2f} h wall on {torch.cuda.device_count()} GPU(s) for {len(train)} \"\n      f\"studies x {n_here} epochs -> {el/n_here*60:.1f} min/epoch. \"\n      f\"The budget is 30 GPU-h/week; plan folds and seeds against this number, not a guess.\")\nprint(f\"best val_noisy {best[0]:.4f} at epoch {best[1]}\"\n      + (\"  (run stopped at the time budget)\" if stopped_early else \"\"))\npd.DataFrame(history)","metadata":{"_uuid":"02e6d0c2-ad43-4532-8323-46e141524f56","_cell_guid":"bc84f867-78b7-44c4-961c-7061897ffef1","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ----------------------------------------------------------------------------------\n# Final read-out, checkpoint, artefacts\n# ----------------------------------------------------------------------------------\nassert os.path.isfile(BEST_PT), (\n    f\"{BEST_PT} is gone -- it is written the first time val_noisy improves and read here. If \"\n    f\"this cell is being re-run in a session that resumed from m2_last.pt without ever \"\n    f\"improving, set CFG['resume']=False and train again rather than shipping the last epoch \"\n    f\"as if it were the selected one.\")\nbest_state = torch.load(BEST_PT, map_location=\"cpu\", weights_only=True)\nunwrap(model).load_state_dict(best_state)\npn, an, un = predict(model, vn_loader, want_attn=True)\npg, ag, ug = predict(model, vg_loader, want_attn=True)\n\nprint(f\"PER-COLUMN AUC -- epoch {best[1]}, {CFG['backbone']}, {'/'.join(CFG['slot_priority'])}\\n\")\nprint(f\"val_noisy ({len(val_noisy)} studies, labels from the extractor, so this measures \"\n      f\"agreement\\nwith the extractor -- it is the selection set, not the truth)\\n\")\nprint(auc_table({\"val_noisy\": pn}, yvn_bin, mvn))\nprint(f\"\\ngold (n={len(val_gold)}, ground truth, never trained or selected on; +/-0.06 per column)\\n\")\nprint(auc_table({\"gold\": pg}, yvg_bin, mvg))\n\ntorch.save(dict(\n    state_dict=best_state, epoch=best[1], val_noisy_mean=best[0], labels=LABELS,\n    cfg={k: CFG[k] for k in (\"backbone\", \"size\", \"slot_priority\", \"attn_hidden\", \"dropout\",\n                             \"per_column_attn\", \"seed\", \"val_bucket\", \"group_by_site\",\n                             \"mask_unresolved_lateral\")},\n    # How this fold was actually split, not how CFG asked for it -- grouping silently falls\n    # back to i.i.d. when the manifest has no site keys, and a later fold or blend has to know\n    # which of the two produced this val_noisy number before comparing against it.\n    split=dict(grouped=bool(use_groups), n_sites=int(n_sites),\n               val_sites=sorted({site_of.get(u, \"\") for u in val_noisy[\"StudyInstanceUID\"]}\n                                - {\"\"}),\n               n_train=len(train), n_val_noisy=len(val_noisy), n_val_gold=len(val_gold)),\n    # the cache geometry this model was trained on; the submission notebook asserts against it\n    cache=dict(slices=cache_shape[0], size=cache_shape[-1]),\n    # constant fallback for test studies with no usable series: the metric is rank-based, so a\n    # column-mean constant is the least-harmful thing to emit for a study you cannot look at\n    fallback=pn.mean(0).astype(np.float32),\n), os.path.join(CFG[\"out_dir\"], \"m2_fold0.pt\"))\n\nfor name, p, u in ((\"val_noisy\", pn, un), (\"gold\", pg, ug)):\n    pd.DataFrame(p, columns=LABELS).assign(StudyInstanceUID=u).to_csv(\n        os.path.join(CFG[\"out_dir\"], f\"m2_{name}_preds.csv\"), index=False)\npd.DataFrame(history).to_csv(os.path.join(CFG[\"out_dir\"], \"m2_history.csv\"), index=False)\n\ncol_n, mean_n, npos_n = per_column_auc(pn, yvn_bin, mvn)\ncol_g, mean_g, npos_g = per_column_auc(pg, yvg_bin, mvg)\npd.DataFrame({\"label\": LABELS,\n              \"auc_val_noisy\": [col_n[k] for k in LABELS],\n              \"n_pos_val_noisy\": [npos_n[k] for k in LABELS],\n              \"auc_gold\": [col_g[k] for k in LABELS],\n              \"n_pos_gold\": [npos_g[k] for k in LABELS]}).to_csv(\n    os.path.join(CFG[\"out_dir\"], \"m2_per_column_auc.csv\"), index=False)\nprint(f\"\\nsaved: m2_fold0.pt, m2_per_column_auc.csv, m2_history.csv, preds -> {CFG['out_dir']}\")\nprint(f\"also m2_last.pt ({os.path.getsize(LAST_PT)/1e6:.0f} MB, optimiser + schedule state) and \"\n      f\"m2_best_state.pt\\n({os.path.getsize(BEST_PT)/1e6:.0f} MB, the selected weights alone). \"\n      f\"Re-running this notebook in the same session\\nresumes from m2_last.pt; delete both to \"\n      f\"start clean. m2_fold0.pt is the only one the\\nsubmission notebook needs -- the other two \"\n      f\"are ~{(os.path.getsize(LAST_PT)+os.path.getsize(BEST_PT))/1e6:.0f} MB of dead weight in \"\n      f\"the dataset you push.\")\n\n# attention over slice index: flat means the pooling learned nothing and this is mean pooling\nprof = an.mean(axis=(0, 2))\nfig, ax = plt.subplots(1, 2, figsize=(11, 3.2))\nax[0].plot(prof)\nax[0].axhline(1 / len(prof), ls=\"--\", c=\"gray\")\nax[0].set_title(\"mean attention vs slice index (dashed = uniform)\")\nax[0].set_xlabel(\"slice (lateral -> medial)\")\nh = pd.DataFrame(history)\nax[1].plot(h[\"epoch\"], h[\"val_noisy\"], marker=\"o\", label=\"val_noisy\")\nax[1].plot(h[\"epoch\"], h[\"gold\"], marker=\"s\", label=\"gold\")\nax[1].axhline(0.5, ls=\"--\", c=\"gray\")\nax[1].set_title(\"mean AUC\")\nax[1].legend()\nplt.tight_layout()\nplt.show()","metadata":{"_uuid":"87e0940e-f4a0-414b-835c-f02af25a404d","_cell_guid":"32453002-e323-4f6a-9f4d-7ca6e14163fa","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## What to do with these numbers\n#\n1. **Save this notebook's output as a Kaggle dataset** -- `m2_fold0.pt` is what\n   `knee_m2_submit.ipynb` loads -- then run the submission notebook and actually submit. Until\n   something has scored, the LB-vs-CV gap is unknown and every next decision is a guess.\n2. **Read `m2_per_column_auc.csv` column by column, never just the mean.** Each column is 1/12\n   of the metric. The expected shape from ANALYSIS.md is that Medial/Lateral OA and Synovitis\n   are hardest and Effusion is the noisiest label; if some *other* column sits at 0.5, suspect\n   the pipeline rather than the model.\n3. **Compare gold against val_noisy per column.** Much lower on gold means the model learned\n   the extractor's mistakes instead of the anatomy -- that is a labels problem (upgrade the\n   4,165 rules rows), not an architecture problem. The reverse, at n=58, is usually just noise.\n4. **Check which split actually ran** -- the cell prints it and the checkpoint records it. A\n   grouped split holds whole sites out, so `val_noisy` is a domain-shift estimate and the\n   OOF->LB gap should be near the +0.05 CLIMB.md expects. If it fell back to i.i.d. because\n   the manifest was missing, `val_noisy` is measuring memorised scanners as well as anatomy\n   and will read optimistic -- do not compare the two kinds of number to each other.\n#\nPer PLAN.md, what comes next follows from that table rather than from a plan written before it.\nCandidates already identified: coronal/axial slots (should help MCL and the OA compartments,\nwhich sagittal sees poorly -- the cache already holds them), better labels for the 4,165 rules\nrows, `per_column_attn=True`, then more folds and seeds with rank-averaging.","metadata":{"_uuid":"48a42acb-246e-42b5-a1c2-0fcdf6c78fbb","_cell_guid":"1e6e9f09-7920-4ed5-9f7a-7a517464de8c","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}}]}