{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\n<style>\n    @import url('https://fonts.googleapis.com/css2?family=Orbitron:wght@400;700&family=Inter:wght@300;400;600&display=swap');\n    \n    .astral-header {\n        background: linear-gradient(135deg, #090B10 0%, #170E30 50%, #20134E 100%);\n        padding: 30px;\n        border-radius: 12px;\n        color: #E2E8F0;\n        font-family: 'Inter', sans-serif;\n        box-shadow: 0 4px 20px rgba(138, 43, 226, 0.2);\n        border: 1px solid #3B1C73;\n        margin-bottom: 20px;\n    }\n    .astral-title {\n        font-family: 'Orbitron', sans-serif;\n        font-size: 36px;\n        margin: 0 0 10px 0;\n        background: linear-gradient(90deg, #9F7AEA, #ED64A6);\n        -webkit-background-clip: text;\n        -webkit-text-fill-color: transparent;\n        text-shadow: 0 0 20px rgba(159, 122, 234, 0.3);\n    }\n    .astral-subtitle {\n        font-size: 16px;\n        font-weight: 300;\n        opacity: 0.9;\n        margin: 0;\n        line-height: 1.5;\n    }\n    .astral-note {\n        border-left: 4px solid #ED64A6;\n        padding-left: 15px;\n        margin-top: 20px;\n        font-style: italic;\n        color: #CBD5E1;\n    }\n</style>\n<div class=\"astral-header\">\n    <h1 class=\"astral-title\">âœ§ The Astral CPU Baseline âœ§</h1>\n    <h3 class=\"astral-subtitle\">Deep learning at the speed of starlight.<br>A highly optimized, CPU-friendly solution for the RSNA 2026 Knee MRI Challenge.</h3>\n    <div class=\"astral-note\">\"To see into the body is to peer into a universe. Let's illuminate it efficiently.\" â€” The Researcher</div>\n</div>\n\nWe stand at the frontier of musculoskeletal imaging. Our task is to decipher the complex, 3D structures of the knee and identify twelve distinct pathologies. However, computational power is finite. The standard approaches rely heavily on monolithic GPU clusters. In this notebook, we take a different path. We will build a **CPU-optimized, blazing-fast pipeline** that sacrifices almost nothing in accuracy while running in a fraction of the time. \n\n### Core Innovations in this Notebook:\n1. **The MobileNetV3 Backbone**: We replace the heavy EfficientNetV2 with `mobilenetv3_large_100`, explicitly designed for low-latency CPU inference.\n2. **Slice & Resolution Reduction**: We sample 12 critical slices at 192px instead of 16 at 224px â€” a ~45% compute reduction with minimal information loss.\n3. **The Polyglot Teacher**: The unlabelled data is vastly larger than our labelled subset. We use a deeply expanded multilingual NLP lexicon to generate high-confidence pseudo-labels from the radiology reports.\n4. **Dynamic Quantization**: At inference time, our models run in `int8` precision, slicing through the test set seamlessly.\n","metadata":{}},{"cell_type":"markdown","source":"## What this notebook changes versus the public EfficientNet 2.5D baseline\n\nThat notebook is a reasonable skeleton and its backbone choice is kept here. These are the issues that\nstop it from training well â€” or from running at all.\n\n| # | Issue in the public baseline | Consequence | Fix here |\n|---|---|---|---|\n| 1 | `timm.create_model('efficientnetv2_s')` | **Crashes** â€” that name is not registered in timm (the real ones are `tf_efficientnetv2_s`, `efficientnetv2_rw_s`) | `resolve_backbone()` picks the first name that actually exists |\n| 2 | `pretrained=False` | Random init on a few thousand studies â€” the single largest score loss | `pretrained=True` when training; weights loaded from a dataset at submit time |\n| 3 | `T.RandomHorizontalFlip(p=0.5)` | Destroys medial/lateral, i.e. 4 of the 12 targets | no flips; laterality normalised instead |\n| 4 | `self.augment` is built but never called in `__getitem__` | No augmentation happens at all | augmentation applied explicitly to the volume |\n| 5 | `sorted(folder.glob('*.dcm'))` | Sorts by SOP UID â€” an essentially random slice order | sorted by `ImagePositionPatient` projected on the slice normal |\n| 6 | Only middle slices are read | Baker's cyst and MCL live at the periphery of the stack | slices sampled uniformly across the volume |\n| 7 | Trains on labelled studies only | Discards most of the dataset | report teacher labels the rest |\n| 8 | `StratifiedGroupKFold(..., groups=StudyInstanceUID)` on one label | Groups are unique, so grouping is a no-op; rare targets end up unbalanced | iterative multilabel stratification + grouping by duplicate report |\n| 9 | `pos_weight = neg/pos` | Up to ~100Ã— on rare labels â€” unstable gradients, and AUC ignores the prior anyway | plain BCE on soft targets |\n| 10 | Every series of every study decoded per item | Dataloader-bound; one epoch takes hours | pre-cached `uint8` volumes |\n| 11 | `MONOCHROME1`, `RescaleSlope`, anisotropic spacing ignored | Some studies come out inverted or stretched | handled inside the loader |\n| 12 | Validation AUC over whatever labels are present | Silently optimistic | validation restricted to hard labels |","metadata":{}},{"cell_type":"markdown","source":"## 0. Configuration and theme\n\nOne switch block drives the whole notebook. On Kaggle it is convenient to split this into three\nnotebooks (`preprocess` â†’ `train` â†’ `submit`), but the code is identical.","metadata":{}},{"cell_type":"code","source":"import os, re, sys, gc, json, math, time, warnings, hashlib\nfrom pathlib import Path\nfrom collections import Counter\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nwarnings.filterwarnings('ignore')\npd.set_option('display.width', 200)\npd.set_option('display.max_columns', 50)\n\n\nclass CFG:\n    # --- paths ---\n    comp_dir    = Path('/kaggle/input/competitions/rsna-knee-abnormality-detection')\n    work_dir    = Path('/kaggle/working')\n    cache_dir   = Path('/kaggle/temp/cache')     # preprocessing output\n    extra_cache = []                             # finished shards attached as datasets\n\n    # --- switches ---\n    RUN_EDA        = True\n    RUN_TEACHER    = True     # report -> labels\n    RUN_PREPROCESS = False    # DICOM -> npz, slow, run in shards\n    RUN_TRAIN      = False\n    RUN_INFER      = True\n\n    # --- preprocessing shards ---\n    SHARD_ID   = 0\n    NUM_SHARDS = 8\n\n    # --- cache geometry ---\n    IMG_SIZE = 192\n    N_SLICES = 12\n    VIEWS    = ['SAG_FS', 'COR_FS', 'AX_FS']\n    N_VIEWS  = 3\n\n    # --- model ---\n    backbone   = ['mobilenetv3_large_100', 'tf_efficientnetv2_b0']\n    pretrained = True          # needs internet -> training notebook only\n    drop_rate  = 0.3\n    grad_ckpt  = True          # ~30% slower, much lower memory\n\n    # --- training ---\n    folds       = 5\n    train_folds = [0]\n    epochs      = 6\n    batch_size  = 2            # studies, i.e. 2*3*16 = 96 images per forward\n    accum       = 4\n    lr          = 3e-4         # head\n    backbone_lr = 1e-4\n    wd          = 1e-2\n    pseudo_w    = 0.5          # weight of report-derived soft labels vs hard labels\n    aux_w       = 0.3          # per-view auxiliary head\n    seed        = 42\n    num_workers = 4\n    amp         = True\n\nTARGETS = ['ACL', 'MCL', 'Medial Meniscus', 'Lateral Meniscus',\n           'Medial OA', 'Lateral OA', 'PF OA', 'Effusion',\n           'Synovitis', \"Baker's\", 'Contusion', 'Fracture']\n\nCFG.cache_dir.mkdir(parents=True, exist_ok=True)\nCFG.work_dir.mkdir(parents=True, exist_ok=True)\n\n# ------------------------------------------------------------------ icefire theme\nICEFIRE  = sns.color_palette('icefire', 12)    # categorical ramp, cold -> hot\nICE      = ICEFIRE[2]                          # cold accent\nFIRE     = ICEFIRE[9]                          # warm accent\nMIDLINE  = '#39414d'\nCMAP     = 'icefire'                                            # diverging: lift, centred maps\nCMAP_SEQ = sns.blend_palette(['#eef1f6', ICE, FIRE], as_cmap=True)   # sequential, same family\n\nsns.set_theme(style='dark', rc={'axes.facecolor': '#090B10', 'figure.facecolor': '#090B10', 'text.color': '#E2E8F0', 'axes.edgecolor': '#3B1C73'})\nmpl.rcParams.update({\n    'figure.dpi': 120, 'savefig.dpi': 120,\n    'axes.spines.top': False, 'axes.spines.right': False,\n    'axes.edgecolor': '#c8cdd4', 'axes.labelcolor': MIDLINE,\n    'axes.titlesize': 12, 'axes.titleweight': 'bold', 'axes.titlepad': 10,\n    'text.color': MIDLINE, 'xtick.color': MIDLINE, 'ytick.color': MIDLINE,\n    'grid.color': '#e8eaee', 'axes.grid': True, 'axes.grid.axis': 'y',\n    'legend.frameon': False, 'figure.facecolor': 'white',\n})\n\ndef grad(n, reverse=False):\n    \"\"\"n colours along the icefire ramp, skipping its near-black centre.\n\n    icefire is diverging: sampling it directly gives a muddy black bar in the middle of any\n    categorical chart. Taking colours from the two lobes keeps the cold->hot reading without it.\n    \"\"\"\n    cm = plt.get_cmap(CMAP)\n    cold = np.linspace(0.02, 0.28, n - n // 2)      # icefire is near-black between ~0.35 and ~0.75\n    warm = np.linspace(0.78, 0.99, n // 2)\n    cols = [cm(p) for p in np.concatenate([cold, warm])][:n]\n    return cols[::-1] if reverse else cols\n\n\ndef seed_everything(seed=42):\n    import random\n    random.seed(seed); np.random.seed(seed); os.environ['PYTHONHASHSEED'] = str(seed)\n    try:\n        import torch\n        torch.manual_seed(seed); torch.cuda.manual_seed_all(seed)\n    except ImportError:\n        pass\n\nseed_everything(CFG.seed)\n\nif not CFG.comp_dir.exists():                       # be forgiving about the mount point\n    for cand in list(Path('/kaggle/input').glob('*knee*')) + \\\n                list(Path('/kaggle/input/competitions').glob('*knee*')):\n        CFG.comp_dir = cand\n        break\nprint('comp_dir :', CFG.comp_dir, '| exists:', CFG.comp_dir.exists())","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Environment check\n\nSeries arrive in a mix of transfer syntaxes, including JPEG 2000 and JPEG Lossless. `pydicom` cannot\ndecode those on its own â€” it needs `pylibjpeg` (+ `pylibjpeg-openjpeg`, `pylibjpeg-libjpeg`) or\n`python-gdcm`. **The submission notebook has no internet**, so those wheels must be attached as a\ndataset. Better to learn that here than through a wall of exceptions three hours in.","metadata":{}},{"cell_type":"code","source":"import pydicom\nprint('pydicom :', pydicom.__version__)\n\ndecoders = {}\nfor mod in ['pylibjpeg', 'openjpeg', 'libjpeg', 'gdcm', 'PIL']:\n    try:\n        __import__(mod)\n        decoders[mod] = 'ok'\n    except Exception as e:\n        decoders[mod] = f'MISSING ({type(e).__name__})'\nfor k, v in decoders.items():\n    print(f'  {k:<10}: {v}')\n\nif 'MISSING' in decoders['openjpeg'] and 'MISSING' in decoders['gdcm']:\n    print('\\n[!] No JPEG-2000 decoder found â€” compressed series will fail to decode.')\n    print('    pip install pylibjpeg pylibjpeg-openjpeg pylibjpeg-libjpeg python-gdcm')\n    print('    and attach the wheels as a dataset for the offline submission notebook.')\n\ntry:\n    import torch\n    gpu = torch.cuda.get_device_name(0) if torch.cuda.is_available() else 'cpu only'\n    print(f'\\ntorch   : {torch.__version__} | cuda: {torch.cuda.is_available()} | {gpu}')\nexcept ImportError:\n    print('\\ntorch not available (EDA-only environment)')","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train        = pd.read_csv(CFG.comp_dir / 'train.csv')\ntrain_series = pd.read_csv(CFG.comp_dir / 'train_series.csv')\ntest         = pd.read_csv(CFG.comp_dir / 'test.csv')\ntest_series  = pd.read_csv(CFG.comp_dir / 'test_series.csv')\nsample_sub   = pd.read_csv(CFG.comp_dir / 'sample_submission.csv')\n\nprint(f'train studies : {len(train):,}')\nprint(f'train series  : {len(train_series):,}  ({len(train_series)/len(train):.2f} per study)')\nprint(f'test studies  : {len(test):,}  (about 1300 in the real test set)')\nprint('\\ntrain columns :', train.columns.tolist())\nprint('test  columns :', test.columns.tolist())\nassert 'Report' not in test.columns, \\\n    'If reports ever appear in the test set the whole strategy changes â€” check this.'\ntrain.head(3)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Part 1 â€” EDA\n\nNot decoration: every plot below answers a question that changes the code further down.\n\n## 1.1 How much of the data is actually labelled","metadata":{}},{"cell_type":"code","source":"has_label  = train[TARGETS].notna().all(axis=1)\nhas_report = train['Report'].fillna('').str.strip().str.len() > 0\n\nprint(f'hard labels        : {has_label.sum():,} / {len(train):,}  ({has_label.mean()*100:.1f}%)')\nprint(f'report present     : {has_report.sum():,} / {len(train):,}  ({has_report.mean()*100:.1f}%)')\nprint(f'neither            : {(~has_label & ~has_report).sum():,}  <- unusable for supervision')\nprint(f'partially labelled : {(train[TARGETS].notna().any(axis=1) & ~has_label).sum():,}')\n\nlab  = train[has_label].copy()\nprev = lab[TARGETS].mean().sort_values(ascending=False)\nsummary = pd.DataFrame({'pos_rate': prev,\n                        'n_pos': lab[TARGETS].sum().reindex(prev.index).astype(int)})\nsummary['n_pos_per_val_fold'] = (summary['n_pos'] / CFG.folds).round().astype(int)\nprint('\\nPrevalence on the labelled subset:')\ndisplay(summary.style.background_gradient(cmap=CMAP_SEQ, subset=['pos_rate'])\n                     .format({'pos_rate': '{:.3f}'}))\nprint('Watch n_pos_per_val_fold: below ~30 the single-fold AUC of that target swings by +-0.05.')","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 2, figsize=(15, 4.6))\n\nax = axes[0]\nax.barh(prev.index[::-1], prev.values[::-1] * 100, color=grad(12), edgecolor='white', height=.75)\nax.set_xlabel('positive rate, %'); ax.set_title('Prevalence of each target (hard labels)')\nax.grid(axis='x'); ax.grid(axis='y', visible=False)\nfor i, v in enumerate(prev.values[::-1]):\n    ax.text(v * 100 + .35, i, f'{v*100:.1f}', va='center', fontsize=8, color=MIDLINE)\n\nax = axes[1]\nnpl = lab[TARGETS].sum(axis=1)\nvals, cnts = np.unique(npl, return_counts=True)\nax.bar(vals, cnts, color=grad(len(vals)), edgecolor='white')\nax.set_xlabel('abnormalities per study'); ax.set_ylabel('studies')\nax.set_title(f'Multi-label density (mean = {npl.mean():.2f})')\nplt.tight_layout(); plt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 1.2 Co-occurrence â€” use lift, not conditional probability\n\n`P(col | row)` mostly reflects the prevalence of the column, so it looks dramatic without saying much.\nLift, `P(A,B) / P(A)P(B)`, isolates the actual association. Pairs well above 1 (typically\n`Medial OA` â†” `Lateral OA`, `ACL` â†” `Contusion`) will be learnt jointly by a shared trunk; isolated\ntargets such as `Fracture` need their own attention.","metadata":{}},{"cell_type":"code","source":"P = lab[TARGETS].mean().values\nJ = (lab[TARGETS].T.values @ lab[TARGETS].values) / len(lab)\nlift = pd.DataFrame(J / (P[:, None] * P[None, :] + 1e-9), index=TARGETS, columns=TARGETS)\n\nfig, ax = plt.subplots(figsize=(10, 8))\nsns.heatmap(lift, annot=True, fmt='.1f', cmap=CMAP, center=1.0, vmin=0, vmax=3,\n            mask=np.eye(12, dtype=bool), linewidths=.6, linecolor='white', ax=ax,\n            annot_kws={'size': 8}, cbar_kws={'label': 'lift = P(A,B) / P(A)P(B)', 'shrink': .8})\nax.set_title('Co-occurrence lift â€” above 1 means the pair appears together more often than chance')\nplt.tight_layout(); plt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 1.3 Series inventory â€” which views a study actually has\n\nThe question that decides the architecture: **for how many studies does each\n(plane Ã— fluid-sensitive Ã— fat-sat) combination exist?** If axial is present for only 60% of studies,\nthe model must handle a missing view through masking rather than crash or read zeros as signal.","metadata":{}},{"cell_type":"code","source":"def view_key(r):\n    plane = str(r['Anatomical_Plane'])[:3].upper()\n    return f\"{plane}_{'FS' if r['Fluid_Sensitive'] == 1 else 'T1'}{'_FAT' if r['Fat_Suppression'] == 1 else ''}\"\n\nts = train_series.copy()\nts['view'] = ts.apply(view_key, axis=1)\n\ncombo = (ts.groupby(['Anatomical_Plane', 'Fluid_Sensitive', 'Fat_Suppression'])\n           .size().reset_index(name='n_series').sort_values('n_series', ascending=False))\ncombo['pct'] = (combo.n_series / len(ts) * 100).round(1)\nprint('Series type combinations:'); display(combo)\n\ncov = ts.assign(one=1).pivot_table(index='StudyInstanceUID', columns='view',\n                                   values='one', aggfunc='sum', fill_value=0)\ncoverage = pd.DataFrame({\n    'studies_with_view': (cov > 0).sum(),\n    'pct_studies': ((cov > 0).mean() * 100).round(1),\n    'mean_series_when_present': cov.replace(0, np.nan).mean().round(2),\n}).sort_values('pct_studies', ascending=False)\n\nfig, axes = plt.subplots(1, 2, figsize=(15, 4.4))\nax = axes[0]\nax.barh(coverage.index[::-1], coverage.pct_studies.values[::-1],\n        color=grad(len(coverage)), edgecolor='white', height=.7)\nax.axvline(80, color=FIRE, ls='--', lw=1.2, label='80% â€” below this, make the view optional')\nax.set_xlabel('% of studies containing the view'); ax.set_title('View coverage per study')\nax.grid(axis='x'); ax.grid(axis='y', visible=False); ax.legend(fontsize=8)\n\nax = axes[1]\nplane_cov = (ts.pivot_table(index='StudyInstanceUID', columns='Anatomical_Plane',\n                            values='Fluid_Sensitive', aggfunc='size', fill_value=0) > 0)\npc = (plane_cov.mean() * 100).sort_values(ascending=False)\nax.bar(pc.index, pc.values, color=grad(len(pc)), edgecolor='white', width=.55)\nax.set_ylim(0, 108); ax.set_ylabel('% of studies'); ax.set_title('Coverage by anatomical plane')\nfor i, v in enumerate(pc.values):\n    ax.text(i, v + 1.8, f'{v:.1f}', ha='center', fontsize=9, color=MIDLINE)\nplt.tight_layout(); plt.show()\n\ndisplay(coverage)\nprint(f'studies with all three planes           : {plane_cov.all(axis=1).mean()*100:.1f}%')\nprint(f\"studies with any fluid-sensitive series : \"\n      f\"{ts.groupby('StudyInstanceUID')['Fluid_Sensitive'].max().mean()*100:.1f}%\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 1.4 DICOM audit â€” tags, transfer syntax, laterality\n\nThe organisers kept 86 tags. Which ones matters a great deal: can we read the knee side\n(`Laterality` / `ImageLaterality`), can we sort slices geometrically (`ImagePositionPatient` +\n`ImageOrientationPatient`), is `PatientID` present (fold leakage if one patient has several studies)?","metadata":{}},{"cell_type":"code","source":"def sample_dcm_paths(root, n_series=60, per_series=2, seed=0):\n    rng = np.random.default_rng(seed)\n    series_dirs = [p for p in root.glob('*/*') if p.is_dir()]\n    if not series_dirs:\n        return []\n    pick = rng.choice(len(series_dirs), size=min(n_series, len(series_dirs)), replace=False)\n    return [f for i in pick for f in sorted(series_dirs[i].glob('*.dcm'))[:per_series]]\n\n\npaths = sample_dcm_paths(CFG.comp_dir / 'train_series')\nprint('scanning files:', len(paths))\n\ntag_counter, rows = Counter(), []\nfor p in paths:\n    try:\n        ds = pydicom.dcmread(str(p), stop_before_pixels=True)\n    except Exception as e:\n        print('read failed:', p.name, e); continue\n    for el in ds:\n        tag_counter[el.keyword or str(el.tag)] += 1\n    rows.append({\n        'Rows': getattr(ds, 'Rows', None), 'Columns': getattr(ds, 'Columns', None),\n        'PixelSpacing': str(getattr(ds, 'PixelSpacing', None)),\n        'SliceThickness': getattr(ds, 'SliceThickness', None),\n        'Photometric': getattr(ds, 'PhotometricInterpretation', None),\n        'Laterality': getattr(ds, 'Laterality', None) or getattr(ds, 'ImageLaterality', None),\n        'SeriesDescription': getattr(ds, 'SeriesDescription', None),\n        'FieldStrength': getattr(ds, 'MagneticFieldStrength', None),\n        'Manufacturer': getattr(ds, 'Manufacturer', None),\n        'TransferSyntax': str(ds.file_meta.TransferSyntaxUID) if hasattr(ds, 'file_meta') else None,\n        'has_IPP': hasattr(ds, 'ImagePositionPatient'),\n        'has_IOP': hasattr(ds, 'ImageOrientationPatient'),\n        'has_PatientID': hasattr(ds, 'PatientID'),\n    })\n\nmeta = pd.DataFrame(rows)\nprint('\\n=== available tags, top 40 by frequency ===')\nprint(pd.Series(tag_counter).sort_values(ascending=False).head(40).to_string())\nfor c in ['Laterality', 'Photometric', 'TransferSyntax', 'Manufacturer',\n          'has_IPP', 'has_IOP', 'has_PatientID']:\n    print(f'\\n--- {c} ---')\n    print(meta[c].value_counts(dropna=False).head(8).to_string())\nprint('\\n--- geometry ---')\nprint(meta[['Rows', 'Columns', 'SliceThickness']].describe().round(2).to_string())","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 3, figsize=(15, 3.8))\nfor ax, col, title in zip(axes, ['Laterality', 'Photometric', 'Manufacturer'],\n                          ['Knee side', 'Photometric interpretation', 'Scanner vendor']):\n    vc = meta[col].fillna('missing').value_counts().head(6)\n    ax.bar(vc.index.astype(str), vc.values, color=grad(max(len(vc), 3)), edgecolor='white', width=.6)\n    ax.set_title(title); ax.tick_params(axis='x', rotation=20, labelsize=8)\nplt.tight_layout(); plt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Audit checklist**\n\n- `Laterality` / `ImageLaterality` populated â†’ the knee side can be normalised directly. If empty,\n  fall back to the sign of X in `ImagePositionPatient` (in the LPS frame a right knee sits at x < 0).\n- `MONOCHROME1` present â†’ those slices must be inverted.\n- Several transfer syntaxes â†’ the JPEG-2000 decoders flagged above are mandatory, not optional.\n- `PatientID` present â†’ group the folds by patient.\n\n## 1.5 Leakage check","metadata":{}},{"cell_type":"code","source":"if meta['has_PatientID'].any():\n    print('PatientID exists in the headers -> collect it during preprocessing and group folds by it.')\nelse:\n    print('PatientID is not among the retained tags.')\n\nrep = train['Report'].fillna('').str.strip()\nlong_rep = rep[rep.str.len() > 30]\ndup = long_rep.duplicated(keep=False)\nprint(f'\\nexactly duplicated non-empty reports: {dup.sum():,} rows '\n      f'({long_rep[dup].nunique():,} unique texts)')\nif dup.sum():\n    print('-> those studies must share a fold; handled in Part 4 via a report hash group.')","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 1.6 Reports â€” length, language, keyword signal\n\nHand-written stop-word heuristics make a poor language detector: `la ` occurs in French, Spanish and\nItalian alike, and an `if/elif` chain bakes in an ordering bias. The estimate below is at least\nsymmetric, and it is only used for orientation â€” the teacher in Part 3 relies on **character n-grams,\nwhich are language-agnostic by construction**.","metadata":{}},{"cell_type":"code","source":"rep_len = train['Report'].fillna('').str.split().str.len()\nprint(rep_len.describe().round(1).to_string())\n\nLANG_HINTS = {\n    'en': [' the ', ' and ', ' with ', ' there is ', ' no evidence '],\n    'nl': [' de ', ' het ', ' en ', ' geen ', ' voorste ', ' kruisband '],\n    'de': [' der ', ' die ', ' das ', ' und ', ' kein ', ' kreuzband '],\n    'fr': [' le ', ' la ', ' des ', ' avec ', ' pas de ', ' croise '],\n    'es': [' el ', ' los ', ' con ', ' sin ', ' rodilla ', ' menisco '],\n    'pt': [' do ', ' da ', ' com ', ' sem ', ' joelho ', ' menisco '],\n    'it': [' il ', ' della ', ' con ', ' senza ', ' ginocchio ', ' menisco '],\n    'tr': [' ve ', ' ile ', ' yok', ' diz ', ' menisk'],\n}\n\ndef guess_lang(t):\n    if not isinstance(t, str) or len(t.strip()) < 15:\n        return 'empty'\n    s = ' ' + t.lower() + ' '\n    scores = {k: sum(s.count(w) for w in v) for k, v in LANG_HINTS.items()}\n    best, val = max(scores.items(), key=lambda kv: kv[1])\n    return best if val > 0 else 'other'\n\ntrain['lang_guess'] = train['Report'].apply(guess_lang)\nlc = train['lang_guess'].value_counts()\n\nfig, axes = plt.subplots(1, 2, figsize=(15, 4))\nax = axes[0]\nax.hist(rep_len.clip(0, rep_len.quantile(.99)), bins=50, color=ICE, edgecolor='white')\nax.axvline(rep_len.median(), color=FIRE, ls='--', lw=1.4,\n           label=f'median = {rep_len.median():.0f} words')\nax.set_xlabel('report length, words'); ax.set_ylabel('studies')\nax.set_title('Report length'); ax.legend(fontsize=8)\n\nax = axes[1]\nax.bar(lc.index, lc.values, color=grad(max(len(lc), 3)), edgecolor='white', width=.6)\nax.set_title('Estimated report language (rough)'); ax.tick_params(axis='x', rotation=30)\nplt.tight_layout(); plt.show()\n\nprint('\\n=== sample report ===')\nif (has_label & has_report).any():\n    print(train.loc[has_label & has_report, 'Report'].iloc[0][:800])","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 1.7 Visual check â€” slices and laterality\n\nOne study across all of its series. This is also where the laterality story gets confirmed by eye: on a\ncoronal slice the medial compartment of a right knee and of a left knee sit on opposite sides of the frame.","metadata":{}},{"cell_type":"code","source":"def read_pixels(path):\n    ds = pydicom.dcmread(str(path))\n    img = ds.pixel_array.astype(np.float32)\n    img = img * float(getattr(ds, 'RescaleSlope', 1) or 1) + \\\n          float(getattr(ds, 'RescaleIntercept', 0) or 0)\n    if getattr(ds, 'PhotometricInterpretation', '') == 'MONOCHROME1':\n        img = img.max() - img\n    lo, hi = np.percentile(img, [0.5, 99.5])\n    return np.clip((img - lo) / (hi - lo + 1e-6), 0, 1), ds\n\n\ndemo_study = train_series['StudyInstanceUID'].iloc[0]\nsers = train_series[train_series.StudyInstanceUID == demo_study]\nfig, axes = plt.subplots(len(sers), 5, figsize=(15, 3 * len(sers)))\naxes = np.atleast_2d(axes)\nfor r, (_, s) in enumerate(sers.iterrows()):\n    files = sorted((CFG.comp_dir / 'train_series' / demo_study / s.SeriesInstanceUID).glob('*.dcm'))\n    idxs = np.linspace(0, len(files) - 1, 5).astype(int) if files else []\n    for c, i in enumerate(idxs):\n        try:\n            img, ds = read_pixels(files[i])\n            axes[r, c].imshow(img, cmap='gray')\n            if c == 0:\n                side = getattr(ds, 'Laterality', None) or getattr(ds, 'ImageLaterality', '?')\n                axes[r, c].set_title(\n                    f'{s.Anatomical_Plane} | FS={s.Fluid_Sensitive} | FAT={s.Fat_Suppression} | side={side}',\n                    fontsize=8, loc='left', color=FIRE)\n        except Exception as e:\n            axes[r, c].text(.5, .5, str(e)[:32], ha='center', fontsize=6)\n        axes[r, c].axis('off'); axes[r, c].grid(False)\nplt.suptitle(f'study {demo_study[:26]}...', y=1.0, fontweight='bold')\nplt.tight_layout(); plt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Part 2 â€” Data model\n\nThree entities and one policy.\n\n```\nStudy â”€â”€< Series â”€â”€< Slice\n  â”‚          â”‚\n  â”‚          â””â”€â”€ plane, fluid_sensitive, fat_sat, laterality, n_slices\n  â””â”€â”€ VIEW SLOTS:  SAG_FS Â· COR_FS Â· AX_FS   (+ presence mask)\n```\n\n**Series selection.** One series per slot:\n\n| Slot | Preferred series | Targets it carries |\n|---|---|---|\n| `SAG_FS` | sagittal fluid-sensitive, fat-sat preferred | ACL, menisci, contusion |\n| `COR_FS` | coronal fluid-sensitive, fat-sat preferred | MCL, OA compartments, contusion |\n| `AX_FS`  | axial fluid-sensitive | PF OA, synovitis, Baker's cyst, effusion |\n\nFallback cascade: exact match â†’ same plane without fat-sat â†’ same plane, anything â†’ slot marked empty\n(mask = 0) and ignored by the model. Ties are broken by slice count.\n\n**Laterality normalisation.** Every volume is mapped onto a right knee: coronal and axial get their X\naxis mirrored when `Laterality == 'L'`; sagittal gets its slice order reversed, because that is its\nmedialâ†”lateral axis. Only after this can `Medial *` and `Lateral *` be learnt at all.\n\n**Cache format.** One `npz` per study: `uint8 (V, D, H, W)` plus a `(V,)` mask. At `V=3, D=16, H=W=224`\nthat is roughly 2.4 MB per study.","metadata":{}},{"cell_type":"code","source":"VIEW_SPEC = {                       # slot -> (plane, fluid_sensitive, preferred fat_sat)\n    'SAG_FS': ('Sagittal', 1, 1),\n    'COR_FS': ('Coronal',  1, 1),\n    'AX_FS':  ('Axial',    1, 0),\n}\n\ndef assign_views(series_df: pd.DataFrame) -> dict:\n    \"\"\"For a single study, return {slot: SeriesInstanceUID or None}.\"\"\"\n    out = {}\n    for slot, (plane, fluid, prefer_fat) in VIEW_SPEC.items():\n        cand = series_df[series_df.Anatomical_Plane == plane]\n        if len(cand) == 0:\n            out[slot] = None\n            continue\n        tiers = [cand[(cand.Fluid_Sensitive == fluid) & (cand.Fat_Suppression == prefer_fat)],\n                 cand[cand.Fluid_Sensitive == fluid],\n                 cand]\n        chosen = next((t for t in tiers if len(t) > 0), None)\n        if chosen is None:\n            out[slot] = None\n            continue\n        if 'n_files' in chosen.columns:\n            chosen = chosen.sort_values('n_files', ascending=False)\n        out[slot] = chosen.iloc[0].SeriesInstanceUID\n    return out\n\n\ndef build_manifest(series_df: pd.DataFrame, root: Path, count_files=False) -> pd.DataFrame:\n    df = series_df.copy()\n    if count_files:\n        df['n_files'] = [len(list((root / r.StudyInstanceUID / r.SeriesInstanceUID).glob('*.dcm')))\n                         for r in df.itertuples()]\n    rows = []\n    for study, g in df.groupby('StudyInstanceUID'):\n        rec = {'StudyInstanceUID': study}\n        rec.update(assign_views(g))\n        rec['n_series'] = len(g)\n        rows.append(rec)\n    return pd.DataFrame(rows)\n\n\n# --- self-test on a synthetic study: sagittal fat-sat wins, axial is absent ---\n_demo = pd.DataFrame([\n    dict(SeriesInstanceUID='a', Anatomical_Plane='Sagittal', Fluid_Sensitive=1, Fat_Suppression=1),\n    dict(SeriesInstanceUID='b', Anatomical_Plane='Sagittal', Fluid_Sensitive=0, Fat_Suppression=0),\n    dict(SeriesInstanceUID='c', Anatomical_Plane='Coronal',  Fluid_Sensitive=1, Fat_Suppression=0),\n])\nassert assign_views(_demo) == {'SAG_FS': 'a', 'COR_FS': 'c', 'AX_FS': None}\nprint('assign_views self-test passed')\n\nmanifest = build_manifest(train_series, CFG.comp_dir / 'train_series')\nfill = (manifest[list(VIEW_SPEC)].notna().mean() * 100).round(1)\n\nfig, ax = plt.subplots(figsize=(7, 3.2))\nax.bar(fill.index, fill.values, color=grad(3), edgecolor='white', width=.5)\nax.set_ylim(0, 112); ax.set_ylabel('% of studies'); ax.set_title('View slot fill rate')\nfor i, v in enumerate(fill.values):\n    ax.text(i, v + 2.5, f'{v:.1f}', ha='center', fontsize=9, color=MIDLINE)\nplt.tight_layout(); plt.show()\n\nprint(f\"studies with no slot at all : {(manifest[list(VIEW_SPEC)].isna().all(axis=1)).sum()}\")\nprint(f\"studies with all three      : {(manifest[list(VIEW_SPEC)].notna().all(axis=1)).mean()*100:.1f}%\")\nmanifest.head()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.1 Preprocessing â€” DICOM to `uint8` volume\n\nThe details that break a naive implementation:\n\n1. **Slice order.** `sorted(glob('*.dcm'))` sorts by SOP UID, which is effectively random. Sort by\n   `ImagePositionPatient` projected onto the slice normal (the cross product of the two\n   `ImageOrientationPatient` vectors), falling back to `InstanceNumber`.\n2. **`MONOCHROME1`** must be inverted.\n3. **Normalisation** is robust and per volume â€” MR intensities are not in Hounsfield units and are not\n   comparable across scanners.\n4. **Anisotropic pixels** are corrected via `PixelSpacing` before the square resize.\n5. **Slice sampling** spans the whole volume rather than the middle window: Baker's cyst and MCL live\n   at the periphery.\n6. **Laterality** is normalised to a right knee.\n\nHeaders are read first with `stop_before_pixels=True`, so only the slices actually kept get decoded â€”\non a 300-slice series that alone is a large saving.","metadata":{}},{"cell_type":"code","source":"import cv2\ncv2.setNumThreads(0)          # essential: OpenCV threads deadlock inside DataLoader workers\n\n\ndef _slice_order(dss):\n    try:\n        iop = np.array(dss[0].ImageOrientationPatient, dtype=float)\n        n = np.cross(iop[:3], iop[3:])\n        pos = [float(np.dot(np.array(d.ImagePositionPatient, dtype=float), n)) for d in dss]\n    except Exception:\n        pos = [float(getattr(d, 'InstanceNumber', i) or i) for i, d in enumerate(dss)]\n    return np.argsort(pos)\n\n\ndef _laterality(dss):\n    for d in dss:\n        v = getattr(d, 'Laterality', None) or getattr(d, 'ImageLaterality', None)\n        if v in ('L', 'R'):\n            return v\n    try:                                     # LPS frame: a right knee sits at x < 0\n        return 'R' if float(dss[0].ImagePositionPatient[0]) < 0 else 'L'\n    except Exception:\n        return 'R'\n\n\ndef load_series_volume(series_dir: Path, plane: str,\n                       size=CFG.IMG_SIZE, n_slices=CFG.N_SLICES):\n    \"\"\"DICOM series -> uint8 array (n_slices, size, size), oriented as a right knee.\"\"\"\n    files = list(series_dir.glob('*.dcm'))\n    if not files:\n        return None\n\n    heads = []\n    for f in files:\n        try:\n            heads.append((f, pydicom.dcmread(str(f), stop_before_pixels=True)))\n        except Exception:\n            continue\n    if not heads:\n        return None\n\n    order = _slice_order([h[1] for h in heads])\n    lat = _laterality([h[1] for h in heads])\n    take = np.linspace(0, len(order) - 1, min(n_slices, len(order))).round().astype(int)\n    sel_files = [heads[order[i]][0] for i in take]\n\n    ps = getattr(heads[0][1], 'PixelSpacing', None)\n    sy, sx = (float(ps[0]), float(ps[1])) if ps else (1.0, 1.0)\n\n    frames = []\n    for f in sel_files:\n        try:\n            ds = pydicom.dcmread(str(f))\n            a = ds.pixel_array.astype(np.float32)\n        except Exception:\n            continue\n        a = a * float(getattr(ds, 'RescaleSlope', 1) or 1) + \\\n            float(getattr(ds, 'RescaleIntercept', 0) or 0)\n        if getattr(ds, 'PhotometricInterpretation', '') == 'MONOCHROME1':\n            a = a.max() - a\n        if a.ndim == 3:                                   # multiframe or RGB\n            a = a[..., 0] if a.shape[-1] <= 4 else a[0]\n        if abs(sy - sx) > 1e-3:\n            a = cv2.resize(a, None, fx=sx / min(sx, sy), fy=sy / min(sx, sy),\n                           interpolation=cv2.INTER_LINEAR)\n        frames.append(cv2.resize(a, (size, size), interpolation=cv2.INTER_AREA))\n    if not frames:\n        return None\n\n    vol = np.stack(frames)\n    lo, hi = np.percentile(vol, [0.5, 99.5])\n    vol = (np.clip((vol - lo) / (hi - lo + 1e-6), 0, 1) * 255).astype(np.uint8)\n\n    if lat == 'L':                                        # normalise to a right knee\n        vol = vol[:, :, ::-1] if plane in ('Coronal', 'Axial') else vol[::-1]\n    vol = np.ascontiguousarray(vol)\n\n    if vol.shape[0] < n_slices:                           # pad by repeating the last slice\n        vol = np.concatenate([vol, np.repeat(vol[-1:], n_slices - vol.shape[0], axis=0)], 0)\n    return vol\n\n\ndef cache_study(study: str, mrow, root: Path, out_dir: Path) -> bool:\n    dst = out_dir / f'{study}.npz'\n    if dst.exists():\n        return True\n    vols = np.zeros((CFG.N_VIEWS, CFG.N_SLICES, CFG.IMG_SIZE, CFG.IMG_SIZE), np.uint8)\n    mask = np.zeros(CFG.N_VIEWS, np.uint8)\n    for vi, slot in enumerate(CFG.VIEWS):\n        sid = mrow.get(slot)\n        if not isinstance(sid, str):\n            continue\n        v = load_series_volume(root / study / sid, VIEW_SPEC[slot][0])\n        if v is not None:\n            vols[vi], mask[vi] = v, 1\n    if mask.sum() == 0:\n        return False\n    np.savez_compressed(dst, vol=vols, mask=mask)\n    return True\n\nprint('preprocessing helpers ready')","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Preprocessing smoke test\n\nCache exactly one study and look at it before committing hours of compute. If the montage is upside\ndown, inverted or scrambled, fix it now rather than after eight shards.","metadata":{}},{"cell_type":"code","source":"_row = manifest.iloc[0]\n_t0 = time.time()\n_ok = cache_study(_row.StudyInstanceUID, _row, CFG.comp_dir / 'train_series', CFG.cache_dir)\n_dt = time.time() - _t0\nprint(f'cached: {_ok} in {_dt:.1f}s -> roughly {len(manifest)*_dt/3600/CFG.NUM_SHARDS:.1f} h per shard')\n\nif _ok:\n    z = np.load(CFG.cache_dir / f'{_row.StudyInstanceUID}.npz')\n    vol, msk = z['vol'], z['mask']\n    print('volume:', vol.shape, vol.dtype, '| mask:', msk,\n          '| intensity mean/std:', round(float(vol.mean()), 1), round(float(vol.std()), 1))\n    show = np.linspace(0, CFG.N_SLICES - 1, 8).astype(int)\n    fig, axes = plt.subplots(CFG.N_VIEWS, 8, figsize=(15, 2.1 * CFG.N_VIEWS))\n    for v in range(CFG.N_VIEWS):\n        for c, s in enumerate(show):\n            axes[v, c].imshow(vol[v, s], cmap='gray', vmin=0, vmax=255)\n            axes[v, c].axis('off'); axes[v, c].grid(False)\n            if c == 0:\n                axes[v, c].set_title(f'{CFG.VIEWS[v]} (present={msk[v]})',\n                                     loc='left', fontsize=9, color=FIRE)\n    plt.suptitle('Cached study â€” every volume normalised to a right knee', y=1.01, fontweight='bold')\n    plt.tight_layout(); plt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if CFG.RUN_PREPROCESS:\n    out_dir = CFG.cache_dir / f'shard{CFG.SHARD_ID}'\n    out_dir.mkdir(parents=True, exist_ok=True)\n    todo = manifest.iloc[CFG.SHARD_ID::CFG.NUM_SHARDS].reset_index(drop=True)\n    print(f'shard {CFG.SHARD_ID}/{CFG.NUM_SHARDS}: {len(todo)} studies -> {out_dir}')\n\n    t0, ok = time.time(), 0\n    root = CFG.comp_dir / 'train_series'\n    for i, r in todo.iterrows():\n        try:\n            ok += cache_study(r.StudyInstanceUID, r, root, out_dir)\n        except Exception as e:\n            print('failed', r.StudyInstanceUID[:16], type(e).__name__, e)\n        if (i + 1) % 100 == 0:\n            el = time.time() - t0\n            print(f'{i+1}/{len(todo)} | ok={ok} | {el/60:.1f} min | '\n                  f'ETA {(el/(i+1)*(len(todo)-i-1))/60:.1f} min', flush=True)\n    print('done:', ok)\n    # then: Save Version -> publish the output as a Kaggle Dataset -> list it in CFG.extra_cache","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"> **In practice:** run this notebook `NUM_SHARDS` times, changing `SHARD_ID` each time, publish every\n> output as its own dataset, and attach all of them to the training notebook. Inference needs no cache â€”\n> ~1300 studies are read straight from DICOM inside the DataLoader workers.\n\n# Part 3 â€” Teacher: report â†’ 12 labels\n\nThis is where most of the baseline's edge comes from. Hard labels are scarce, reports are nearly\nuniversal, and the teacher turns the second into the first.\n\nTwo components:\n\n1. **A multilingual lexicon with negation scope.** Interpretable, works on languages that barely appear\n   in the labelled subset, and catches \"ACL intact\" or \"pas de rupture\".\n2. **Character n-gram TF-IDF (3â€“5) plus logistic regression per label.** Character n-grams are\n   language-agnostic and pick up morphology on their own: `menisc`, `kruisband`, `Kreuzband` and\n   `meniskÃ¼s` all become features without any translation step.\n\nThe final soft label mixes both, and a hard label always wins where it exists. The teacher is scored\nhonestly by out-of-fold AUC on the labelled subset: a target at 0.95+ produces trustworthy\npseudo-labels, one at 0.70 needs its weight cut.","metadata":{}},{"cell_type":"code","source":"LEXICON = {\n    'ACL': dict(\n        pos=[r'\\bacl\\b', r'anterior cruciate', r'ligament(?:um)? crois[eÃ©] ant[eÃ©]rieur', r'\\blca\\b',\n             r'vorder\\w* kreuzband', r'voorste kruisband', r'\\bvkb\\b',\n             r'ligamento cruzado anterior', r'legamento crociato anteriore',\n             r'Ã¶n Ã§apraz baÄŸ', r'\\bÃ¶Ã§b\\b'],\n        cond=[r'tear|rupture|torn|ruptu\\w*|scheur|riss|rotura|rottura|lesi[oÃ³]n|d[eÃ©]chirure|yÄ±rtÄ±k|kopma']),\n    'MCL': dict(\n        pos=[r'\\bmcl\\b', r'medial collateral', r'ligament(?:um)? lat[eÃ©]ral interne',\n             r'inner\\w* seitenband', r'mediale?s? kollateralband', r'mediale collaterale band',\n             r'ligamento colateral medial', r'legamento collaterale mediale', r'medial kollateral'],\n        cond=[r'tear|rupture|torn|sprain|ruptu\\w*|riss|distors|lesi[oÃ³]n|d[eÃ©]chirure|yÄ±rtÄ±k']),\n    'Medial Meniscus': dict(\n        pos=[r'medial menisc\\w+', r'm[eÃ©]nisque interne', r'innenmeniskus', r'mediale meniscus',\n             r'menisco (?:medial|interno)', r'iÃ§ meniskÃ¼s', r'\\bmm\\b'],\n        cond=[r'tear|torn|ruptu\\w*|riss|scheur|rotura|lesi[oÃ³]n|d[eÃ©]chirure|yÄ±rtÄ±k|degenerat']),\n    'Lateral Meniscus': dict(\n        pos=[r'lateral menisc\\w+', r'm[eÃ©]nisque externe', r'aussenmeniskus|auÃŸenmeniskus',\n             r'laterale meniscus', r'menisco (?:lateral|externo)', r'dÄ±ÅŸ meniskÃ¼s', r'\\blm\\b'],\n        cond=[r'tear|torn|ruptu\\w*|riss|scheur|rotura|lesi[oÃ³]n|d[eÃ©]chirure|yÄ±rtÄ±k|degenerat']),\n    'Medial OA': dict(\n        pos=[r'medial (?:compartment|tibiofemoral|femorotibial)', r'compartiment interne',\n             r'mediale?n? (?:kompartiment|gelenkspalt)', r'mediale compartiment',\n             r'compartimento medial'],\n        cond=[r'osteoarthrit|arthros|gonarthros|arthrose|artrosis|artrose|cartilage loss|'\n              r'chondropath|chondral|knorpel|kraakbeen|kÄ±kÄ±rdak|osteofit|osteophyt|joint space narrow']),\n    'Lateral OA': dict(\n        pos=[r'lateral (?:compartment|tibiofemoral|femorotibial)', r'compartiment externe',\n             r'lateral\\w* (?:kompartiment|gelenkspalt)', r'laterale compartiment',\n             r'compartimento lateral'],\n        cond=[r'osteoarthrit|arthros|gonarthros|arthrose|artrosis|artrose|cartilage loss|'\n              r'chondropath|chondral|knorpel|kraakbeen|kÄ±kÄ±rdak|osteofit|osteophyt|joint space narrow']),\n    'PF OA': dict(\n        pos=[r'patellofemoral', r'f[eÃ©]moro-?patellaire', r'retropatell\\w+', r'patellofemoraal',\n             r'patelofemoral', r'femoropatelar', r'patellar cartilage', r'patella\\w* knorpel'],\n        cond=[r'osteoarthrit|arthros|arthrose|artrosis|chondropath|chondral|chondromalac|'\n              r'cartilage|knorpel|kraakbeen|osteophyt|osteofit']),\n    'Effusion': dict(\n        pos=[r'effusion', r'joint fluid', r'[Ã©e]panchement', r'gelenkerguss|erguss', r'gewrichtsvocht',\n             r'derrame(?: articular)?', r'versamento', r'eklem s[Ä±i]v[Ä±i]s[Ä±i]|efÃ¼zyon', r'hydrops'],\n        cond=[]),\n    'Synovitis': dict(\n        pos=[r'synovit\\w+', r'synovial (?:thickening|proliferation|hypertroph)', r'synovialitis',\n             r'sinovitis', r'sinovite', r'synovite', r'sinovit'],\n        cond=[]),\n    \"Baker's\": dict(\n        pos=[r\"baker'?s? cyst\", r'popliteal cyst', r'kyste (?:de )?baker|kyste poplit[eÃ©]',\n             r'baker[- ]?zyste|poplitealzyste', r'bakercyste', r'quiste de baker',\n             r'cisti di baker', r'baker kisti', r'gastrocnemio-?semimembranosus'],\n        cond=[]),\n    'Contusion': dict(\n        pos=[r'contusion', r'bone bruise', r'bone marrow (?:edema|oedema)', r'knochenmark[soÃ¶]dem',\n             r'beenmerg[oÃ¶]edeem', r'[oÃ³]edema (?:[oÃ³]seo|de m[eÃ©]dula)', r'contus[aÃ£]o',\n             r'contusione', r'kemik ili[gÄŸ]i [oÃ¶]dem', r'trabecular (?:edema|injury)'],\n        cond=[]),\n    'Fracture': dict(\n        pos=[r'fracture', r'fraktur', r'fractuur', r'fractura', r'frattura', r'k[Ä±i]r[Ä±i][kg]',\n             r'avulsion', r'segond'],\n        cond=[]),\n}\n\nNEG_CUES = [\n    r'\\bno\\b', r'\\bnot\\b', r'\\bwithout\\b', r'\\bintact\\b', r'\\bnormal\\w*', r'\\bunremarkable\\b',\n    r'\\bnegative\\b', r'\\bfree of\\b', r'\\bruled out\\b', r'\\bgeen\\b', r'\\bkein\\w*', r'\\bohne\\b',\n    r'\\bintakt\\b', r\"\\bpas d[eu']\", r'\\bsans\\b', r'\\baucun\\w*', r'\\babsence\\b', r'\\bsin\\b',\n    r'\\bsem\\b', r'\\bsenza\\b', r'\\bnessun\\w*', r'\\byok(?:tur)?\\b', r'\\bdo[gÄŸ]al\\b',\n    r'\\bregelrecht\\b', r'\\bunauff[aÃ¤]llig\\b', r'\\bnegatif\\b',\n]\nNEG_RE = re.compile('|'.join(NEG_CUES), re.I)\nCLAUSE_SPLIT = re.compile(r'[.;:\\n]')\n\n\ndef rule_score(text: str) -> dict:\n    \"\"\"Rule-based estimate per label: 0.03 / 0.10 / 0.15 / 0.90.\n\n    Key detail: negation is searched only inside the current clause. Otherwise\n    \"Kein Erguss. Innenmeniskus Riss\" marks the meniscal tear as negated.\n    \"\"\"\n    if not isinstance(text, str) or not text.strip():\n        return {t: np.nan for t in TARGETS}\n    t = ' ' + re.sub(r'\\s+', ' ', text.lower()) + ' '\n    out = {}\n    for label, spec in LEXICON.items():\n        pos_re = re.compile('|'.join(spec['pos']), re.I)\n        cond_re = re.compile('|'.join(spec['cond']), re.I) if spec['cond'] else None\n        hits = negated = conditioned = 0\n        for m in pos_re.finditer(t):\n            hits += 1\n            left = CLAUSE_SPLIT.split(t[max(0, m.start() - 80):m.start()])[-1]\n            right = CLAUSE_SPLIT.split(t[m.end():m.end() + 110])[0]\n            if NEG_RE.search(left) or NEG_RE.search(right[:40]):\n                negated += 1\n            if cond_re is None or cond_re.search(left) or cond_re.search(right):\n                conditioned += 1\n        if hits == 0:\n            out[label] = 0.03                       # never mentioned\n        elif conditioned == 0:\n            out[label] = 0.15                       # mentioned, no finding word nearby\n        elif negated >= conditioned:\n            out[label] = 0.10                       # explicitly negated\n        else:\n            out[label] = 0.90                       # asserted\n    return out","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- self-test: rerun this after every edit to the lexicon ---\nSANITY = {\n    'EN pos': 'Complete tear of the anterior cruciate ligament. Medial meniscus posterior horn tear. '\n              'Small joint effusion.',\n    'EN neg': 'The anterior cruciate ligament is intact. No meniscal tear. No joint effusion.',\n    'NL pos': 'Ruptuur van de voorste kruisband. Scheur mediale meniscus achterhoorn. Bakercyste aanwezig.',\n    'DE mix': 'Vorderes Kreuzband intakt. Kein Erguss. Innenmeniskus Riss im Hinterhorn.',\n    'FR mix': \"Rupture du ligament croisÃ© antÃ©rieur. Pas d'Ã©panchement articulaire. Kyste de Baker.\",\n    'ES pos': 'Rotura del menisco medial. Derrame articular moderado. Artrosis del compartimento medial.',\n}\n_s = pd.DataFrame({k: rule_score(v) for k, v in SANITY.items()}).T[TARGETS]\n\nassert _s.loc['EN pos', ['ACL', 'Medial Meniscus', 'Effusion']].min() > .5\nassert _s.loc['EN neg', ['ACL', 'Medial Meniscus', 'Effusion']].max() < .5\nassert _s.loc['DE mix', 'Medial Meniscus'] > .5 and _s.loc['DE mix', 'Effusion'] < .5\nassert _s.loc['FR mix', 'Effusion'] < .5 and _s.loc['FR mix', \"Baker's\"] > .5\nassert _s.loc['NL pos', ['ACL', 'Medial Meniscus', \"Baker's\"]].min() > .5\nprint('rule_score self-test passed â€” clause-scoped negation and French elision both handled')\ndisplay(_s.style.background_gradient(cmap=CMAP_SEQ, vmin=0, vmax=1).format('{:.2f}'))","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if CFG.RUN_TEACHER:\n    from sklearn.metrics import roc_auc_score, average_precision_score\n\n    t0 = time.time()\n    rules = pd.DataFrame([rule_score(x) for x in train['Report']], index=train.index)[TARGETS]\n    print(f'rules computed in {time.time()-t0:.1f}s')\n\n    ev, m = [], has_label & has_report\n    for c in TARGETS:\n        y, p = train.loc[m, c].values, rules.loc[m, c].values\n        if len(np.unique(y)) < 2:\n            continue\n        ev.append({'label': c, 'n_pos': int(y.sum()),\n                   'rule_auc': roc_auc_score(y, p), 'rule_ap': average_precision_score(y, p)})\n    ev = pd.DataFrame(ev).sort_values('rule_auc')\n    print('\\nRule quality on the labelled subset:')\n    display(ev.style.background_gradient(cmap=CMAP_SEQ, subset=['rule_auc'], vmin=.5, vmax=1)\n              .format({'rule_auc': '{:.3f}', 'rule_ap': '{:.3f}'}))","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Rules are the floor. The main teacher is character n-gram TF-IDF plus logistic regression, trained on\nthe labelled subset and scored out of fold. The rule outputs are appended as extra columns so the model\ngets a hint on languages that are rare in the labelled part.","metadata":{}},{"cell_type":"code","source":"if CFG.RUN_TEACHER:\n    from sklearn.feature_extraction.text import TfidfVectorizer\n    from sklearn.linear_model import LogisticRegression\n    from sklearn.model_selection import StratifiedKFold\n    from scipy import sparse\n\n    txt = train['Report'].fillna('').str.lower().values\n    fit_mask = (has_label & has_report).values\n\n    vec_char = TfidfVectorizer(analyzer='char_wb', ngram_range=(3, 5), min_df=5,\n                               max_features=300_000, sublinear_tf=True)\n    vec_word = TfidfVectorizer(analyzer='word', ngram_range=(1, 2), min_df=3,\n                               max_features=200_000, sublinear_tf=True)\n    X = sparse.hstack([vec_char.fit_transform(txt),\n                       vec_word.fit_transform(txt),\n                       sparse.csr_matrix(rules.values.astype(np.float32))]).tocsr()\n    print('feature matrix:', X.shape)\n\n    Xfit = X[fit_mask]\n    teacher = pd.DataFrame(index=train.index, columns=TARGETS, dtype=np.float32)\n    oof_rows = []\n\n    for c in TARGETS:\n        y = train.loc[fit_mask, c].values.astype(int)\n        if y.sum() < 10:\n            teacher[c] = rules[c]\n            continue\n        oof = np.zeros(len(y))\n        for tr, va in StratifiedKFold(5, shuffle=True, random_state=CFG.seed).split(Xfit, y):\n            lr = LogisticRegression(C=4.0, max_iter=2000, class_weight='balanced')\n            lr.fit(Xfit[tr], y[tr])\n            oof[va] = lr.predict_proba(Xfit[va])[:, 1]\n        teacher[c] = LogisticRegression(C=4.0, max_iter=2000, class_weight='balanced') \\\n                        .fit(Xfit, y).predict_proba(X)[:, 1]\n        oof_rows.append({'label': c, 'n_pos': int(y.sum()), 'text_auc': roc_auc_score(y, oof)})\n\n    teach_ev = pd.DataFrame(oof_rows).merge(ev[['label', 'rule_auc']], on='label', how='left')\n    teach_ev['gain'] = teach_ev.text_auc - teach_ev.rule_auc\n    teach_ev = teach_ev.sort_values('text_auc')\n    print(f'macro text AUC: {teach_ev.text_auc.mean():.4f}')\n\n    fig, ax = plt.subplots(figsize=(9, 4.4))\n    yy = np.arange(len(teach_ev))\n    ax.hlines(yy, teach_ev.rule_auc, teach_ev.text_auc, color='#cfd4da', lw=2, zorder=1)\n    ax.scatter(teach_ev.rule_auc, yy, s=55, color=ICE, label='rules only', zorder=2)\n    ax.scatter(teach_ev.text_auc, yy, s=55, color=FIRE, label='rules + TF-IDF LR', zorder=2)\n    ax.axvline(.85, color=MIDLINE, ls=':', lw=1, label='0.85 â€” below this, down-weight pseudo-labels')\n    ax.set_yticks(yy); ax.set_yticklabels(teach_ev.label); ax.set_xlim(.45, 1.02)\n    ax.set_xlabel('out-of-fold AUC, report -> label')\n    ax.grid(axis='x'); ax.grid(axis='y', visible=False)\n    ax.set_title('Teacher quality per target'); ax.legend(fontsize=8, loc='lower left')\n    plt.tight_layout(); plt.show()\n    display(teach_ev.round(3))","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if CFG.RUN_TEACHER:\n    # final target table for the student: hard label where available, teacher output otherwise\n    soft = teacher.copy()\n    soft.loc[has_label, TARGETS] = train.loc[has_label, TARGETS].values\n    soft.loc[~has_report & ~has_label, TARGETS] = np.nan        # unusable studies\n\n    targets_df = pd.concat([train[['StudyInstanceUID']], soft], axis=1)\n    targets_df['is_hard'] = has_label.astype(int)\n    targets_df['w'] = np.where(has_label, 1.0, CFG.pseudo_w)\n    targets_df = targets_df.dropna(subset=TARGETS)\n    targets_df.to_csv(CFG.work_dir / 'targets.csv', index=False)\n\n    print(f'training studies after the teacher: {len(targets_df):,} '\n          f'(hard {targets_df.is_hard.sum():,} / soft {(1-targets_df.is_hard).sum():,})')\n\n    cmp_df = pd.DataFrame({\n        'hard_prevalence': train.loc[has_label, TARGETS].mean(),\n        'soft_mean': targets_df.loc[targets_df.is_hard == 0, TARGETS].mean(),\n    })\n    fig, ax = plt.subplots(figsize=(9, 4))\n    yy = np.arange(12)\n    ax.barh(yy - .2, cmp_df.hard_prevalence, height=.38, color=ICE, label='hard prevalence')\n    ax.barh(yy + .2, cmp_df.soft_mean, height=.38, color=FIRE, label='mean soft label')\n    ax.set_yticks(yy); ax.set_yticklabels(cmp_df.index); ax.invert_yaxis()\n    ax.set_xlabel('rate'); ax.set_title('Hard prevalence versus teacher output on unlabelled studies')\n    ax.grid(axis='x'); ax.grid(axis='y', visible=False); ax.legend(fontsize=8)\n    plt.tight_layout(); plt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"> **Diagnostic:** a large gap between the two bars means either the labelled subset was enriched for\n> abnormalities rather than sampled at random, or the teacher is biased. The first is more likely, and\n> it is one more reason to validate the student on hard labels only.\n\n# Part 4 â€” Validation\n\nThree requirements:\n\n1. **Multilabel stratification**, otherwise `Fracture` and `Synovitis` vanish from some folds.\n2. **Grouping** by patient or duplicated report, against leakage.\n3. **Scoring on hard labels only.** Measuring AUC against pseudo-labels measures agreement with the\n   teacher, not with the truth.","metadata":{}},{"cell_type":"code","source":"def multilabel_stratified_kfold(Y: np.ndarray, n_splits=5, seed=42) -> np.ndarray:\n    \"\"\"Greedy iterative stratification (Sechidis et al.), no external dependency.\"\"\"\n    rng = np.random.default_rng(seed)\n    n, L = Y.shape\n    fold_of = np.full(n, -1)\n    desired = Y.sum(0)[None, :] / n_splits * np.ones((n_splits, 1))\n    fold_count = np.zeros(n_splits)\n    target_size = n / n_splits\n    order = rng.permutation(n)\n    order = order[np.argsort(Y[order].sum(1), kind='stable')[::-1]]     # label-rich rows first\n    for i in order:\n        lbl = np.where(Y[i] > 0)[0]\n        if len(lbl) == 0:\n            f = int(np.argmin(fold_count - target_size))\n        else:\n            rarest = lbl[np.argmin(Y[:, lbl].sum(0))]\n            need = desired[:, rarest]\n            cand = np.where(need == need.max())[0]\n            f = int(cand[np.argmin(fold_count[cand])])\n        fold_of[i] = f\n        fold_count[f] += 1\n        desired[f, lbl] -= 1\n    return fold_of\n\n\n# --- self-test: folds must be balanced in size and in rare positives ---\n_rng = np.random.default_rng(0)\n_Y = (_rng.random((2000, 12)) <\n      np.array([.2, .1, .25, .12, .15, .08, .2, .3, .05, .09, .13, .02])).astype(int)\n_f = multilabel_stratified_kfold(_Y, 5, 42)\n_sizes = np.bincount(_f)\n_rare = pd.DataFrame(_Y).groupby(_f).sum()[11]\nassert _sizes.max() - _sizes.min() <= 1, _sizes\nassert _rare.max() - _rare.min() <= 2, _rare.to_dict()\nprint(f'stratification self-test passed | fold sizes {_sizes.tolist()} | '\n      f'rarest-label positives {_rare.tolist()}')","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if CFG.RUN_TEACHER:\n    tdf = targets_df.copy()\n    tdf['fold'] = multilabel_stratified_kfold((tdf[TARGETS].values > .5).astype(int),\n                                              CFG.folds, CFG.seed)\n\n    # group by duplicated report when PatientID is unavailable\n    rep_hash = train.set_index('StudyInstanceUID')['Report'].fillna('').apply(\n        lambda s: hashlib.md5(s.strip().encode()).hexdigest() if len(s.strip()) > 30 else None)\n    tdf['grp'] = tdf.StudyInstanceUID.map(rep_hash)\n    g = tdf.dropna(subset=['grp']).groupby('grp')['fold'].nunique()\n    leaky = g[g > 1].index\n    if len(leaky):\n        first = tdf.dropna(subset=['grp']).groupby('grp')['fold'].first()\n        tdf.loc[tdf.grp.isin(leaky), 'fold'] = tdf.loc[tdf.grp.isin(leaky), 'grp'].map(first)\n        print(f'repaired {len(leaky)} groups that spanned several folds')\n\n    tdf.drop(columns=['grp']).to_csv(CFG.work_dir / 'folds.csv', index=False)\n    print('fold sizes:', tdf.fold.value_counts().sort_index().to_dict())\n\n    hard_by_fold = tdf[tdf.is_hard == 1].groupby('fold')[TARGETS].sum().astype(int)\n    fig, ax = plt.subplots(figsize=(11, 3.4))\n    sns.heatmap(hard_by_fold, annot=True, fmt='d', cmap=CMAP_SEQ, linewidths=.6, linecolor='white',\n                ax=ax, annot_kws={'size': 8}, cbar_kws={'shrink': .8})\n    ax.set_title('Hard positives per fold â€” the real precision of your validation')\n    plt.tight_layout(); plt.show()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Part 5 â€” Student: EfficientNetV2 2.5D with attention pooling\n\n```\n(B, V, D, H, W) uint8\n  â””â”€ 2.5D: every slice d becomes channels (d-1, d, d+1)\n      â””â”€ EfficientNetV2-S, shared across views â†’ (BÂ·VÂ·D, C)\n          â””â”€ + learned view embedding\n              â””â”€ attention pool over D  â†’ (B, V, C)\n                  â””â”€ masked attention pool over V â†’ (B, C)\n                      â””â”€ linear head â†’ 12 logits\n```\n\nWhy this shape:\n\n- one shared backbone across views â€” fewer parameters, more images per parameter;\n- attention over slices, because a finding occupies two or three slices out of sixteen and a mean\n  washes it out (this is precisely what `feats.mean(dim=2)` costs in the public baseline);\n- a mask over views, so a study without an axial series does not feed the pool with zeros;\n- an auxiliary per-view head, which stabilises the early epochs.\n\n**Augmentation:** shift, scale, rotation up to Â±10Â°, brightness and contrast jitter.\n**No horizontal flip** â€” it swaps medial and lateral.","metadata":{}},{"cell_type":"code","source":"if CFG.RUN_TRAIN or CFG.RUN_INFER:\n    import torch\n    import torch.nn as nn\n    import torch.nn.functional as F\n    from torch.utils.data import Dataset, DataLoader\n    import timm\n\n    DEVICE = 'cuda' if torch.cuda.is_available() else 'cpu'\n    print(f'device: {DEVICE} | timm {timm.__version__} | torch {torch.__version__}')\n\n    # ---- AMP compatibility across torch versions ----------------------------\n    _NEW_AMP = hasattr(torch.amp, 'GradScaler')\n\n    def make_scaler():\n        return (torch.amp.GradScaler('cuda', enabled=CFG.amp) if _NEW_AMP\n                else torch.cuda.amp.GradScaler(enabled=CFG.amp))\n\n    def autocast():\n        enabled = CFG.amp and torch.cuda.is_available()\n        return (torch.amp.autocast('cuda', enabled=enabled) if _NEW_AMP\n                else torch.cuda.amp.autocast(enabled=enabled))\n\n    # ---- backbone name resolution -------------------------------------------\n    def resolve_backbone(candidates, pretrained):\n        \"\"\"'efficientnetv2_s' is NOT a registered timm name â€” pick the first one that is.\"\"\"\n        available = set(timm.list_models())\n        with_weights = set(timm.list_models(pretrained=True))\n        for name in candidates:\n            base = name.split('.')[0]\n            if pretrained and name in with_weights:\n                return name\n            if not pretrained and (name in available or base in available):\n                return name if name in available else base\n        for name in candidates:                       # last resort: drop the weight tag\n            if name.split('.')[0] in available:\n                print(f'[!] falling back to {name.split(\".\")[0]} without a specific weight tag')\n                return name.split('.')[0]\n        raise RuntimeError(f'None of {candidates} exists in timm {timm.__version__}')\n\n    BACKBONE = resolve_backbone(CFG.backbone, CFG.pretrained)\n    print('backbone resolved to:', BACKBONE)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if CFG.RUN_TRAIN or CFG.RUN_INFER:\n\n    class KneeDataset(Dataset):\n        \"\"\"Reads pre-cached npz volumes. Missing studies yield zeros with a zero mask.\"\"\"\n\n        def __init__(self, df, cache_dirs, train_mode=True, targets=None):\n            self.df = df.reset_index(drop=True)\n            self.train_mode = train_mode\n            self.targets = targets\n            self.index = {}                       # study -> path, built once, not per item\n            for d in cache_dirs:\n                d = Path(d)\n                if d.exists():\n                    for p in d.glob('*.npz'):\n                        self.index.setdefault(p.stem, p)\n            present = sum(s in self.index for s in self.df.StudyInstanceUID)\n            print(f'  cache index: {len(self.index):,} studies | '\n                  f'{present:,} of {len(self.df):,} rows resolved')\n\n        def __len__(self):\n            return len(self.df)\n\n        def _augment(self, vol):\n            \"\"\"vol (V, D, H, W) uint8 -> float32 with geometric and photometric jitter. No flips.\"\"\"\n            V, D, H, W = vol.shape\n            out = np.empty_like(vol)\n            for v in range(V):\n                ang = np.random.uniform(-10, 10)\n                sc = np.random.uniform(0.9, 1.1)\n                tx, ty = np.random.uniform(-.06, .06, 2) * W\n                M = cv2.getRotationMatrix2D((W / 2, H / 2), ang, sc)\n                M[0, 2] += tx; M[1, 2] += ty\n                for d in range(D):\n                    out[v, d] = cv2.warpAffine(vol[v, d], M, (W, H), borderMode=cv2.BORDER_CONSTANT)\n            x = out.astype(np.float32) / 255.\n            return np.clip(x * np.random.uniform(.85, 1.15) + np.random.uniform(-.08, .08), 0, 1)\n\n        def __getitem__(self, i):\n            r = self.df.iloc[i]\n            p = self.index.get(r.StudyInstanceUID)\n            if p is None:\n                vol = np.zeros((CFG.N_VIEWS, CFG.N_SLICES, CFG.IMG_SIZE, CFG.IMG_SIZE), np.uint8)\n                mask = np.zeros(CFG.N_VIEWS, np.float32)\n            else:\n                z = np.load(p)\n                vol, mask = z['vol'], z['mask'].astype(np.float32)\n            x = self._augment(vol) if self.train_mode else vol.astype(np.float32) / 255.\n            item = {'x': torch.from_numpy(np.ascontiguousarray(x)),\n                    'mask': torch.from_numpy(mask)}\n            if self.targets is not None:\n                item['y'] = torch.tensor(r[self.targets].values.astype(np.float32))\n                item['w'] = torch.tensor(float(r.get('w', 1.0)))\n            return item\n\n\n    class AttnPool(nn.Module):\n        def __init__(self, dim, hidden=128):\n            super().__init__()\n            self.net = nn.Sequential(nn.Linear(dim, hidden), nn.Tanh(), nn.Linear(hidden, 1))\n\n        def forward(self, x, mask=None):                    # x: (B, N, C)\n            a = self.net(x).squeeze(-1)\n            if mask is not None:\n                a = a.masked_fill(mask < .5, -1e4)          # -1e4, not -inf: fp16 safe\n            a = a.softmax(-1)\n            return (x * a.unsqueeze(-1)).sum(1), a\n\n\n    class KneeNet(nn.Module):\n        def __init__(self, backbone=None, n_views=CFG.N_VIEWS, n_out=12,\n                     pretrained=CFG.pretrained, drop=CFG.drop_rate, grad_ckpt=CFG.grad_ckpt):\n            super().__init__()\n            self.enc = timm.create_model(backbone or BACKBONE, pretrained=pretrained,\n                                         num_classes=0, in_chans=3, drop_rate=drop)\n            if grad_ckpt:\n                try:\n                    self.enc.set_grad_checkpointing(True)\n                except Exception:\n                    pass\n            with torch.no_grad():\n                C = self.enc(torch.zeros(1, 3, 128, 128)).shape[-1]\n            self.view_emb = nn.Parameter(torch.zeros(n_views, C))\n            nn.init.normal_(self.view_emb, std=.02)\n            self.slice_pool = AttnPool(C)\n            self.view_pool = AttnPool(C)\n            self.drop = nn.Dropout(drop)\n            self.head = nn.Linear(C, n_out)\n            self.aux_head = nn.Linear(C, n_out)\n\n        @staticmethod\n        def to_25d(x):                                      # (B,V,D,H,W) -> (B,V,D,3,H,W)\n            prev = torch.cat([x[:, :, :1], x[:, :, :-1]], 2)\n            nxt = torch.cat([x[:, :, 1:], x[:, :, -1:]], 2)\n            return torch.stack([prev, x, nxt], 3)\n\n        def forward(self, x, mask):\n            B, V, D, H, W = x.shape\n            f = self.enc(self.to_25d(x).reshape(B * V * D, 3, H, W)).reshape(B, V, D, -1)\n            f = f + self.view_emb[None, :, None, :]\n            fv, _ = self.slice_pool(f.reshape(B * V, D, -1))\n            fv = fv.reshape(B, V, -1)\n            aux = self.aux_head(self.drop(fv))               # (B, V, 12)\n            fs, w_view = self.view_pool(fv, mask)            # (B, C)\n            return self.head(self.drop(fs)), aux, w_view\n\n    print('dataset and model defined')\n\n\n\nimport torch.quantization\n\ndef quantize_model(model):\n    \"\"\"Dynamic int8 quantization for CPU-fast inference.\"\"\"\n    model.eval()\n    quantized_model = torch.quantization.quantize_dynamic(\n        model, {nn.Linear}, dtype=torch.qint8\n    )\n    return quantized_model\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Model smoke test\n\nBuild the network, push a random batch of the exact training shape through it, and run one backward\npass. This catches an unknown backbone name, a shape mismatch or an out-of-memory in seconds, instead\nof twenty minutes into the first epoch. The masked view is checked explicitly: its attention weight\nmust be zero.","metadata":{}},{"cell_type":"code","source":"if CFG.RUN_TRAIN or CFG.RUN_INFER:\n    _m = KneeNet(pretrained=False).to(DEVICE)\n    print(f'parameters: {sum(p.numel() for p in _m.parameters())/1e6:.1f}M '\n          f'| feature dim: {_m.enc.num_features}')\n\n    _x = torch.rand(CFG.batch_size, CFG.N_VIEWS, CFG.N_SLICES,\n                    CFG.IMG_SIZE, CFG.IMG_SIZE, device=DEVICE)\n    _msk = torch.tensor([[1., 1., 0.]] * CFG.batch_size, device=DEVICE)   # axial missing on purpose\n    _y = (torch.rand(CFG.batch_size, 12, device=DEVICE) > .7).float()\n\n    _t0 = time.time()\n    with autocast():\n        _logit, _aux, _wv = _m(_x, _msk)\n        _loss = F.binary_cross_entropy_with_logits(_logit, _y)\n    _loss.backward()\n    if DEVICE == 'cuda':\n        torch.cuda.synchronize()\n\n    assert _logit.shape == (CFG.batch_size, 12), _logit.shape\n    assert _aux.shape == (CFG.batch_size, CFG.N_VIEWS, 12), _aux.shape\n    assert float(_wv[:, 2].abs().max()) < 1e-3, 'a masked view still receives attention weight'\n    print(f'forward + backward ok in {time.time()-_t0:.2f}s | logits {tuple(_logit.shape)} '\n          f'| view weights {_wv[0].detach().float().cpu().numpy().round(3)}')\n    if DEVICE == 'cuda':\n        print(f'peak GPU memory: {torch.cuda.max_memory_allocated()/2**30:.2f} GB of '\n              f'{torch.cuda.get_device_properties(0).total_memory/2**30:.1f} GB')\n        print('If that is close to the limit: lower CFG.batch_size or CFG.N_SLICES, raise CFG.accum.')\n\n    del _m, _x, _msk, _y, _logit, _aux, _wv, _loss\n    gc.collect()\n    if DEVICE == 'cuda':\n        torch.cuda.empty_cache()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if CFG.RUN_TRAIN:\n    from sklearn.metrics import roc_auc_score\n\n    def macro_auc(y, p):\n        aucs = {}\n        for i, c in enumerate(TARGETS):\n            yi = (y[:, i] > .5).astype(int)\n            aucs[c] = np.nan if yi.min() == yi.max() else roc_auc_score(yi, p[:, i])\n        return float(np.nanmean(list(aucs.values()))), aucs\n\n\n    def train_fold(fold, folds_df, cache_dirs):\n        tr = folds_df[folds_df.fold != fold]\n        va = folds_df[(folds_df.fold == fold) & (folds_df.is_hard == 1)]   # hard labels only\n        print(f'\\n=== fold {fold} | train {len(tr):,} | val (hard) {len(va):,} ===')\n\n        dl_tr = DataLoader(KneeDataset(tr, cache_dirs, True, TARGETS), batch_size=CFG.batch_size,\n                           shuffle=True, num_workers=CFG.num_workers, pin_memory=True,\n                           drop_last=True, persistent_workers=CFG.num_workers > 0)\n        dl_va = DataLoader(KneeDataset(va, cache_dirs, False, TARGETS), batch_size=CFG.batch_size,\n                           shuffle=False, num_workers=CFG.num_workers, pin_memory=True)\n\n        model = KneeNet().to(DEVICE)\n        bb = [p for n, p in model.named_parameters() if n.startswith('enc.')]\n        hd = [p for n, p in model.named_parameters() if not n.startswith('enc.')]\n        opt = torch.optim.AdamW([{'params': bb, 'lr': CFG.backbone_lr},\n                                 {'params': hd, 'lr': CFG.lr}], weight_decay=CFG.wd)\n        steps = max(1, len(dl_tr) // CFG.accum) * CFG.epochs\n        sched = torch.optim.lr_scheduler.OneCycleLR(opt, max_lr=[CFG.backbone_lr, CFG.lr],\n                                                    total_steps=steps, pct_start=.1)\n        scaler = make_scaler()\n\n        hist, best = [], -1\n        for ep in range(CFG.epochs):\n            model.train(); t0, run, nb = time.time(), 0., 0\n            opt.zero_grad(set_to_none=True)\n            for it, b in enumerate(dl_tr):\n                x, m = b['x'].to(DEVICE, non_blocking=True), b['mask'].to(DEVICE)\n                y, w = b['y'].to(DEVICE), b['w'].to(DEVICE)\n                with autocast():\n                    logit, aux, _ = model(x, m)\n                    l_main = F.binary_cross_entropy_with_logits(logit, y, reduction='none').mean(1)\n                    l_aux = F.binary_cross_entropy_with_logits(\n                        aux, y[:, None, :].expand_as(aux), reduction='none').mean(-1)\n                    l_aux = (l_aux * m).sum(1) / m.sum(1).clamp(min=1)\n                    loss = ((l_main + CFG.aux_w * l_aux) * w).mean() / CFG.accum\n                scaler.scale(loss).backward()\n                if (it + 1) % CFG.accum == 0:\n                    scaler.unscale_(opt)\n                    torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)\n                    scaler.step(opt); scaler.update()\n                    opt.zero_grad(set_to_none=True)\n                    if sched.last_epoch < steps - 1:\n                        sched.step()\n                run += loss.item() * CFG.accum; nb += 1\n\n            model.eval(); P, Y = [], []\n            with torch.no_grad(), autocast():\n                for b in dl_va:\n                    logit, _, _ = model(b['x'].to(DEVICE), b['mask'].to(DEVICE))\n                    P.append(logit.float().sigmoid().cpu().numpy()); Y.append(b['y'].numpy())\n            P, Y = np.concatenate(P), np.concatenate(Y)\n            auc, per = macro_auc(Y, P)\n            hist.append({'epoch': ep, 'loss': run / max(nb, 1), 'auc': auc, **per})\n            print(f'  epoch {ep} | loss {run/max(nb,1):.4f} | macro AUC {auc:.4f} '\n                  f'| {time.time()-t0:.0f}s', flush=True)\n            if auc > best:\n                best = auc\n                torch.save(model.state_dict(), CFG.work_dir / f'model_f{fold}.pt')\n\n        hist = pd.DataFrame(hist)\n        fig, axes = plt.subplots(1, 2, figsize=(14, 3.8))\n        axes[0].plot(hist.epoch, hist.loss, marker='o', color=ICE)\n        axes[0].set_title(f'fold {fold} â€” training loss'); axes[0].set_xlabel('epoch')\n        axes[1].plot(hist.epoch, hist.auc, marker='o', color=FIRE)\n        axes[1].axhline(best, ls='--', lw=1, color=MIDLINE, label=f'best {best:.4f}')\n        axes[1].set_title(f'fold {fold} â€” validation macro AUC')\n        axes[1].set_xlabel('epoch'); axes[1].legend(fontsize=8)\n        plt.tight_layout(); plt.show()\n\n        best_row = hist.loc[hist.auc.idxmax(), TARGETS].astype(float).sort_values()\n        fig, ax = plt.subplots(figsize=(9, 3.6))\n        ax.barh(best_row.index, best_row.values, color=grad(12), edgecolor='white', height=.7)\n        ax.axvline(.5, color=MIDLINE, ls=':', lw=1)\n        ax.set_xlim(.4, 1.0); ax.set_title(f'fold {fold} â€” per-target AUC at the best epoch')\n        ax.grid(axis='x'); ax.grid(axis='y', visible=False)\n        plt.tight_layout(); plt.show()\n\n        del model; gc.collect()\n        if DEVICE == 'cuda':\n            torch.cuda.empty_cache()\n        return best\n\n\n    folds_df = pd.read_csv(CFG.work_dir / 'folds.csv')\n    cache_dirs = [CFG.cache_dir, CFG.cache_dir / f'shard{CFG.SHARD_ID}'] + list(CFG.extra_cache)\n    scores = [train_fold(f, folds_df, cache_dirs) for f in CFG.train_folds]\n    print('\\nCV:', np.round(scores, 4), '| mean', round(float(np.mean(scores)), 4))","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Part 6 â€” Inference and submission\n\nThe submission notebook has no internet: `pretrained=False`, weights come from your own dataset, and\nDICOM is decoded on the fly by the very same functions â€” the preprocessing must be **bit-identical** to\ntraining, which is exactly why `load_series_volume` is shared rather than reimplemented here.\n\nAbout 1300 studies Ã— 3 views Ã— 16 slices â‰ˆ 62k images. On a T4 with EfficientNetV2-S and AMP that is\nroughly 30â€“45 minutes including DICOM decoding, comfortably inside the 9-hour limit.\n\nFold predictions are combined by **rank**, not by probability: the metric is rank-based and separate\nfolds are not calibrated against each other.","metadata":{}},{"cell_type":"code","source":"if CFG.RUN_INFER:\n    class TestDataset(Dataset):\n        \"\"\"Decodes DICOM on the fly â€” no cache needed for ~1300 studies.\"\"\"\n\n        def __init__(self, manifest_df, root):\n            self.df = manifest_df.reset_index(drop=True)\n            self.root = Path(root)\n\n        def __len__(self):\n            return len(self.df)\n\n        def __getitem__(self, i):\n            r = self.df.iloc[i]\n            vol = np.zeros((CFG.N_VIEWS, CFG.N_SLICES, CFG.IMG_SIZE, CFG.IMG_SIZE), np.uint8)\n            mask = np.zeros(CFG.N_VIEWS, np.float32)\n            for vi, slot in enumerate(CFG.VIEWS):\n                sid = r.get(slot)\n                if not isinstance(sid, str):\n                    continue\n                try:\n                    v = load_series_volume(self.root / r.StudyInstanceUID / sid, VIEW_SPEC[slot][0])\n                except Exception:\n                    v = None\n                if v is not None:\n                    vol[vi], mask[vi] = v, 1\n            return {'x': torch.from_numpy(vol.astype(np.float32) / 255.),\n                    'mask': torch.from_numpy(mask), 'idx': i}\n\n    test_manifest = test.merge(build_manifest(test_series, CFG.comp_dir / 'test_series'),\n                               on='StudyInstanceUID', how='left')\n    print('test manifest:', test_manifest.shape, '| studies with no usable view:',\n          int(test_manifest[list(VIEW_SPEC)].isna().all(axis=1).sum()))\n\n    dl = DataLoader(TestDataset(test_manifest, CFG.comp_dir / 'test_series'),\n                    batch_size=CFG.batch_size, shuffle=False, num_workers=CFG.num_workers)\n\n    ckpts = sorted(CFG.work_dir.glob('model_f*.pt')) or \\\n            sorted(Path('/kaggle/input').glob('*/model_f*.pt'))\n    print('checkpoints found:', len(ckpts), [c.name for c in ckpts])\n\n    ranked = np.zeros((len(test_manifest), 12), np.float32)\n    for ck in ckpts:\n        model = KneeNet(pretrained=False).to(DEVICE)\n        model.load_state_dict(torch.load(ck, map_location=DEVICE))\n        model.eval()\n        model = quantize_model(model)\n        fold_pred = np.zeros_like(ranked)\n        with torch.no_grad(), autocast():\n            for b in dl:\n                logit, _, _ = model(b['x'].to(DEVICE), b['mask'].to(DEVICE))\n                fold_pred[b['idx'].numpy()] = logit.float().sigmoid().cpu().numpy()\n        ranked += pd.DataFrame(fold_pred).rank(pct=True).values.astype(np.float32)\n        del model; gc.collect()\n        if DEVICE == 'cuda':\n            torch.cuda.empty_cache()\n    ranked /= max(1, len(ckpts))\n\n    sub = pd.DataFrame(ranked, columns=TARGETS)\n    sub.insert(0, 'StudyInstanceUID', test_manifest.StudyInstanceUID.values)\n    sub = sample_sub[['StudyInstanceUID']].merge(sub, on='StudyInstanceUID', how='left')\n    sub[TARGETS] = sub[TARGETS].fillna(0.5)\n    sub = sub[sample_sub.columns.tolist()]\nelse:\n    sub = sample_sub.copy()                 # always emit a valid file\n    print('RUN_INFER is off â€” writing the sample submission')\n\nsub.to_csv('submission.csv', index=False)\nassert list(sub.columns) == list(sample_sub.columns), 'column order must match sample_submission'\nassert len(sub) == len(sample_sub) and int(sub[TARGETS].isna().sum().sum()) == 0\nprint('submission.csv written:', sub.shape)\ndisplay(sub.head())","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Part 7 â€” Where to go from here\n\nOrdered by gain per hour of work.\n\n**1. Improve the teacher first â€” cheap, and it raises the ceiling for everything else.**\nPseudo-labels are the upper bound on the student. Check the per-target out-of-fold AUC from Part 3; for\nthe weak ones (`Synovitis` and `Lateral OA` are the usual suspects) add lexicon terms and inspect the\nheaviest n-grams of the logistic regression â€” they show exactly which words are missing. A multilingual\nencoder such as `xlm-roberta-base`, fine-tuned on the labelled subset, typically adds another 0.02â€“0.05\nAUC over TF-IDF on long reports.\n\n**2. Verify that laterality normalisation actually worked.**\nDiagnostic: train, then compare the AUC of `Medial OA` against `Lateral OA`. If both sit near 0.6 while\n`Effusion` is high, the sides are still mixed and the model cannot separate the compartments. Confirm on\na handful of coronal slices by eye.\n\n**3. Resolution and slice count.** 224 Ã— 16 is a compromise. Menisci and cartilage are small structures;\nmoving to 320 Ã— 32 usually helps the meniscus and OA targets noticeably, at a proportional cost in time.\n\n**4. More series per view.** One series per slot right now. A multiple-instance variant â€” every series\nof a study, with attention across series â€” removes the selection heuristic and brings T1/PD information\nback in.\n\n**5. Target-specific treatment.** `Fracture` and `Baker's` are local and rare; a dedicated head with\naggressive positive oversampling helps more than a larger backbone.\n\n**6. Ensembling.** Different backbones (EfficientNetV2-S, ConvNeXt-T, MaxViT-T) and different seeds,\ncombined by rank. On macro-AUC this reliably adds 0.01â€“0.02.\n\n**7. The efficiency track.** The metric charges you per second. Recipe: `tf_efficientnetv2_b0` or\n`resnet34`, 160â€“192 px, 12 slices, two views instead of three (sagittal plus coronal), fp16, one fold,\na larger batch. That usually costs 0.01â€“0.015 macro AUC for a 3â€“4Ã— speed-up, a favourable trade under\nthe efficiency formula.\n\n**What not to do**\n\n- Image plus text fusion at inference â€” there are no reports in the test set.\n- Horizontal flip in augmentation or TTA â€” it destroys the medial/lateral distinction.\n- Validating on pseudo-labels â€” that measures agreement with the teacher, not accuracy.\n- `pos_weight` of 100 on rare labels â€” unstable gradients, and macro-AUC ignores the prior anyway.\n- Probability calibration or threshold tuning â€” the metric is rank-based, so it changes nothing.","metadata":{}}]}