{"cells":[{"cell_type":"markdown","metadata":{},"source":"# RSNA Knee MRI — EDA and Working Starter\n\nEverything you need to get moving in this competition: what the data looks like, how to actually load it, a working DICOM pipeline, and a valid baseline `submission.csv` — plus three setup issues that cost me several failed runs, documented here so they don't cost you any.\n\n## ⚠️ Issue 1 — the data is not at the usual path\n\nThis competition mounts under a **new layout**: `/kaggle/input/competitions/rsna-knee-abnormality-detection/` (note the extra `competitions/` level). Hardcoding the classic `/kaggle/input/<slug>/` gives `FileNotFoundError`. The robust fix is a glob (used below), which works in both layouts.\n\n## ⚠️ Issue 2 — the P100 GPU no longer works with the default PyTorch\n\nThe current PyTorch build in Kaggle's image has dropped support for the Tesla **P100** (compute capability 6.0) — the first `.cuda()` op throws `CUDA error: no kernel image is available for execution on the device`. **Select the T4 x2 accelerator, not P100**, in Notebook Settings (or pass `--accelerator NvidiaTeslaT4` if pushing via the CLI/API)."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import os, glob, warnings\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nwarnings.filterwarnings('ignore')\n\n# Issue 1 fix: locate the data wherever it's mounted\nPATH = os.path.dirname(glob.glob('/kaggle/input/**/train.csv', recursive=True)[0])\nprint('data root:', PATH)\n\ntrain = pd.read_csv(os.path.join(PATH, 'train.csv'))\nseries = pd.read_csv(os.path.join(PATH, 'train_series.csv'))\nLABELS = list(train.columns[2:14])\nprint(f'{len(train)} studies, {len(series)} series, {len(LABELS)} target labels')"},{"cell_type":"markdown","metadata":{},"source":"## The task\n\nOne row per **study** (a knee MRI exam). Each study has several **series** (different pulse sequences / anatomical planes), stored as DICOM slices. We predict 12 independent probabilities per study — one per abnormality — scored by **macro-averaged AUC**.\n\nA detail worth noticing: `train.csv` contains a `Report` column with the radiology report (in one of ~9 languages) for every **training** study. The hidden test set won't have reports — as Issue 3 below explains, these reports are not a bonus, they're where your training labels have to come from."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"prev = train[LABELS].mean().sort_values()\nfig, ax = plt.subplots(figsize=(8, 4.5))\nax.barh(prev.index, prev.values * 100, color='#4477AA')\nfor i, v in enumerate(prev.values):\n    ax.text(v*100 + 0.5, i, f'{v*100:.0f}%', va='center', fontsize=9)\nax.set_xlabel('prevalence (%)'); ax.set_title(f'Label prevalence (n={int(train[LABELS].notna().all(axis=1).sum())} annotated studies)')\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","metadata":{},"source":"Labels range from ~15% (MCL) to ~60% (Effusion) — no extreme rarity, which is friendly for AUC.\n\n## ⚠️ Issue 3 — only ~58 of 4,407 studies actually have labels\n\nThis is the big one. Run `train[LABELS].isna().sum()` and you'll find **~99% of the label cells are NaN** — only a small annotated subset carries the 12 targets. Pandas `mean()` silently skips NaN, so naive prevalence plots (including an earlier version of this one) look like they cover the full corpus when they don't. Practical consequences:\n\n- You **cannot** train directly on `train.csv` labels at scale — a naive `BCEWithLogitsLoss` on raw label arrays will propagate NaN and silently destroy training.\n- The intended path is **deriving targets from the `Report` column** (present for all training studies, absent at test time) — via multilingual rule extraction or an LLM pass — then training images against those derived labels, holding out the annotated subset as a clean validation set.\n- Reports are multilingual (my rough count: ~37% English, then Spanish, Greek, Cyrillic-script, German, Croatian/Serbian, Turkish, Dutch, French) and ~130 studies share byte-identical template reports — keep identical-report groups on the same side of your train/val split or you leak."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"co = train[LABELS].corr()\nfig, ax = plt.subplots(figsize=(7.5, 6))\nim = ax.imshow(co, cmap='RdBu_r', vmin=-1, vmax=1)\nax.set_xticks(range(len(LABELS))); ax.set_xticklabels(LABELS, rotation=45, ha='right', fontsize=8)\nax.set_yticks(range(len(LABELS))); ax.set_yticklabels(LABELS, fontsize=8)\nfor i in range(len(LABELS)):\n    for j in range(len(LABELS)):\n        ax.text(j, i, f'{co.iloc[i,j]:.2f}', ha='center', va='center', fontsize=6)\nplt.colorbar(im, shrink=0.8); ax.set_title('Label correlations')\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","metadata":{},"source":"The osteoarthritis cluster (Medial/Lateral/PF OA) and Effusion↔Synovitis correlate strongly — a multi-label head sharing a backbone should exploit this for free."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"print('series per study:')\nprint(series.groupby('StudyInstanceUID').size().describe().round(1))\ncomp = series.groupby(['Anatomical_Plane', 'Fluid_Sensitive']).size().unstack()\ncomp.plot(kind='bar', figsize=(7, 3.5), color=['#BBBBBB', '#4477AA'])\nplt.title('Series composition: plane x fluid-sensitivity'); plt.ylabel('series count')\nplt.legend(['not fluid-sensitive', 'fluid-sensitive']); plt.xticks(rotation=0)\nplt.tight_layout(); plt.show()\n\nhas_sagfs = series[(series['Anatomical_Plane']=='Sagittal') & (series['Fluid_Sensitive']==1)]['StudyInstanceUID'].nunique()\nprint(f'studies with a sagittal fluid-sensitive series: {has_sagfs}/{train.StudyInstanceUID.nunique()}')"},{"cell_type":"markdown","metadata":{},"source":"## Loading DICOM volumes — working code\n\nMedian 5 series per study. Sagittal fluid-sensitive series (present in nearly every study) are the radiologist's workhorse for ACL/meniscus/effusion — a strong single-series choice for a first model. Below: a loader that reads a series, orders slices by `InstanceNumber`, windows intensities by percentile, and samples a fixed number of slices."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"import pydicom\nimport cv2\n\ndef load_volume(study_uid, series_uid, n_slices=16, size=224, split='train_series'):\n    sdir = os.path.join(PATH, split, study_uid, series_uid)\n    slices = []\n    for f in glob.glob(os.path.join(sdir, '*.dcm')):\n        ds = pydicom.dcmread(f)\n        slices.append((int(getattr(ds, 'InstanceNumber', 0)), ds.pixel_array.astype(np.float32)))\n    slices.sort(key=lambda x: x[0])\n    take = np.linspace(0, len(slices)-1, n_slices).round().astype(int)\n    out = np.zeros((n_slices, size, size), dtype=np.uint8)\n    for i, t in enumerate(take):\n        img = slices[t][1]\n        lo, hi = np.percentile(img, [1, 99])\n        img = np.clip((img - lo) / max(hi - lo, 1e-6), 0, 1)\n        out[i] = (cv2.resize(img, (size, size)) * 255).astype(np.uint8)\n    return out\n\n# visualize one study's sagittal fluid-sensitive series\nex_study = has_ex = series[(series['Anatomical_Plane']=='Sagittal') & (series['Fluid_Sensitive']==1)].iloc[0]\nvol = load_volume(ex_study['StudyInstanceUID'], ex_study['SeriesInstanceUID'])\nfig, axes = plt.subplots(2, 8, figsize=(16, 4.2))\nfor i, ax in enumerate(axes.flat):\n    ax.imshow(vol[i], cmap='gray'); ax.axis('off')\nrow = train[train['StudyInstanceUID']==ex_study['StudyInstanceUID']]\npos = [c for c in LABELS if row.iloc[0][c]==1] if len(row) else []\nplt.suptitle(f\"Sagittal fluid-sensitive, 16 sampled slices | positives: {', '.join(pos) or 'none'}\")\nplt.tight_layout(); plt.show()"},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"# slice counts and native resolutions across a sample of series\nrng = np.random.RandomState(0)\nsample = series.sample(30, random_state=0)\nstats = []\nfor _, r in sample.iterrows():\n    files = glob.glob(os.path.join(PATH, 'train_series', r['StudyInstanceUID'], r['SeriesInstanceUID'], '*.dcm'))\n    if files:\n        ds = pydicom.dcmread(files[0])\n        stats.append({'n_slices': len(files), 'rows': ds.Rows, 'cols': ds.Columns,\n                      'plane': r['Anatomical_Plane']})\nsdf = pd.DataFrame(stats)\nprint(sdf.groupby('plane')[['n_slices','rows','cols']].median())\nprint('\\nslice count range:', sdf.n_slices.min(), '-', sdf.n_slices.max())"},{"cell_type":"markdown","metadata":{},"source":"## Valid baseline submission\n\nConstant per-label train prevalence — scores 0.5 AUC by construction, but validates the end-to-end hidden-test pipeline (always do this before investing in models)."},{"cell_type":"code","execution_count":null,"metadata":{},"outputs":[],"source":"test = pd.read_csv(os.path.join(PATH, 'test.csv'))\nsub = pd.DataFrame({'StudyInstanceUID': test['StudyInstanceUID']})\nfor c in LABELS:\n    sub[c] = train[c].mean()\nsub.to_csv('submission.csv', index=False)\nsub"},{"cell_type":"markdown","metadata":{},"source":"## Modeling roadmap (what I'm building on top of this)\n\n0. **Label derivation first** (see Issue 3): extract the 12 targets from the multilingual reports; validate the extractor against the annotated subset.\n1. **2.5D CNN**: sagittal fluid-sensitive series → 16 slices → pretrained 2D backbone per slice → pool → 12 sigmoid heads trained on derived labels. Simple, fast, strong baseline.\n2. **Multi-series fusion**: axial + coronal towers for the OA / patellofemoral labels that sagittal alone under-serves.\n3. **Report distillation**: use the train-only Spanish reports to train a text teacher, distill into the image model.\n4. Per-label attention pooling over slices (different abnormalities live on different slices).\n\nIf this saved you time, an upvote helps others find it. Questions welcome in the comments — happy to help debug any of the three issues."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11"}},"nbformat":4,"nbformat_minor":5}