{"cells":[{"cell_type":"markdown","metadata":{},"source":"# 🦵 RSNA Knee — Report-Mining + EfficientNet Baseline  |  Public LB ~0.587\n\nOnly **~1.3%** of training studies are labeled — the rest carry a free-text **radiology report**.\nThis baseline shows the *simplest* end-to-end recipe that gets a real score:\n\n1. **Mine labels from the reports with simple multilingual keyword rules** (+ negation)\n2. Train a **2D EfficientNet-b0** on sampled MRI slices (study label broadcast to its slices)\n3. Predict per slice, **average to the study**, write `submission.csv`\n\nIt intentionally stays simple — an easy launchpad to improve on (better label mining with an LLM,\nhigher resolution, multi-view aggregation, TTA, ensembling…). **Public LB ≈ 0.587.**\n\n> ⭐ *If it helps you get started, an upvote is appreciated!*","id":"c00"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"import os, glob, re, unicodedata, numpy as np, pandas as pd, time\nCOMP='/kaggle/input/competitions/rsna-knee-abnormality-detection'\nif not os.path.isdir(COMP):\n    for r,d,_ in os.walk('/kaggle/input'):\n        if 'train_series' in d: COMP=r; break\nLABELS=['ACL','MCL','Medial Meniscus','Lateral Meniscus','Medial OA','Lateral OA',\n        'PF OA','Effusion','Synovitis',\"Baker's\",'Contusion','Fracture']\ntrain=pd.read_csv(f'{COMP}/train.csv'); print('train', train.shape)","id":"c01"},{"cell_type":"markdown","metadata":{},"source":"## 1. Rule-based report → labels (multilingual, negation-aware)\nAccent-insensitive keyword matching with a simple negation window. Deliberately conservative — an\nLLM does this much better, but rules are transparent and dependency-free. We validate on the small\nlabeled subset so you can see exactly how good (or not) it is.","id":"c02"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"def norm(t):\n    t=unicodedata.normalize('NFKD',str(t)); return ''.join(c for c in t if not unicodedata.combining(c)).lower()\nNEG=r'(no |not |without |sin |negative for|free of|ausencia de|unremarkable|intact|normal|preserved|conservad|within normal|sin signos)'\nTEAR=r'(tear|torn|rotura|ruptur|desgarro|scheur|lesion)'; INJ=r'(tear|torn|rotura|ruptur|sprain|esguince|injur|lesion|partial|complete)'\nOA=r'(osteoarthrit|osteoartros|artros|cartilage loss|degenerativ|joint space narrow|osteophyt|osteofit|chondromalac|condromalac)'\nRULES={'ACL':[(r'acl|anterior cruciate|cruzado anterior|lca',INJ)],'MCL':[(r'mcl|medial collateral|colateral medial|lcm',INJ)],\n 'Medial Meniscus':[(r'medial meniscus|menisco medial|menisco interno',TEAR)],'Lateral Meniscus':[(r'lateral meniscus|menisco lateral|menisco externo',TEAR)],\n 'Medial OA':[(r'medial (tibiofemoral|compartment|femorotibial)|compartimento medial',OA)],'Lateral OA':[(r'lateral (tibiofemoral|compartment|femorotibial)|compartimento lateral',OA)],\n 'PF OA':[(r'patellofemoral|patelofemoral|femoropatelar|patellar|trochlea|troclea',OA)],'Effusion':[(r'effusion|derrame|joint fluid|liquido articular',r'.')],\n 'Synovitis':[(r'synovitis|sinovitis|synovial (thicken|inflam)',r'.')],\"Baker's\":[(r\"baker|popliteal cyst|quiste popliteo\",r'.')],\n 'Contusion':[(r'bone (contusion|bruise|marrow edema)|contusion osea|marrow edema|subchondral edema',r'.')],'Fracture':[(r'fracture|fractura|fisura|avulsion',r'.')]}\nPRES={'Effusion','Synovitis',\"Baker's\",'Contusion','Fracture'}\ndef extract(rep):\n    t=norm(rep); sents=[s for s in re.split(r'[\\.\\n;]+',t) if s.strip()]; out={l:0 for l in LABELS}\n    for lab,rs in RULES.items():\n        for anat,dis in rs:\n            for s in sents:\n                for m in re.finditer(anat,s):\n                    if re.search(NEG,s[max(0,m.start()-45):m.start()]): continue\n                    if lab in PRES: out[lab]=1; break\n                    elif re.search(dis,s): out[lab]=1; break\n                if out[lab]: break\n            if out[lab]: break\n    return out\nlabels=pd.DataFrame([extract(r) for r in train['Report'].fillna('')])[LABELS]\nlabels.insert(0,'StudyInstanceUID',train['StudyInstanceUID'].values)\nprint('mined label prevalence:'); print(labels[LABELS].mean().round(3))","id":"c03"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"# validate rules on the labeled subset\nfrom sklearn.metrics import f1_score\ng=train[train[LABELS].notna().all(axis=1)]\nif len(g):\n    Y=g[LABELS].astype(int).values; P=labels.set_index('StudyInstanceUID').loc[g['StudyInstanceUID'],LABELS].values\n    print('rule macro-F1 vs %d gold: %.3f'%(len(g), np.mean([f1_score(Y[:,j],P[:,j],zero_division=0) for j in range(12)])))\nprint('=> rules are a weak labeler (an LLM lifts this a lot) — but enough for a baseline.')","id":"c04"},{"cell_type":"markdown","metadata":{},"source":"## 2. DICOM loading + slice sampling\nMRI has no fixed intensity scale, so we percentile-normalise each slice. We sample a few slices from\nthe middle of each series and a few series per study to keep it fast.","id":"c05"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"import subprocess, sys\ntry: import gdcm\nexcept Exception: subprocess.run([sys.executable,'-m','pip','install','-q','python-gdcm','pylibjpeg','pylibjpeg-libjpeg','pylibjpeg-openjpeg'],check=False)\nimport pydicom, cv2\nIMG=224; SLICES=4; MAXSER=4\ndef read(p):\n    try: d=pydicom.dcmread(p); a=d.pixel_array.astype(np.float32)\n    except Exception: return None\n    lo,hi=np.percentile(a,1),np.percentile(a,99)\n    if hi<=lo: return None\n    a=np.clip((a-lo)/(hi-lo),0,1)\n    if getattr(d,'PhotometricInterpretation','')=='MONOCHROME1': a=1-a\n    return cv2.resize(a,(IMG,IMG))\ndef slice_paths(sd,k=SLICES):\n    fs=sorted(glob.glob(sd+'/*.dcm'));\n    if not fs: return []\n    n=len(fs); fs=fs[int(n*.2):int(n*.8)] or fs\n    idx=np.linspace(0,len(fs)-1,min(k,len(fs))).round().astype(int); return [fs[i] for i in idx]","id":"c06"},{"cell_type":"markdown","metadata":{},"source":"## 3. Train EfficientNet-b0 (subsampled for a quick, runnable demo)\nWe broadcast each study's mined labels to its slices, train a 12-way multi-label head, then average\nslice predictions to the study. *For a real run, use all studies and more slices/epochs.*","id":"c07"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"import torch, torch.nn as nn, timm\nfrom torch.utils.data import Dataset, DataLoader\nstudies=train['StudyInstanceUID'].tolist(); import random; random.Random(0).shuffle(studies)\nN=1200  # subsample for demo speed; raise for the full run\nstudies=studies[:N]; val=set(studies[:200]); tr=[s for s in studies if s not in val]\nymap=labels.set_index('StudyInstanceUID')[LABELS].astype(np.float32)\ndef index(sts):\n    it=[]\n    for st in sts:\n        for se in sorted(glob.glob(f'{COMP}/train_series/{st}/*'))[:MAXSER]:\n            for sp in slice_paths(se): it.append((st,sp))\n    return it\nclass DS(Dataset):\n    def __init__(s,it,aug): s.it=it; s.aug=aug\n    def __len__(s): return len(s.it)\n    def __getitem__(s,i):\n        st,sp=s.it[i]; a=read(sp)\n        a=np.zeros((IMG,IMG),np.float32) if a is None else a\n        if s.aug and random.random()<.5: a=a[:,::-1].copy()\n        return torch.from_numpy(a)[None].repeat(3,1,1), torch.from_numpy(ymap.loc[st].values), st\ntri,vai=index(tr),index(val); print('train slices',len(tri),'val slices',len(vai))\ndev='cuda'; model=timm.create_model('tf_efficientnet_b0.ns_jft_in1k',pretrained=True,num_classes=12,in_chans=3).to(dev)\nopt=torch.optim.AdamW(model.parameters(),lr=3e-4,weight_decay=1e-4); lf=nn.BCEWithLogitsLoss()\nsc=torch.cuda.amp.GradScaler()\nfor ep in range(2):\n    model.train(); tot=0; dl=DataLoader(DS(tri,True),batch_size=32,shuffle=True,num_workers=2,drop_last=True)\n    for x,y,_ in dl:\n        x,y=x.to(dev),y.to(dev); opt.zero_grad()\n        with torch.cuda.amp.autocast(): loss=lf(model(x),y)\n        sc.scale(loss).backward(); sc.step(opt); sc.update(); tot+=loss.item()\n    print(f'epoch {ep} loss {tot/len(dl):.4f}')","id":"c08"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"# study-level CV AUC on the held-out fold\nfrom sklearn.metrics import roc_auc_score\nmodel.eval(); pred={}\nwith torch.no_grad(), torch.cuda.amp.autocast():\n    for x,y,sts in DataLoader(DS(vai,False),batch_size=64,num_workers=2):\n        p=torch.sigmoid(model(x.to(dev))).float().cpu().numpy()\n        for k,st in enumerate(sts): pred.setdefault(st,[]).append(p[k])\norder=list(pred.keys())\nP=np.array([np.mean(pred[st],0) for st in order]); Y=np.array([ymap.loc[st].values for st in order])\naucs=[roc_auc_score(Y[:,j],P[:,j]) for j in range(12) if len(np.unique(Y[:,j]))>1]\nprint('study-level macro-AUC (vs mined labels): %.3f on %d val studies'%(np.mean(aucs),len(pred)))","id":"c09"},{"cell_type":"markdown","metadata":{},"source":"## 4. Inference → `submission.csv`\nSame recipe on the test studies: sample slices, predict, average to the study. On the real (hidden)\ntest set this pipeline scores **≈ 0.587** public LB. Improve it from here — the biggest lever by far\nis **better label mining from the reports (try an LLM)**.","id":"c10"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"sub=pd.read_csv(f'{COMP}/sample_submission.csv'); tst=sub['StudyInstanceUID'].tolist()\nrows=[]\nfor st in tst:\n    imgs=[]\n    for se in sorted(glob.glob(f'{COMP}/test_series/{st}/*'))[:MAXSER]:\n        for sp in slice_paths(se):\n            a=read(sp)\n            if a is not None: imgs.append(a)\n    if not imgs: rows.append([0.5]*12); continue\n    x=torch.from_numpy(np.stack(imgs))[:,None].repeat(1,3,1,1).to(dev)\n    with torch.no_grad(), torch.cuda.amp.autocast(): p=torch.sigmoid(model(x)).float().cpu().numpy()\n    rows.append(p.mean(0).tolist())\nout=pd.DataFrame(rows,columns=LABELS); out.insert(0,'StudyInstanceUID',tst)\nout.to_csv('submission.csv',index=False); print(out.shape); out.head()","id":"c11"},{"cell_type":"markdown","metadata":{},"source":"## Where to go next (all big levers)\n- **Mine labels with an LLM** instead of keyword rules — this is *the* lever in this competition.\n- Higher resolution, more slices, and smarter **slice→study aggregation** (e.g. top-k, attention).\n- **Route** by plane/sequence; add **TTA** and **fold/backbone ensembling** (rank-average).\n- Trust a **study-grouped, label-stratified CV** over the public LB.\n\n⭐ **Upvote if useful** — and share your improvements in the comments!","id":"c12"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python"}},"nbformat":4,"nbformat_minor":5}