{"cells":[{"cell_type":"markdown","id":"40693d55","metadata":{},"source":"# RSNA Knee Abnormality Detection — Exploratory Data Analysis\n\nGoal: build a solid mental model of this dataset — label coverage, report\nlanguage and structure, and DICOM series composition — before committing\nto a modeling approach for the 12 knee-MRI findings this competition asks\nfor (ACL, MCL, medial/lateral meniscus, medial/lateral/PF osteoarthritis,\neffusion, synovitis, Baker's cyst, contusion, fracture).\n"},{"cell_type":"code","execution_count":1,"id":"18fc5ca0","metadata":{"execution":{"iopub.execute_input":"2026-08-12T22:46:31.37698Z","iopub.status.busy":"2026-08-12T22:46:31.37639Z","iopub.status.idle":"2026-08-12T22:46:33.483164Z","shell.execute_reply":"2026-08-12T22:46:33.482039Z"}},"outputs":[],"source":"import os\nimport re\nimport subprocess\nimport sys\nfrom collections import Counter, defaultdict\n\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\n\nsns.set_theme(style=\"whitegrid\")\npd.set_option(\"display.max_colwidth\", 120)\n\n# A couple of packages used below aren't guaranteed to be preinstalled;\n# quiet-install them (no-op if already present).\nsubprocess.run(\n    [sys.executable, \"-m\", \"pip\", \"install\", \"-q\",\n     \"langdetect\", \"pylibjpeg\", \"pylibjpeg-libjpeg\", \"pylibjpeg-openjpeg\", \"python-gdcm\"],\n    check=False,\n)\n\n# The competition dataset is mounted somewhere under /kaggle/input, but not\n# necessarily at a flat /kaggle/input/<competition-slug>/ path - it can land\n# nested (e.g. under a \"competitions/<slug>/\" folder). Search for whichever\n# directory actually holds train.csv rather than assuming a fixed path.\nKAGGLE_INPUT_ROOT = \"/kaggle/input\"\nprint(f\"{KAGGLE_INPUT_ROOT} top-level contents: {sorted(os.listdir(KAGGLE_INPUT_ROOT))}\")\n\nDATA_DIR = None\nfor dirpath, dirnames, filenames in os.walk(KAGGLE_INPUT_ROOT):\n    if \"train.csv\" in filenames and \"train_series.csv\" in filenames:\n        DATA_DIR = dirpath\n        break  # stop before descending into train_series/ (24k+ subfolders)\n\nif DATA_DIR is None:\n    raise FileNotFoundError(\n        f\"No train.csv/train_series.csv pair found anywhere under {KAGGLE_INPUT_ROOT}. \"\n        \"Make sure the competition dataset is attached to this notebook.\"\n    )\n\nWORK_DIR = \"/kaggle/working\"\nos.makedirs(WORK_DIR, exist_ok=True)\n\nLABEL_COLS = [\n    \"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\", \"Medial OA\",\n    \"Lateral OA\", \"PF OA\", \"Effusion\", \"Synovitis\", \"Baker's\", \"Contusion\", \"Fracture\",\n]\n\nprint(f\"DATA_DIR={DATA_DIR}\")\nprint(f\"WORK_DIR={WORK_DIR}\")\n"},{"cell_type":"markdown","id":"6ff483fd","metadata":{},"source":"## 1. Load metadata\n\nLoad the study-level labels/reports, series-level descriptors, and the\nfiles provided for the test set.\n"},{"cell_type":"code","execution_count":2,"id":"ba778aa6","metadata":{"execution":{"iopub.execute_input":"2026-08-12T22:46:33.486454Z","iopub.status.busy":"2026-08-12T22:46:33.486051Z","iopub.status.idle":"2026-08-12T22:46:33.606053Z","shell.execute_reply":"2026-08-12T22:46:33.604821Z"}},"outputs":[],"source":"train = pd.read_csv(os.path.join(DATA_DIR, \"train.csv\"))\ntrain_series = pd.read_csv(os.path.join(DATA_DIR, \"train_series.csv\"))\ntest = pd.read_csv(os.path.join(DATA_DIR, \"test.csv\"))\ntest_series = pd.read_csv(os.path.join(DATA_DIR, \"test_series.csv\"))\nsample_sub = pd.read_csv(os.path.join(DATA_DIR, \"sample_submission.csv\"))\n\nfor name, df in [(\"train\", train), (\"train_series\", train_series),\n                  (\"test\", test), (\"test_series\", test_series),\n                  (\"sample_submission\", sample_sub)]:\n    print(f\"{name:20s} shape={df.shape}  columns={list(df.columns)}\")\n"},{"cell_type":"code","execution_count":3,"id":"4d69bcab","metadata":{"execution":{"iopub.execute_input":"2026-08-12T22:46:33.608784Z","iopub.status.busy":"2026-08-12T22:46:33.608534Z","iopub.status.idle":"2026-08-12T22:46:33.623991Z","shell.execute_reply":"2026-08-12T22:46:33.622858Z"}},"outputs":[],"source":"train.head(3)\n"},{"cell_type":"markdown","id":"c7f0b283","metadata":{},"source":"Note: the competition's dataset description lists a `PatientSex` column on\n`train.csv` — it isn't actually present in the file. It *does* exist as a\nDICOM tag (`Patient's Sex`, 0010,0040) though, confirmed below in the\nDICOM header section. Worth rechecking the forum periodically since the\ncompetition is very new (started 7 days ago) and this kind of\ndescription/data drift is common early on.\n"},{"cell_type":"markdown","id":"e9634260","metadata":{},"source":"## 2. Label coverage — the core constraint of this competition\n\n`train.csv` has one row per study, but per the dataset description only a\nsmall subset actually carries the 12 ground-truth labels. Let's quantify\nexactly how small.\n"},{"cell_type":"code","execution_count":4,"id":"22895b79","metadata":{"execution":{"iopub.execute_input":"2026-08-12T22:46:33.626708Z","iopub.status.busy":"2026-08-12T22:46:33.626437Z","iopub.status.idle":"2026-08-12T22:46:33.635919Z","shell.execute_reply":"2026-08-12T22:46:33.634897Z"}},"outputs":[],"source":"n_filled = train[LABEL_COLS].notna().sum(axis=1)\nhas_any_label = n_filled > 0\n\nprint(f\"Total studies:                 {len(train)}\")\nprint(f\"Studies with labels:           {has_any_label.sum()}  ({has_any_label.mean():.1%})\")\nprint(f\"Studies with NO labels:        {(~has_any_label).sum()}  ({(~has_any_label).mean():.1%})\")\nprint()\nprint(\"Distribution of # non-null labels per study (should be all-or-nothing: 0 or 12):\")\nprint(n_filled.value_counts().sort_index())\n"},{"cell_type":"markdown","id":"2d412155","metadata":{},"source":"**58 of 4407 studies (1.3%) are labeled, and it's strictly all-or-nothing —\nevery labeled study has all 12 columns filled.** The other 4349 studies\nhave a report but no structured labels.\n\nThis is the defining fact of the competition: with 58 labeled studies you\ncannot train *or reliably validate* a 12-way classifier directly. Any\nviable strategy has to convert some combination of (a) the unlabeled\nreports and (b) image-report structure into extra training signal for the\nother 4349 studies.\n"},{"cell_type":"code","execution_count":5,"id":"7a9b79e4","metadata":{"execution":{"iopub.execute_input":"2026-08-12T22:46:33.638848Z","iopub.status.busy":"2026-08-12T22:46:33.638551Z","iopub.status.idle":"2026-08-12T22:46:33.966813Z","shell.execute_reply":"2026-08-12T22:46:33.96518Z"}},"outputs":[],"source":"labeled = train.loc[has_any_label]\n\npos_rate = labeled[LABEL_COLS].mean().sort_values(ascending=False)\npos_count = labeled[LABEL_COLS].sum().sort_values(ascending=False)\n\nfig, ax = plt.subplots(figsize=(9, 5))\npos_rate.plot(kind=\"barh\", ax=ax, color=\"#4C72B0\")\nax.set_xlabel(\"Positive rate among the 58 labeled studies\")\nax.set_title(\"Label prevalence (labeled subset only — likely NOT representative of\\nfull-population prevalence; this subset looks enriched for findings)\")\nax.invert_yaxis()\nfor i, (label, rate) in enumerate(pos_rate.items()):\n    ax.text(rate + 0.01, i, f\"{int(pos_count[label])}\", va=\"center\")\nplt.tight_layout()\nplt.show()\n\nprint(f\"Average # positive labels per labeled study: {labeled[LABEL_COLS].sum(axis=1).mean():.2f} / 12\")\n"},{"cell_type":"code","execution_count":6,"id":"b5ffec52","metadata":{"execution":{"iopub.execute_input":"2026-08-12T22:46:33.969589Z","iopub.status.busy":"2026-08-12T22:46:33.969367Z","iopub.status.idle":"2026-08-12T22:46:34.336518Z","shell.execute_reply":"2026-08-12T22:46:34.335185Z"}},"outputs":[],"source":"co_occurrence = labeled[LABEL_COLS].corr()\n\nfig, ax = plt.subplots(figsize=(8, 7))\nsns.heatmap(co_occurrence, annot=True, fmt=\".2f\", cmap=\"coolwarm\", center=0,\n            square=True, ax=ax, cbar_kws={\"shrink\": 0.8})\nax.set_title(\"Label co-occurrence (Pearson corr, n=58 labeled studies)\\nsmall-n — treat as a rough hint, not a stable estimate\")\nplt.tight_layout()\nplt.show()\n"},{"cell_type":"markdown","id":"56273ccc","metadata":{},"source":"## 3. Radiology report text\n\nReports are the main lever we have for turning the 4349 unlabeled studies\ninto usable training signal — but only for *training*: `test.csv` has no\n`Report` column, so whatever we do with text has to end up baked into an\nimage-only model or into silver labels used at training time.\n"},{"cell_type":"code","execution_count":7,"id":"19718ebf","metadata":{"execution":{"iopub.execute_input":"2026-08-12T22:46:34.340371Z","iopub.status.busy":"2026-08-12T22:46:34.339887Z","iopub.status.idle":"2026-08-12T22:46:34.502025Z","shell.execute_reply":"2026-08-12T22:46:34.500098Z"}},"outputs":[],"source":"report_len = train[\"Report\"].fillna(\"\").str.len()\n\nfig, ax = plt.subplots(figsize=(8, 4))\nsns.histplot(report_len, bins=50, ax=ax, color=\"#4C72B0\")\nax.set_xlabel(\"Report length (characters)\")\nax.set_title(\"Report length distribution (n=4407)\")\nplt.tight_layout()\nplt.show()\n\nprint(report_len.describe())\nprint()\nprint(\"Missing/empty reports:\", (report_len == 0).sum())\n"},{"cell_type":"code","execution_count":8,"id":"083f3e3a","metadata":{"execution":{"iopub.execute_input":"2026-08-12T22:46:34.508184Z","iopub.status.busy":"2026-08-12T22:46:34.507511Z","iopub.status.idle":"2026-08-12T22:46:34.730515Z","shell.execute_reply":"2026-08-12T22:46:34.729511Z"}},"outputs":[],"source":"# Language ID. Cached to report_langs.csv since langdetect over 4407 reports\n# takes a little while - reuse the cache if present.\nlang_cache_path = os.path.join(WORK_DIR, \"report_langs.csv\")\n\nif os.path.exists(lang_cache_path):\n    langs = pd.read_csv(lang_cache_path)\nelse:\n    from langdetect import DetectorFactory, detect\n    DetectorFactory.seed = 0\n\n    def safe_detect(text):\n        try:\n            return detect(text[:500])\n        except Exception:\n            return \"unknown\"\n\n    langs = train[[\"StudyInstanceUID\"]].copy()\n    langs[\"lang\"] = train[\"Report\"].apply(safe_detect)\n    langs.to_csv(lang_cache_path, index=False)\n\ntrain_lang = train.merge(langs, on=\"StudyInstanceUID\")\n\nfig, axes = plt.subplots(1, 2, figsize=(13, 4.5))\ntrain_lang[\"lang\"].value_counts().plot(kind=\"bar\", ax=axes[0], color=\"#4C72B0\")\naxes[0].set_title(f\"Report language — full corpus (n={len(train_lang)})\")\naxes[0].set_ylabel(\"# studies\")\n\ntrain_lang.loc[has_any_label.values, \"lang\"].value_counts().plot(kind=\"bar\", ax=axes[1], color=\"#DD8452\")\naxes[1].set_title(\"Report language — labeled subset (n=58)\")\nplt.tight_layout()\nplt.show()\n\nprint(f\"{train_lang['lang'].nunique()} distinct languages detected in the full corpus\")\n"},{"cell_type":"markdown","id":"a57fbbb2","metadata":{},"source":"Majority non-English (~61%), 9 languages detected. Any report-mining\napproach (rule-based or LLM) has to be multilingual by design, not\nEnglish-first with translation as an afterthought. The labeled subset\ncovers all major languages but only 2-3 examples for some (nl, de, bg,\nel) — thin for per-language precision/recall estimates on a labeler, but\nenough for a sanity spot-check.\n"},{"cell_type":"code","execution_count":9,"id":"25675fad","metadata":{"execution":{"iopub.execute_input":"2026-08-12T22:46:34.733329Z","iopub.status.busy":"2026-08-12T22:46:34.733082Z","iopub.status.idle":"2026-08-12T22:46:34.853419Z","shell.execute_reply":"2026-08-12T22:46:34.852638Z"}},"outputs":[],"source":"# De-identification placeholder tokens\nall_text = \" \".join(train[\"Report\"].tolist())\ntokens = re.findall(r\"\\[[A-Z ]+\\]\", all_text)\ntoken_counts = Counter(tokens).most_common(15)\n\nfig, ax = plt.subplots(figsize=(7, 4))\nlabels_, counts_ = zip(*token_counts)\nax.barh(labels_, counts_, color=\"#55A868\")\nax.invert_yaxis()\nax.set_title(\"De-identification placeholder tokens found in reports\")\nplt.tight_layout()\nplt.show()\n"},{"cell_type":"code","execution_count":10,"id":"e1ea51bd","metadata":{"execution":{"iopub.execute_input":"2026-08-12T22:46:34.856151Z","iopub.status.busy":"2026-08-12T22:46:34.855929Z","iopub.status.idle":"2026-08-12T22:46:34.889933Z","shell.execute_reply":"2026-08-12T22:46:34.888582Z"}},"outputs":[],"source":"dup_count = train[\"Report\"].duplicated().sum()\nprint(f\"Exact duplicate reports: {dup_count} / {len(train)}\")\nprint(f\"Unique reports: {train['Report'].nunique()}\")\nprint()\nprint(\"Most-repeated exact report strings (institutional 'normal' templates):\")\nfor text, cnt in train[\"Report\"].value_counts().head(5).items():\n    preview = text[:140].replace(\"\\n\", \" \")\n    print(f\"  [{cnt:>3}x] {preview}...\")\n"},{"cell_type":"markdown","id":"a1eaab50","metadata":{},"source":"A meaningful chunk of reports are exact-duplicate institutional templates\nfor normal studies (one Turkish \"everything normal\" template alone covers\n37 studies). That's good news for weak labeling: a cheap template/negative\nmatch per language+site can confidently knock out a large fraction of the\nnegative-everything cases, concentrating labeling effort (rule-based or\nLLM) on reports that actually describe findings.\n"},{"cell_type":"markdown","id":"2ecb3370","metadata":{},"source":"## 4. Series-level metadata\n\nEach study is a *bag* of MRI series (different planes/sequences). The\nmodel has to consume a variable, heterogeneous set of series per study —\nthere's no fixed protocol.\n"},{"cell_type":"code","execution_count":11,"id":"a5d55561","metadata":{"execution":{"iopub.execute_input":"2026-08-12T22:46:34.893164Z","iopub.status.busy":"2026-08-12T22:46:34.892916Z","iopub.status.idle":"2026-08-12T22:46:35.022928Z","shell.execute_reply":"2026-08-12T22:46:35.022073Z"}},"outputs":[],"source":"series_per_study = train_series.groupby(\"StudyInstanceUID\").size()\n\nfig, ax = plt.subplots(figsize=(7, 4))\nsns.histplot(series_per_study, bins=range(1, 16), ax=ax, color=\"#4C72B0\")\nax.set_xlabel(\"# series per study\")\nax.set_title(f\"Series per study (median={series_per_study.median():.0f}, \"\n             f\"min={series_per_study.min()}, max={series_per_study.max()})\")\nplt.tight_layout()\nplt.show()\n"},{"cell_type":"code","execution_count":12,"id":"8af8d155","metadata":{"execution":{"iopub.execute_input":"2026-08-12T22:46:35.02594Z","iopub.status.busy":"2026-08-12T22:46:35.025617Z","iopub.status.idle":"2026-08-12T22:46:35.247591Z","shell.execute_reply":"2026-08-12T22:46:35.245828Z"}},"outputs":[],"source":"fig, axes = plt.subplots(1, 2, figsize=(12, 4))\n\ntrain_series[\"Anatomical_Plane\"].value_counts().plot(kind=\"bar\", ax=axes[0], color=\"#4C72B0\")\naxes[0].set_title(\"Anatomical plane\")\n\ncrosstab = pd.crosstab(train_series[\"Fluid_Sensitive\"], train_series[\"Fat_Suppression\"])\nsns.heatmap(crosstab, annot=True, fmt=\"d\", cmap=\"Blues\", ax=axes[1])\naxes[1].set_title(\"Fluid_Sensitive vs Fat_Suppression\")\nplt.tight_layout()\nplt.show()\n\nprint(\"Fluid_Sensitive and Fat_Suppression are perfectly correlated in this data\")\nprint(\"(always both 0 or both 1) - effectively one binary flag, not two independent axes.\")\n"},{"cell_type":"code","execution_count":13,"id":"c24354c7","metadata":{"execution":{"iopub.execute_input":"2026-08-12T22:46:35.252157Z","iopub.status.busy":"2026-08-12T22:46:35.251804Z","iopub.status.idle":"2026-08-12T22:46:35.329026Z","shell.execute_reply":"2026-08-12T22:46:35.3276Z"}},"outputs":[],"source":"combo = train_series[\"Anatomical_Plane\"] + \"_\" + train_series[\"Fluid_Sensitive\"].astype(str) + train_series[\"Fat_Suppression\"].astype(str)\nfingerprint = combo.groupby(train_series[\"StudyInstanceUID\"]).apply(lambda x: tuple(sorted(x)))\n\nprint(f\"{fingerprint.nunique()} distinct per-study series-composition patterns across {len(fingerprint)} studies\")\nprint()\nprint(\"Top 10 most common patterns:\")\ntop = fingerprint.value_counts().head(10)\nfor pattern, cnt in top.items():\n    print(f\"  [{cnt:>4}x] {pattern}\")\n"},{"cell_type":"markdown","id":"72533df1","metadata":{},"source":"Modal protocol (~40% of studies): 5 series covering Axial/Coronal/Sagittal\nx {fluid-sensitive-fatsat, non-fs} - a standard 3-plane knee MRI protocol.\nBut 159 distinct fingerprints total, min 3 / max 14 series per study.\n**Model architecture needs to pool over an arbitrary bag of series, not\nassume fixed slots.**\n"},{"cell_type":"markdown","id":"656b0e71","metadata":{},"source":"## 5. DICOM header / pixel inspection (one sample study)\n\nPull the full DICOM series for one of the labeled studies, so the\nheader/pixel spot-check below can be cross-referenced against known\nground-truth labels. This is a spot-check on real acquisition metadata\nand pixel data, not a representative sample — treat it as \"what does one\nstudy look like\" rather than a population-level claim.\n"},{"cell_type":"code","execution_count":14,"id":"fb390455","metadata":{"execution":{"iopub.execute_input":"2026-08-12T22:46:35.331779Z","iopub.status.busy":"2026-08-12T22:46:35.331498Z","iopub.status.idle":"2026-08-12T22:46:35.878126Z","shell.execute_reply":"2026-08-12T22:46:35.876998Z"}},"outputs":[],"source":"import glob\n\nimport pydicom\n\nsample_study_uid = labeled[\"StudyInstanceUID\"].iloc[0]\nsample_files = sorted(glob.glob(os.path.join(DATA_DIR, \"train_series\", sample_study_uid, \"*\", \"*.dcm\")))\nsample_labels = labeled.loc[labeled[\"StudyInstanceUID\"] == sample_study_uid, LABEL_COLS].iloc[0]\nprint(f\"study {sample_study_uid}: {len(sample_files)} DICOM files\")\nprint(\"ground-truth labels:\", sample_labels[sample_labels == 1].index.tolist())\n\nseries = defaultdict(list)\nfor fpath in sample_files:\n    ds = pydicom.dcmread(fpath, stop_before_pixels=True)\n    series[ds.SeriesInstanceUID].append((fpath, ds))\n\nprint(f\"{len(series)} series\")\n"},{"cell_type":"code","execution_count":15,"id":"45335779","metadata":{"execution":{"iopub.execute_input":"2026-08-12T22:46:35.880779Z","iopub.status.busy":"2026-08-12T22:46:35.880481Z","iopub.status.idle":"2026-08-12T22:46:35.893831Z","shell.execute_reply":"2026-08-12T22:46:35.89305Z"}},"outputs":[],"source":"rows = []\nfor sid, items in series.items():\n    _, ds = items[0]\n    tsyn = ds.file_meta.TransferSyntaxUID if hasattr(ds, \"file_meta\") else None\n    rows.append({\n        \"series_uid_tail\": sid[-12:],\n        \"n_slices\": len(items),\n        \"SeriesDescription\": getattr(ds, \"SeriesDescription\", None),\n        \"Rows\": getattr(ds, \"Rows\", None),\n        \"Columns\": getattr(ds, \"Columns\", None),\n        \"PixelSpacing_mm\": tuple(round(float(x), 3) for x in ds.PixelSpacing) if hasattr(ds, \"PixelSpacing\") else None,\n        \"SliceThickness_mm\": getattr(ds, \"SliceThickness\", None),\n        \"TransferSyntax\": tsyn.name if tsyn else None,\n        \"MagneticField_T\": getattr(ds, \"MagneticFieldStrength\", None),\n        \"Manufacturer\": getattr(ds, \"Manufacturer\", None),\n    })\n\nheader_df = pd.DataFrame(rows)\nheader_df\n"},{"cell_type":"markdown","id":"434c7856","metadata":{},"source":"Even *within this one study*, image matrix size varies across series (see\nthe table above for the actual sizes seen in this run) - resizing/\nnormalization has to happen per-series, no dataset-wide fixed input size\nassumption. The dataset description says transfer syntax varies too\n(uncompressed Explicit VR LE, JPEG Lossless, JPEG 2000, Implicit VR LE),\nwhich is why `pydicom` + `pylibjpeg`/`gdcm` plugins are installed above -\nneeded to decode whichever ones this particular study happens to use.\n"},{"cell_type":"code","execution_count":16,"id":"78008bfd","metadata":{"execution":{"iopub.execute_input":"2026-08-12T22:46:35.896672Z","iopub.status.busy":"2026-08-12T22:46:35.896449Z","iopub.status.idle":"2026-08-12T22:46:35.903335Z","shell.execute_reply":"2026-08-12T22:46:35.902206Z"}},"outputs":[],"source":"# Patient-level tags (present in DICOM even though PatientSex isn't a train.csv column)\n_, ds0 = next(iter(series.values()))[0], pydicom.dcmread(next(iter(series.values()))[0][0])\nprint(\"Patient's Sex:\", getattr(ds0, \"PatientSex\", None))\nprint(\"Patient ID (pseudonymized):\", getattr(ds0, \"PatientID\", None))\nprint(\"Body Part Examined:\", getattr(ds0, \"BodyPartExamined\", None))\nprint(\"Laterality:\", repr(getattr(ds0, \"Laterality\", None)))\n"},{"cell_type":"code","execution_count":17,"id":"2c31960d","metadata":{"execution":{"iopub.execute_input":"2026-08-12T22:46:35.906426Z","iopub.status.busy":"2026-08-12T22:46:35.905979Z","iopub.status.idle":"2026-08-12T22:46:36.516895Z","shell.execute_reply":"2026-08-12T22:46:36.514603Z"}},"outputs":[],"source":"fig, axes = plt.subplots(1, len(series), figsize=(4 * len(series), 4.5))\nif len(series) == 1:\n    axes = [axes]\n\nfor ax, (sid, items) in zip(axes, series.items()):\n    fpath, ds = items[len(items) // 2]  # a middle slice\n    ds_full = pydicom.dcmread(fpath)\n    arr = ds_full.pixel_array.astype(float)\n    ax.imshow(arr, cmap=\"gray\")\n    ax.set_title(getattr(ds_full, \"SeriesDescription\", sid[-8:]), fontsize=9)\n    ax.axis(\"off\")\n\nplt.suptitle(\"One mid-stack slice from each series in the sample study\")\nplt.tight_layout()\nplt.show()\n"},{"cell_type":"markdown","id":"1533a102","metadata":{},"source":"## 6. Summary — implications for strategy\n\n- **58/4407 (1.3%) labeled studies, all-or-nothing.** Not enough to train\n  or validate a 12-way classifier directly. The competition is really\n  about how well you can mine the other 4349 reports (and/or the\n  image-report pairing itself) for training signal.\n- **Reports are training-time only** — `test.csv` has no `Report` column,\n  so text can inform silver labels or contrastive pretraining, but the\n  deployed model has to be image-only.\n- **9 languages, majority non-English.** Report mining must be\n  multilingual from the start.\n- **Reports are heavily templated for normal cases** — cheap\n  template-matching can likely resolve a large chunk of the\n  negative-everything studies, concentrating harder labeling effort\n  (rule-based or LLM) on the reports that describe actual findings.\n- **The labeled-58's class balance looks enriched for abnormal findings**\n  (e.g. 60% Effusion, 31% Fracture) relative to plausible general\n  prevalence — don't treat it as a population prior.\n- **Series composition is heterogeneous** (159 distinct patterns, 3-14\n  series/study) and **image resolution varies even within a study** — the\n  model needs to pool over a variable bag of series/slices, not assume\n  fixed input slots or a fixed image size.\n- **Training will need GPU compute** — this EDA doesn't, but the\n  downstream image model (and any contrastive image-report pretraining)\n  will, so budget for it up front.\n"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.13.14"}},"nbformat":4,"nbformat_minor":5}