{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceType":"competition","sourceId":154281}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# RSNA Knee Abnormality Detection — Exploratory Data Analysis\n### Objective\n\nThe goal of this competition is to predict 12 knee abnormalities from MRI studies. The training data also include multilingual radiology reports, which provide supervision for studies without structured target labels.\nBefore building a model, we need to understand:\n\n- what each row in the metadata represents\n- how patients, studies, series and images are related\n- how the abnormality labels are distributed\n- which imaging modalities and views are available\n- whether the image data contain quality problems\n- whether the training and test data have different distributions\n- whether there are duplicate or related studies that could cause data leakage\n- how the radiology reports can provide weak supervision for studies without structured labels\n\nThis notebook focuses on understanding the dataset. Modelling and validation decisions will be based on the findings from this analysis.\n\n> Important: medical images from the same patient or study must not be split across training and validation sets. We will investigate these relationships carefully before modelling.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"# Setup","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 1. Imports and reproducibility\n# ============================================================\n\nfrom pathlib import Path\nfrom collections import Counter\nimport os\nimport warnings\nimport importlib.metadata as metadata\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nwarnings.filterwarnings(\"ignore\", category=FutureWarning)\n\nRANDOM_STATE = 42\nDATA_ROOT = Path(\"/kaggle/input\")\n\npd.set_option(\"display.max_columns\", 100)\npd.set_option(\"display.max_rows\", 100)\npd.set_option(\"display.float_format\", lambda value: f\"{value:,.4f}\")\n\nsns.set_theme(\n    style=\"whitegrid\",\n    context=\"notebook\",\n    palette=\"deep\"\n)\n\nprint(f\"NumPy:      {metadata.version('numpy')}\")\nprint(f\"Pandas:     {metadata.version('pandas')}\")\nprint(f\"Matplotlib: {metadata.version('matplotlib')}\")\nprint(f\"Seaborn:    {metadata.version('seaborn')}\")\nprint(f\"Random seed: {RANDOM_STATE}\")\nprint(f\"Input root:  {DATA_ROOT}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:32.270825Z","iopub.execute_input":"2026-09-19T13:55:32.271193Z","iopub.status.idle":"2026-09-19T13:55:32.667573Z","shell.execute_reply.started":"2026-09-19T13:55:32.271158Z","shell.execute_reply":"2026-09-19T13:55:32.666759Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# 2. Initial input inventory\n# ============================================================\n\nif not DATA_ROOT.exists():\n    raise FileNotFoundError(f\"Could not find {DATA_ROOT}\")\n\ndef describe_entry(path):\n    \"\"\"Create a compact summary of a file or directory.\"\"\"\n    if path.is_file():\n        return {\n            \"name\": path.name,\n            \"type\": \"file\",\n            \"size_mb\": path.stat().st_size / 1024**2,\n            \"children\": np.nan,\n        }\n\n    children = list(path.iterdir())\n\n    return {\n        \"name\": path.name,\n        \"type\": \"directory\",\n        \"size_mb\": np.nan,\n        \"children\": len(children),\n    }\n\ninput_summary = pd.DataFrame(\n    [describe_entry(path) for path in sorted(DATA_ROOT.iterdir())]\n)\n\ndisplay(input_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:32.669163Z","iopub.execute_input":"2026-09-19T13:55:32.669425Z","iopub.status.idle":"2026-09-19T13:55:32.705534Z","shell.execute_reply.started":"2026-09-19T13:55:32.6694Z","shell.execute_reply":"2026-09-19T13:55:32.704374Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Understanding the competition file structure\n\nWe first inspect the directory structure and file types before loading any data.\n\nThis is important because medical-imaging competitions often contain several levels of organisation, such as:\n\n- patients\n- studies\n- series\n- image files\n- tabular metadata or labels","metadata":{"_kg_hide-input":true}},{"cell_type":"code","source":"# ============================================================\n# 3. Inspect competition directory structure\n# ============================================================\n\nCOMPETITIONS_ROOT = DATA_ROOT / \"competitions\"\n\ndef preview_directory(path, limit=50):\n    \"\"\"\n    Display only the first few entries in a directory.\n    This avoids scanning the entire image dataset.\n    \"\"\"\n    rows = []\n\n    with os.scandir(path) as entries:\n        for index, entry in enumerate(entries):\n            if index >= limit:\n                break\n\n            entry_path = Path(entry.path)\n            is_directory = entry.is_dir(follow_symlinks=False)\n\n            rows.append({\n                \"name\": entry.name,\n                \"type\": \"directory\" if is_directory else \"file\",\n                \"size_mb\": (\n                    np.nan\n                    if is_directory\n                    else entry.stat(follow_symlinks=False).st_size / 1024**2\n                ),\n            })\n\n    return pd.DataFrame(rows)\n\n\nprint(f\"Competition directory: {COMPETITIONS_ROOT}\")\n\ncompetition_directories = sorted(\n    path for path in COMPETITIONS_ROOT.iterdir()\n    if path.is_dir()\n)\n\nprint(\"\\nCompetition folders:\")\nfor path in competition_directories:\n    print(f\"  - {path.name}\")\n\nprint(\"\\nPreview of competition directory:\")\ndisplay(preview_directory(COMPETITIONS_ROOT))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:32.706624Z","iopub.execute_input":"2026-09-19T13:55:32.70692Z","iopub.status.idle":"2026-09-19T13:55:32.724837Z","shell.execute_reply.started":"2026-09-19T13:55:32.706892Z","shell.execute_reply":"2026-09-19T13:55:32.723764Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# 4. Preview the competition files\n# ============================================================\n\nif len(competition_directories) != 1:\n    print(\"More than one competition folder was found.\")\n    print(\"Please inspect the table above before continuing.\")\nelse:\n    COMPETITION_ROOT = competition_directories[0]\n\n    print(f\"Selected competition folder: {COMPETITION_ROOT}\")\n    display(preview_directory(COMPETITION_ROOT, limit=100))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:32.72692Z","iopub.execute_input":"2026-09-19T13:55:32.72727Z","iopub.status.idle":"2026-09-19T13:55:32.751224Z","shell.execute_reply.started":"2026-09-19T13:55:32.72723Z","shell.execute_reply":"2026-09-19T13:55:32.750537Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Loading the tabular data\n\nThe competition provides several CSV files describing the studies and imaging series.\n\nWe will initially work with the tabular files only. The image directories are much larger, so we will inspect them separately after understanding the metadata.","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 5. Load the tabular competition files\n# ============================================================\n\nTRAIN_PATH = COMPETITION_ROOT / \"train.csv\"\nTEST_PATH = COMPETITION_ROOT / \"test.csv\"\nTRAIN_SERIES_PATH = COMPETITION_ROOT / \"train_series.csv\"\nTEST_SERIES_PATH = COMPETITION_ROOT / \"test_series.csv\"\nSAMPLE_SUBMISSION_PATH = COMPETITION_ROOT / \"sample_submission.csv\"\n\ntable_paths = {\n    \"train\": TRAIN_PATH,\n    \"test\": TEST_PATH,\n    \"train_series\": TRAIN_SERIES_PATH,\n    \"test_series\": TEST_SERIES_PATH,\n    \"sample_submission\": SAMPLE_SUBMISSION_PATH,\n}\n\nmissing_files = [\n    str(path)\n    for path in table_paths.values()\n    if not path.exists()\n]\n\nif missing_files:\n    raise FileNotFoundError(\n        \"The following expected files were not found:\\n\"\n        + \"\\n\".join(missing_files)\n    )\n\ntrain = pd.read_csv(TRAIN_PATH)\ntest = pd.read_csv(TEST_PATH)\ntrain_series = pd.read_csv(TRAIN_SERIES_PATH)\ntest_series = pd.read_csv(TEST_SERIES_PATH)\nsample_submission = pd.read_csv(SAMPLE_SUBMISSION_PATH)\n\ntables = {\n    \"train\": train,\n    \"test\": test,\n    \"train_series\": train_series,\n    \"test_series\": test_series,\n    \"sample_submission\": sample_submission,\n}\n\nprint(\"All tabular files loaded successfully.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:32.752111Z","iopub.execute_input":"2026-09-19T13:55:32.752346Z","iopub.status.idle":"2026-09-19T13:55:33.157889Z","shell.execute_reply.started":"2026-09-19T13:55:32.752324Z","shell.execute_reply":"2026-09-19T13:55:33.157041Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Initial table overview\n\nThe first question is how many rows and columns each table contains.\n\nThe number of rows does not necessarily equal the number of patients or studies. A single study may contain multiple imaging series, and a single series may contain multiple images.","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 6. Compare table sizes\n# ============================================================\n\ntable_overview = pd.DataFrame([\n    {\n        \"table\": name,\n        \"rows\": dataframe.shape[0],\n        \"columns\": dataframe.shape[1],\n        \"memory_mb\": dataframe.memory_usage(deep=True).sum() / 1024**2,\n    }\n    for name, dataframe in tables.items()\n])\n\ndisplay(table_overview)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:33.158912Z","iopub.execute_input":"2026-09-19T13:55:33.159328Z","iopub.status.idle":"2026-09-19T13:55:33.194567Z","shell.execute_reply.started":"2026-09-19T13:55:33.159291Z","shell.execute_reply":"2026-09-19T13:55:33.193728Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Column and data-quality overview\n\nWe now inspect the columns, data types, number of unique values and missing values.\n\nThis will help us identify:\n\n- identifier columns;\n- target columns;\n- categorical metadata;\n- numerical metadata;\n- columns that are constant or nearly constant; and\n- possible data-quality problems.","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 7. Inspect columns and missing values\n# ============================================================\n\nschema_rows = []\n\nfor table_name, dataframe in tables.items():\n    for column in dataframe.columns:\n        schema_rows.append({\n            \"table\": table_name,\n            \"column\": column,\n            \"dtype\": str(dataframe[column].dtype),\n            \"unique_values\": dataframe[column].nunique(dropna=False),\n            \"missing_values\": dataframe[column].isna().sum(),\n            \"missing_percent\": (\n                dataframe[column].isna().mean() * 100\n            ),\n        })\n\nschema = (\n    pd.DataFrame(schema_rows)\n    .sort_values([\"table\", \"column\"])\n    .reset_index(drop=True)\n)\n\ndisplay(schema)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:33.195537Z","iopub.execute_input":"2026-09-19T13:55:33.195789Z","iopub.status.idle":"2026-09-19T13:55:33.279988Z","shell.execute_reply.started":"2026-09-19T13:55:33.195766Z","shell.execute_reply":"2026-09-19T13:55:33.279385Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# 8. Preview each table\n# ============================================================\n\nfor table_name, dataframe in tables.items():\n    print(f\"\\n{table_name.upper()}\")\n    display(dataframe.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:33.280834Z","iopub.execute_input":"2026-09-19T13:55:33.281061Z","iopub.status.idle":"2026-09-19T13:55:33.325575Z","shell.execute_reply.started":"2026-09-19T13:55:33.28104Z","shell.execute_reply":"2026-09-19T13:55:33.324908Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Labels and target coverage\n\nThe training table contains 4,407 studies, but the abnormality labels are not available for every study.\n\nThis distinction is important:\n\n- a value of `0` means the abnormality was not identified;\n- a value of `1` means the abnormality was identified; and\n- a missing value means that the study is unlabelled.\n\nTherefore, missing labels must not be interpreted as normal studies.","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 9. Identify target columns and label coverage\n# ============================================================\n\nID_COLUMN = \"StudyInstanceUID\"\n\nTARGET_COLUMNS = [\n    column\n    for column in sample_submission.columns\n    if column != ID_COLUMN\n]\n\nprint(\"Target columns:\")\nprint(TARGET_COLUMNS)\n\nlabels_per_row = train[TARGET_COLUMNS].notna().sum(axis=1)\n\nlabel_status = pd.Series(\n    np.select(\n        [\n            labels_per_row == 0,\n            labels_per_row == len(TARGET_COLUMNS),\n        ],\n        [\n            \"Unlabelled\",\n            \"Fully labelled\",\n        ],\n        default=\"Partially labelled\",\n    ),\n    name=\"label_status\",\n)\n\nlabel_status_summary = (\n    label_status\n    .value_counts()\n    .rename_axis(\"label_status\")\n    .reset_index(name=\"number_of_studies\")\n)\n\nlabel_status_summary[\"percentage\"] = (\n    label_status_summary[\"number_of_studies\"]\n    / len(train)\n    * 100\n)\n\ndisplay(label_status_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:33.326504Z","iopub.execute_input":"2026-09-19T13:55:33.326891Z","iopub.status.idle":"2026-09-19T13:55:33.352257Z","shell.execute_reply.started":"2026-09-19T13:55:33.326856Z","shell.execute_reply":"2026-09-19T13:55:33.351571Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Key findings\n\nOnly **58 of the 4,407 training studies (1.32%)** contain expert-structured labels for the 12 target abnormalities. The remaining **4,349 studies (98.68%)** contain MRI data and a radiology report but no structured target values.\n\nThere are no partially labelled studies: each study either has all 12 structured targets or none of them.\n\nThis is an intentional feature of the competition rather than a missing-data error. The 58 labelled studies form a small **gold-label subset**, while the reports provide weak supervision for the much larger unlabelled subset.\n\nA successful solution will therefore need to derive training signals from the reports—for example through multilingual NLP, clinical rules or an LLM-based label extractor—and then train an image model using those report-derived labels.\n\nMissing targets must never be converted directly to zero: `NaN` means unlabelled, whereas `0` means an abnormality was assessed as absent.","metadata":{}},{"cell_type":"markdown","source":"## Label distribution\n\nWe now examine the abnormality labels among the labelled studies only.\n\nThe prevalence figures below describe the labelled subset. They should not be interpreted as prevalence estimates for all 4,407 training studies because most studies do not have labels.","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 10. Target distribution among labelled studies\n# ============================================================\n\nfully_labelled_mask = labels_per_row == len(TARGET_COLUMNS)\nlabelled_train = train.loc[fully_labelled_mask].copy()\n\ntarget_summary_rows = []\n\nfor target in TARGET_COLUMNS:\n    values = labelled_train[target].dropna()\n\n    target_summary_rows.append({\n        \"target\": target,\n        \"labelled_studies\": len(values),\n        \"positive_cases\": int((values == 1).sum()),\n        \"negative_cases\": int((values == 0).sum()),\n        \"positive_percentage\": (values == 1).mean() * 100,\n        \"unique_values\": sorted(values.unique().tolist()),\n    })\n\ntarget_summary = (\n    pd.DataFrame(target_summary_rows)\n    .sort_values(\"positive_percentage\", ascending=False)\n    .reset_index(drop=True)\n)\n\ndisplay(target_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:33.355305Z","iopub.execute_input":"2026-09-19T13:55:33.355619Z","iopub.status.idle":"2026-09-19T13:55:33.380406Z","shell.execute_reply.started":"2026-09-19T13:55:33.355585Z","shell.execute_reply":"2026-09-19T13:55:33.379567Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# 11. Validate target values\n# ============================================================\n\nobserved_target_values = set(\n    labelled_train[TARGET_COLUMNS]\n    .to_numpy()\n    .ravel()\n)\n\nexpected_target_values = {0.0, 1.0}\nunexpected_target_values = observed_target_values - expected_target_values\n\nassert not unexpected_target_values, (\n    f\"Unexpected target values found: {unexpected_target_values}\"\n)\n\nprint(\"Target validation passed.\")\nprint(\"All labelled target values are binary: 0 or 1.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:33.381489Z","iopub.execute_input":"2026-09-19T13:55:33.381844Z","iopub.status.idle":"2026-09-19T13:55:33.388188Z","shell.execute_reply.started":"2026-09-19T13:55:33.381808Z","shell.execute_reply":"2026-09-19T13:55:33.387387Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Key findings\n\nAll observed targets are binary, with values of either `0` or `1`. This confirms that each abnormality is formulated as a separate binary classification target.\n\nThe positive rates vary considerably across the 12 abnormalities:\n\n- **Effusion** is the most frequent, appearing in 35 of 58 labelled studies (60.3%).\n- **Synovitis**, **Medial Meniscus injury** and **ACL injury** are also relatively common.\n- **MCL injury** is the least frequent, appearing in only 9 studies (15.5%).\n- **Lateral OA** and **Baker's cyst** are similarly uncommon.\n\nAcross all targets, there are 240 positive labels among 58 studies, equivalent to approximately **4.14 positive abnormalities per study**. The targets are therefore not mutually exclusive: an individual study can contain several abnormalities.\n\nBecause the labelled sample is very small, these percentages are uncertain and should not be interpreted as estimates of prevalence in the full dataset or the general population. The rare targets will also make validation particularly challenging. For example, a five-fold split would contain only around one or two positive MCL cases in each validation fold.","metadata":{"_kg_hide-input":false}},{"cell_type":"markdown","source":"# Multi-label target analysis\n\nEach study can contain more than one abnormality. We will examine both the balance of the individual targets and the relationships between them.\n\nAll results in this section refer only to the 58 gold-labelled studies.","metadata":{}},{"cell_type":"markdown","source":"## Individual target balance\n\nThe chart below compares the positive rate of each target within the gold-labelled subset.\n\nBecause only 58 studies are available, these rates are descriptive rather than reliable population prevalence estimates.","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 12. Visualise individual target prevalence\n# ============================================================\n\nprevalence_plot = (\n    target_summary\n    .sort_values(\"positive_percentage\")\n    .reset_index(drop=True)\n)\n\nfig, ax = plt.subplots(figsize=(10, 7))\n\nbars = ax.barh(\n    prevalence_plot[\"target\"],\n    prevalence_plot[\"positive_percentage\"],\n    color=sns.color_palette(\"deep\")[0],\n)\n\nfor bar, positive_cases, percentage in zip(\n    bars,\n    prevalence_plot[\"positive_cases\"],\n    prevalence_plot[\"positive_percentage\"],\n):\n    ax.text(\n        bar.get_width() + 1,\n        bar.get_y() + bar.get_height() / 2,\n        f\"{positive_cases}/58 ({percentage:.1f}%)\",\n        va=\"center\",\n    )\n\nax.set_xlim(0, 70)\nax.set_xlabel(\"Positive studies (%)\")\nax.set_ylabel(\"\")\nax.set_title(\"Target prevalence in the gold-labelled subset\")\nax.grid(axis=\"x\", alpha=0.25)\nax.grid(axis=\"y\", visible=False)\n\nsns.despine()\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:33.389253Z","iopub.execute_input":"2026-09-19T13:55:33.389535Z","iopub.status.idle":"2026-09-19T13:55:33.724743Z","shell.execute_reply.started":"2026-09-19T13:55:33.389509Z","shell.execute_reply":"2026-09-19T13:55:33.723816Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# 13. Count positive labels per study\n# ============================================================\n\ngold_labels = (\n    labelled_train[TARGET_COLUMNS]\n    .astype(int)\n    .copy()\n)\n\npositive_labels_per_study = gold_labels.sum(axis=1)\n\npositive_label_summary = pd.DataFrame({\n    \"metric\": [\n        \"Gold-labelled studies\",\n        \"Total positive labels\",\n        \"Mean positives per study\",\n        \"Median positives per study\",\n        \"Minimum positives in one study\",\n        \"Maximum positives in one study\",\n        \"Studies with no positive targets\",\n        \"Studies with at least one positive target\",\n    ],\n    \"value\": [\n        len(gold_labels),\n        int(positive_labels_per_study.sum()),\n        positive_labels_per_study.mean(),\n        positive_labels_per_study.median(),\n        int(positive_labels_per_study.min()),\n        int(positive_labels_per_study.max()),\n        int((positive_labels_per_study == 0).sum()),\n        int((positive_labels_per_study > 0).sum()),\n    ],\n})\n\ndisplay(positive_label_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:33.725977Z","iopub.execute_input":"2026-09-19T13:55:33.726315Z","iopub.status.idle":"2026-09-19T13:55:33.738999Z","shell.execute_reply.started":"2026-09-19T13:55:33.726263Z","shell.execute_reply":"2026-09-19T13:55:33.738168Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# 14. Visualise positive labels per study\n# ============================================================\n\nlabel_count_distribution = (\n    positive_labels_per_study\n    .value_counts()\n    .sort_index()\n    .reindex(\n        range(\n            int(positive_labels_per_study.min()),\n            int(positive_labels_per_study.max()) + 1,\n        ),\n        fill_value=0,\n    )\n    .rename_axis(\"positive_labels\")\n    .reset_index(name=\"number_of_studies\")\n)\n\nfig, ax = plt.subplots(figsize=(10, 5))\n\nbars = ax.bar(\n    label_count_distribution[\"positive_labels\"],\n    label_count_distribution[\"number_of_studies\"],\n    color=sns.color_palette(\"deep\")[0],\n)\n\nax.bar_label(bars, padding=3)\n\nax.set_xlabel(\"Number of positive abnormalities\")\nax.set_ylabel(\"Number of studies\")\nax.set_title(\"Number of abnormalities per gold-labelled study\")\nax.set_xticks(label_count_distribution[\"positive_labels\"])\nax.grid(axis=\"y\", alpha=0.25)\nax.grid(axis=\"x\", visible=False)\n\nsns.despine()\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:33.740243Z","iopub.execute_input":"2026-09-19T13:55:33.740974Z","iopub.status.idle":"2026-09-19T13:55:33.955799Z","shell.execute_reply.started":"2026-09-19T13:55:33.740945Z","shell.execute_reply":"2026-09-19T13:55:33.954838Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Key findings\n\nThe gold-labelled studies contain an average of **4.14 positive abnormalities per study**, with a median of four and a range from one to nine.\n\nEvery one of the 58 studies has at least one positive target. There are no completely negative or normal studies in the gold-labelled subset.\n\nThis suggests that the gold subset is enriched for abnormal cases. It should therefore be used cautiously for estimating prevalence or validating performance on normal studies. Although every individual target has both positive and negative examples, the subset cannot directly measure how well a model distinguishes a completely normal knee from one containing any abnormality.\n\nThe high number of findings per study also confirms that this is a genuinely multi-label problem. Models must be able to identify several co-existing abnormalities rather than select a single diagnosis.","metadata":{}},{"cell_type":"markdown","source":"## Target co-occurrence\n\nSome knee abnormalities may frequently appear together. The following matrix counts the number of gold-labelled studies in which each pair of targets is positive.\n\nThe diagonal represents the total number of positive cases for each individual target.","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 15. Calculate target co-occurrence\n# ============================================================\n\ncooccurrence_matrix = gold_labels.T.dot(gold_labels)\n\nfig, ax = plt.subplots(figsize=(11, 9))\n\nsns.heatmap(\n    cooccurrence_matrix,\n    annot=True,\n    fmt=\"d\",\n    cmap=\"Blues\",\n    square=True,\n    linewidths=0.5,\n    cbar_kws={\"label\": \"Number of studies\"},\n    ax=ax,\n)\n\nax.set_title(\"Co-occurrence of abnormalities in gold-labelled studies\")\nax.set_xlabel(\"\")\nax.set_ylabel(\"\")\n\nplt.xticks(rotation=45, ha=\"right\")\nplt.yticks(rotation=0)\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:33.956837Z","iopub.execute_input":"2026-09-19T13:55:33.957164Z","iopub.status.idle":"2026-09-19T13:55:34.475745Z","shell.execute_reply.started":"2026-09-19T13:55:33.957119Z","shell.execute_reply":"2026-09-19T13:55:34.474811Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# 16. Summarise the most frequent target pairs\n# ============================================================\n\npair_rows = []\n\nfor first_index, first_target in enumerate(TARGET_COLUMNS):\n    for second_target in TARGET_COLUMNS[first_index + 1:]:\n        first_positive = gold_labels[first_target] == 1\n        second_positive = gold_labels[second_target] == 1\n\n        both_positive = int(\n            (first_positive & second_positive).sum()\n        )\n\n        either_positive = int(\n            (first_positive | second_positive).sum()\n        )\n\n        pair_rows.append({\n            \"first_target\": first_target,\n            \"second_target\": second_target,\n            \"both_positive\": both_positive,\n            \"percentage_of_gold_studies\": (\n                both_positive / len(gold_labels) * 100\n            ),\n            \"jaccard_similarity\": (\n                both_positive / either_positive\n                if either_positive > 0\n                else np.nan\n            ),\n        })\n\ntarget_pairs = (\n    pd.DataFrame(pair_rows)\n    .sort_values(\n        [\"both_positive\", \"jaccard_similarity\"],\n        ascending=False,\n    )\n    .reset_index(drop=True)\n)\n\ndisplay(target_pairs.head(15))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:34.476824Z","iopub.execute_input":"2026-09-19T13:55:34.477146Z","iopub.status.idle":"2026-09-19T13:55:34.513984Z","shell.execute_reply.started":"2026-09-19T13:55:34.477094Z","shell.execute_reply":"2026-09-19T13:55:34.513208Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Key findings\n\nThe strongest observed overlap is between **Effusion and Synovitis**. Both are present in 22 studies, representing 37.9% of the gold subset. Their Jaccard similarity of 0.55 means that 55% of studies containing either finding contain both.\n\nOther frequent combinations include:\n\n- Medial Meniscus injury with Effusion: 19 studies;\n- ACL injury with Effusion: 18 studies;\n- Lateral Meniscus injury with Effusion: 16 studies;\n- Medial Meniscus injury with Synovitis: 15 studies; and\n- PF OA with Synovitis: 14 studies.\n\nSome less frequent targets also show notable proportional overlap. For example, ACL injury and Contusion have a Jaccard similarity of 0.43, while Medial Meniscus injury and Medial OA have a similarity of 0.41.\n\nRaw co-occurrence counts are partly influenced by individual target prevalence. Effusion appears in many of the leading pairs because it is the most common target. The Jaccard similarity provides additional context by measuring overlap relative to how often either target occurs.\n\nThese relationships are exploratory. With only 58 studies, they should not be interpreted as precise or causal clinical associations.","metadata":{}},{"cell_type":"markdown","source":"# Radiology report analysis\n\nOnly 58 studies contain expert-structured target labels. For the remaining 4,349 studies, the accompanying radiology reports are the main source of supervision.\n\nBefore attempting to extract weak labels, we need to understand:\n\n- report completeness and uniqueness;\n- report length and structure;\n- differences between gold-labelled and report-only studies;\n- repeated or templated reports;\n- the languages represented in the dataset; and\n- how consistently the 12 target findings are described.\n\nWe begin with general text characteristics before examining languages and condition-specific terminology.","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 17. Prepare the radiology reports\n# ============================================================\n\nreport_data = train[\n    [ID_COLUMN, \"Report\"]\n].copy()\n\nreport_data[\"label_status\"] = np.where(\n    fully_labelled_mask,\n    \"Gold-labelled\",\n    \"Report-only\",\n)\n\nreport_data[\"normalized_report\"] = (\n    report_data[\"Report\"]\n    .fillna(\"\")\n    .str.replace(r\"\\s+\", \" \", regex=True)\n    .str.strip()\n    .str.lower()\n)\n\nreport_data[\"report_characters\"] = (\n    report_data[\"Report\"]\n    .fillna(\"\")\n    .str.len()\n)\n\nreport_data[\"report_words\"] = (\n    report_data[\"Report\"]\n    .fillna(\"\")\n    .str.split()\n    .str.len()\n)\n\nreport_data[\"report_lines\"] = (\n    report_data[\"Report\"]\n    .fillna(\"\")\n    .str.count(\"\\n\")\n    .add(1)\n)\n\nreport_data[\"exact_duplicate\"] = (\n    report_data[\"Report\"]\n    .duplicated(keep=False)\n)\n\nreport_data[\"normalized_duplicate\"] = (\n    report_data[\"normalized_report\"]\n    .duplicated(keep=False)\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:34.515016Z","iopub.execute_input":"2026-09-19T13:55:34.515335Z","iopub.status.idle":"2026-09-19T13:55:34.962189Z","shell.execute_reply.started":"2026-09-19T13:55:34.515298Z","shell.execute_reply":"2026-09-19T13:55:34.961379Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Report coverage and uniqueness\n\nMissing or empty reports would provide no weak-supervision signal. Repeated reports may represent reporting templates, genuine duplicate text or closely related examinations.","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 18. Summarise report coverage\n# ============================================================\n\nreport_coverage = pd.DataFrame({\n    \"metric\": [\n        \"Training studies\",\n        \"Missing reports\",\n        \"Empty reports\",\n        \"Unique exact reports\",\n        \"Unique normalized reports\",\n        \"Rows belonging to an exact duplicate group\",\n        \"Rows belonging to a normalized duplicate group\",\n    ],\n    \"value\": [\n        len(report_data),\n        int(report_data[\"Report\"].isna().sum()),\n        int(report_data[\"normalized_report\"].eq(\"\").sum()),\n        int(report_data[\"Report\"].nunique(dropna=True)),\n        int(report_data[\"normalized_report\"].nunique()),\n        int(report_data[\"exact_duplicate\"].sum()),\n        int(report_data[\"normalized_duplicate\"].sum()),\n    ],\n})\n\ndisplay(report_coverage)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:34.963179Z","iopub.execute_input":"2026-09-19T13:55:34.963504Z","iopub.status.idle":"2026-09-19T13:55:35.011092Z","shell.execute_reply.started":"2026-09-19T13:55:34.963477Z","shell.execute_reply":"2026-09-19T13:55:35.010199Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Report availability and uniqueness findings\n\nAll 4,407 training studies have a non-empty report, meaning every report-only study has some potential source of weak supervision.\n\nMost reports are unique. However, 177 rows belong to an exact duplicate group, increasing to 204 after differences in whitespace and capitalisation are removed. Report duplication is therefore present but affects only a small proportion of the dataset.","metadata":{}},{"cell_type":"markdown","source":"## Report length\n\nReport length provides a simple indication of how much textual information is available for weak-label extraction.\n\nWe compare gold-labelled and report-only studies to check whether the small gold subset contains systematically different reports.","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 19. Summarise report length\n# ============================================================\n\ndef report_length_summary(group):\n    return pd.Series({\n        \"studies\": len(group),\n        \"mean_words\": group[\"report_words\"].mean(),\n        \"median_words\": group[\"report_words\"].median(),\n        \"q1_words\": group[\"report_words\"].quantile(0.25),\n        \"q3_words\": group[\"report_words\"].quantile(0.75),\n        \"maximum_words\": group[\"report_words\"].max(),\n        \"mean_characters\": group[\"report_characters\"].mean(),\n        \"median_characters\": group[\"report_characters\"].median(),\n    })\n\nreport_length_by_status = (\n    report_data\n    .groupby(\"label_status\", sort=False)\n    .apply(report_length_summary, include_groups=False)\n    .reset_index()\n)\n\ndisplay(report_length_by_status)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:35.012088Z","iopub.execute_input":"2026-09-19T13:55:35.012352Z","iopub.status.idle":"2026-09-19T13:55:35.038069Z","shell.execute_reply.started":"2026-09-19T13:55:35.012329Z","shell.execute_reply":"2026-09-19T13:55:35.037168Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# 20. Visualise report length\n# ============================================================\n\nword_limit_99 = report_data[\"report_words\"].quantile(0.99)\n\nplot_reports = report_data.loc[\n    report_data[\"report_words\"] <= word_limit_99\n].copy()\n\nfig, axes = plt.subplots(\n    1,\n    2,\n    figsize=(14, 5),\n    gridspec_kw={\"width_ratios\": [2, 1]},\n)\n\nsns.histplot(\n    data=plot_reports,\n    x=\"report_words\",\n    bins=50,\n    color=sns.color_palette(\"deep\")[0],\n    ax=axes[0],\n)\n\naxes[0].set_title(\n    \"Report length distribution\\n\"\n    \"(displayed up to the 99th percentile)\"\n)\naxes[0].set_xlabel(\"Number of words\")\naxes[0].set_ylabel(\"Number of reports\")\n\nsns.boxplot(\n    data=report_data,\n    x=\"label_status\",\n    y=\"report_words\",\n    showfliers=False,\n    color=sns.color_palette(\"deep\")[0],\n    ax=axes[1],\n)\n\naxes[1].set_title(\"Report length by label status\")\naxes[1].set_xlabel(\"\")\naxes[1].set_ylabel(\"Number of words\")\n\nsns.despine()\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:35.039097Z","iopub.execute_input":"2026-09-19T13:55:35.039421Z","iopub.status.idle":"2026-09-19T13:55:35.482111Z","shell.execute_reply.started":"2026-09-19T13:55:35.039394Z","shell.execute_reply":"2026-09-19T13:55:35.481332Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Report length Findings\n\nReport lengths vary considerably and have a long right tail.\n\nGold-labelled reports are generally longer than report-only reports. Their median length is 161.5 words, compared with 128 words for report-only studies. However, the distributions overlap substantially.\n\nThis difference provides further evidence that the 58 gold-labelled studies may be a selected subset rather than a representative sample of all training studies.","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 21. Inspect repeated report text\n# ============================================================\n\nduplicate_report_summary = (\n    report_data\n    .groupby(\"normalized_report\")\n    .agg(\n        number_of_studies=(ID_COLUMN, \"size\"),\n        gold_labelled_studies=(\n            \"label_status\",\n            lambda values: (values == \"Gold-labelled\").sum(),\n        ),\n        example_report=(\"Report\", \"first\"),\n    )\n    .query(\"number_of_studies > 1\")\n    .sort_values(\"number_of_studies\", ascending=False)\n    .reset_index(drop=True)\n)\n\nduplicate_report_summary[\"report_excerpt\"] = (\n    duplicate_report_summary[\"example_report\"]\n    .str.replace(r\"\\s+\", \" \", regex=True)\n    .str.slice(0, 250)\n)\n\ndisplay(\n    duplicate_report_summary[\n        [\n            \"number_of_studies\",\n            \"gold_labelled_studies\",\n            \"report_excerpt\",\n        ]\n    ].head(15)\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:35.483361Z","iopub.execute_input":"2026-09-19T13:55:35.483726Z","iopub.status.idle":"2026-09-19T13:55:35.882831Z","shell.execute_reply.started":"2026-09-19T13:55:35.483688Z","shell.execute_reply":"2026-09-19T13:55:35.881949Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Repeated reports\n\nThe largest repeated-report group contains 37 studies, with several smaller groups appearing in Turkish, Spanish and English.\n\nThese repetitions may represent reporting templates or genuinely duplicated text. Most of the largest groups contain no gold-labelled studies.\n\nWhen validating a future report-based label extractor, studies with identical normalized reports should remain in the same fold. Otherwise, the same report text could appear in both training and validation data and produce an overly optimistic result.","metadata":{}},{"cell_type":"markdown","source":"## Report languages\n\nThe training reports are multilingual. This matters because terminology, negation and reporting conventions differ between languages, making report-to-label extraction more difficult.\n\nWe first estimate the language of each report. These predictions are exploratory rather than definitive, because medical terms and abbreviations can confuse general-purpose language-detection tools.","metadata":{}},{"cell_type":"code","source":"%pip install -q langdetect","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:35.883858Z","iopub.execute_input":"2026-09-19T13:55:35.884127Z","iopub.status.idle":"2026-09-19T13:55:42.40042Z","shell.execute_reply.started":"2026-09-19T13:55:35.884098Z","shell.execute_reply":"2026-09-19T13:55:42.399393Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# 22. Detect and summarise report languages\n# ============================================================\n\nfrom langdetect import (\n    DetectorFactory,\n    LangDetectException,\n    detect_langs,\n)\n\nDetectorFactory.seed = RANDOM_STATE\n\n\ndef detect_report_language(text):\n    \"\"\"Return the most likely language and estimated confidence.\"\"\"\n    try:\n        prediction = detect_langs(text)[0]\n\n        return pd.Series({\n            \"language_code\": prediction.lang,\n            \"language_confidence\": float(prediction.prob),\n        })\n\n    except LangDetectException:\n        return pd.Series({\n            \"language_code\": \"unknown\",\n            \"language_confidence\": np.nan,\n        })\n\n\nlanguage_predictions = (\n    report_data[\"Report\"]\n    .apply(detect_report_language)\n)\n\nreport_data[\n    [\"language_code\", \"language_confidence\"]\n] = language_predictions\n\n\nLANGUAGE_NAMES = {\n    \"af\": \"Afrikaans\",\n    \"ar\": \"Arabic\",\n    \"ca\": \"Catalan\",\n    \"cs\": \"Czech\",\n    \"da\": \"Danish\",\n    \"de\": \"German\",\n    \"en\": \"English\",\n    \"es\": \"Spanish\",\n    \"fi\": \"Finnish\",\n    \"fr\": \"French\",\n    \"hu\": \"Hungarian\",\n    \"it\": \"Italian\",\n    \"ja\": \"Japanese\",\n    \"ko\": \"Korean\",\n    \"nl\": \"Dutch\",\n    \"no\": \"Norwegian\",\n    \"pl\": \"Polish\",\n    \"pt\": \"Portuguese\",\n    \"ro\": \"Romanian\",\n    \"ru\": \"Russian\",\n    \"sv\": \"Swedish\",\n    \"tr\": \"Turkish\",\n    \"uk\": \"Ukrainian\",\n    \"zh-cn\": \"Chinese\",\n    \"bg\": \"Bulgarian\",\n    \"el\": \"Greek\",\n    \"hr\": \"Croatian\",\n}\n\nreport_data[\"language\"] = (\n    report_data[\"language_code\"]\n    .map(LANGUAGE_NAMES)\n    .fillna(report_data[\"language_code\"])\n)\n\nlanguage_summary = (\n    report_data\n    .groupby(\n        [\"language_code\", \"language\"],\n        dropna=False,\n    )\n    .agg(\n        reports=(ID_COLUMN, \"size\"),\n        gold_labelled_reports=(\n            \"label_status\",\n            lambda values: (values == \"Gold-labelled\").sum(),\n        ),\n        median_words=(\"report_words\", \"median\"),\n        mean_confidence=(\"language_confidence\", \"mean\"),\n    )\n    .reset_index()\n)\n\nlanguage_summary[\"percentage\"] = (\n    language_summary[\"reports\"]\n    / len(report_data)\n    * 100\n)\n\nlanguage_summary = (\n    language_summary\n    .sort_values(\"reports\", ascending=False)\n    .reset_index(drop=True)\n)\n\ndisplay(language_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:55:42.402093Z","iopub.execute_input":"2026-09-19T13:55:42.402411Z","iopub.status.idle":"2026-09-19T13:56:09.335827Z","shell.execute_reply.started":"2026-09-19T13:55:42.402378Z","shell.execute_reply":"2026-09-19T13:56:09.335058Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Key findings\n\nNine languages were detected. English is the largest group with 1,736 reports (39.4%), but most reports—60.6%—are written in another language.\n\nSpanish and Turkish are the next most common, followed by Croatian, Greek, German, Bulgarian, Dutch and French. Report length also differs by language: the median ranges from 85 words in Turkish to 215 words in French.\n\nThe 58 gold-labelled studies cover eight of the nine detected languages, but coverage is extremely limited. English has 28 gold reports, while most other languages have only two to ten. None of the 81 French reports has a gold label.\n\nConsequently, the gold subset cannot provide a reliable language-specific evaluation of report-derived labels. In particular, French label extraction cannot be checked directly against any gold-labelled example.\n\nThe gold subset also contains a higher proportion of English reports than the full training set. Language composition may therefore explain part of the difference previously observed between gold and report-only report lengths.","metadata":{}},{"cell_type":"markdown","source":"### Language examples\n\nWe inspect one randomly selected report from each detected language. This provides a simple manual check that the automatic language classifications are reasonable.","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 23. Preview one report from each detected language\n# ============================================================\n\nlanguage_examples = (\n    report_data\n    .groupby(\"language\", group_keys=False)\n    .sample(n=1, random_state=RANDOM_STATE)\n    .copy()\n)\n\nlanguage_examples[\"report_excerpt\"] = (\n    language_examples[\"Report\"]\n    .str.replace(r\"\\s+\", \" \", regex=True)\n    .str.slice(0, 300)\n)\n\nlanguage_examples = (\n    language_examples\n    .merge(\n        language_summary[[\"language\", \"reports\"]],\n        on=\"language\",\n        how=\"left\",\n    )\n    .sort_values(\"reports\", ascending=False)\n)\n\ndisplay(\n    language_examples[\n        [\n            \"language\",\n            \"language_confidence\",\n            \"label_status\",\n            \"report_excerpt\",\n        ]\n    ].reset_index(drop=True)\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:09.336904Z","iopub.execute_input":"2026-09-19T13:56:09.337795Z","iopub.status.idle":"2026-09-19T13:56:09.364759Z","shell.execute_reply.started":"2026-09-19T13:56:09.337768Z","shell.execute_reply":"2026-09-19T13:56:09.364042Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Language detection check\n\nThe sampled excerpts match their predicted languages across all nine groups. Combined with the high confidence values, this suggests that automatic language detection is sufficiently reliable for summarising the dataset.\n\nThis does not make the predictions definitive—particularly for closely related languages—but no obvious classification errors were found in this small manual check.\n\nThe language diversity confirms that weak-label extraction must be multilingual. Report terminology and negation cannot safely be handled using an English-only keyword system.","metadata":{}},{"cell_type":"markdown","source":"## Report structure\n\nRadiology reports may contain separate sections for technique, findings and conclusions. These sections could help a future label extractor distinguish clinically important conclusions from background information.\n\nWe begin with language-independent structural indicators: line breaks, colons and short heading-like lines. These are simple heuristics rather than definitive section labels.","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 24. Summarise report structure\n# ============================================================\n\ndef contains_heading_like_line(text):\n    \"\"\"\n    Identify short lines that resemble headings.\n\n    A line is considered heading-like if it ends with a colon,\n    or if it is a short uppercase line.\n    \"\"\"\n    for line in str(text).splitlines():\n        line = line.strip()\n\n        if not line:\n            continue\n\n        ends_with_colon = (\n            len(line) <= 80\n            and line.endswith(\":\")\n        )\n\n        short_uppercase_line = (\n            3 <= len(line) <= 80\n            and len(line.split()) <= 8\n            and line.isupper()\n        )\n\n        if ends_with_colon or short_uppercase_line:\n            return True\n\n    return False\n\n\nreport_data[\"has_line_break\"] = (\n    report_data[\"Report\"].str.contains(\"\\n\", regex=False)\n)\n\nreport_data[\"colon_count\"] = (\n    report_data[\"Report\"].str.count(\":\")\n)\n\nreport_data[\"has_heading_like_line\"] = (\n    report_data[\"Report\"].apply(contains_heading_like_line)\n)\n\n\nstructure_summary = (\n    report_data\n    .groupby(\"language\")\n    .agg(\n        reports=(ID_COLUMN, \"size\"),\n        median_lines=(\"report_lines\", \"median\"),\n        reports_with_line_breaks=(\"has_line_break\", \"mean\"),\n        reports_with_colons=(\"colon_count\", lambda values: (values > 0).mean()),\n        median_colons=(\"colon_count\", \"median\"),\n        reports_with_heading_like_lines=(\"has_heading_like_line\", \"mean\"),\n    )\n    .reset_index()\n)\n\npercentage_columns = [\n    \"reports_with_line_breaks\",\n    \"reports_with_colons\",\n    \"reports_with_heading_like_lines\",\n]\n\nstructure_summary[percentage_columns] *= 100\n\nstructure_summary = (\n    structure_summary\n    .sort_values(\"reports\", ascending=False)\n    .reset_index(drop=True)\n)\n\ndisplay(structure_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:09.365821Z","iopub.execute_input":"2026-09-19T13:56:09.36658Z","iopub.status.idle":"2026-09-19T13:56:09.43222Z","shell.execute_reply.started":"2026-09-19T13:56:09.366551Z","shell.execute_reply":"2026-09-19T13:56:09.431544Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — report formatting varies by language\n\nReport structure differs substantially across languages. French reports are consistently multi-line and section-like, while Turkish and Dutch reports are almost entirely single-line despite frequently containing colons. Croatian reports use many line breaks but almost no colons.\n\nThese differences likely reflect reporting templates and institutions as well as language. A single rule-based section parser would therefore be unreliable across the dataset; future report processing should account for language-specific formatting.","metadata":{}},{"cell_type":"markdown","source":"### Common report headings\n\nThe previous summary showed large formatting differences between languages. We now inspect the most frequently detected heading-like lines to understand what those structural indicators represent.","metadata":{}},{"cell_type":"code","source":"# 25. Inspect common heading-like lines\n\nheading_rows = []\n\nfor language, report in report_data[[\"language\", \"Report\"]].itertuples(index=False):\n    for line in str(report).splitlines():\n        line = \" \".join(line.split())\n\n        if not line:\n            continue\n\n        ends_with_colon = len(line) <= 80 and line.endswith(\":\")\n        short_uppercase_line = (\n            3 <= len(line) <= 80\n            and len(line.split()) <= 8\n            and line.isupper()\n        )\n\n        if ends_with_colon or short_uppercase_line:\n            heading_rows.append({\n                \"language\": language,\n                \"heading\": line,\n                \"normalized_heading\": line.casefold(),\n            })\n\nheading_data = pd.DataFrame(heading_rows)\n\ncommon_headings = (\n    heading_data\n    .groupby([\"language\", \"normalized_heading\"], as_index=False)\n    .agg(\n        occurrences=(\"heading\", \"size\"),\n        example=(\"heading\", \"first\"),\n    )\n    .sort_values(\n        [\"language\", \"occurrences\"],\n        ascending=[True, False],\n    )\n    .groupby(\"language\", group_keys=False)\n    .head(5)\n    .drop(columns=\"normalized_heading\")\n    .reset_index(drop=True)\n)\n\ndisplay(common_headings)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:09.433233Z","iopub.execute_input":"2026-09-19T13:56:09.433538Z","iopub.status.idle":"2026-09-19T13:56:09.609827Z","shell.execute_reply.started":"2026-09-19T13:56:09.433501Z","shell.execute_reply":"2026-09-19T13:56:09.608916Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — headings are useful but language-dependent\n\nThe heuristic identifies genuine recurring sections in English, Spanish, French and German, including Findings, Impression and Conclusion. These sections may be particularly useful because the final diagnostic summary is often more concise than the full report.\n\nHowever, several languages produce no common headings, and entries such as [DATE]. show that the heuristic can generate false positives. Section-aware extraction could therefore help for some report templates, but it should not be required for every report.","metadata":{}},{"cell_type":"markdown","source":"## Reports and gold labels\n\nThe 58 fully labelled studies provide the only direct evidence of how the radiology reports relate to the competition targets. We begin by placing each gold-labelled report beside its positive targets.\n\nThe examples below are qualitative rather than a test of report–label agreement. A diagnosis may be expressed using synonyms, abbreviations or negation, and many reports contain several positive targets.","metadata":{}},{"cell_type":"code","source":"# Link gold labels to their reports\n\ngold_report_data = (\n    report_data.loc[\n        report_data[\"label_status\"] == \"Gold-labelled\"\n    ]\n    .merge(\n        train[[ID_COLUMN] + TARGET_COLUMNS],\n        on=ID_COLUMN,\n        how=\"left\",\n        validate=\"one_to_one\",\n    )\n)\n\ngold_report_data[\"positive_targets\"] = (\n    gold_report_data[TARGET_COLUMNS]\n    .apply(\n        lambda row: \", \".join(\n            target\n            for target in TARGET_COLUMNS\n            if row[target] == 1\n        ),\n        axis=1,\n    )\n)\n\ndef create_report_excerpt(report, maximum_characters=400):\n    clean_report = \" \".join(str(report).split())\n\n    if len(clean_report) <= maximum_characters:\n        return clean_report\n\n    excerpt_length = (maximum_characters - 5) // 2\n\n    return (\n        clean_report[:excerpt_length]\n        + \" ... \"\n        + clean_report[-excerpt_length:]\n    )\n\ngold_report_data[\"report_excerpt\"] = (\n    gold_report_data[\"Report\"]\n    .apply(create_report_excerpt)\n)\n\ngold_report_examples = (\n    gold_report_data\n    .groupby(\"language\", group_keys=False)\n    .sample(n=1, random_state=RANDOM_STATE)\n    .sort_values(\"language\")\n    [[\"language\", \"positive_targets\", \"report_excerpt\"]]\n    .reset_index(drop=True)\n)\n\ndisplay(gold_report_examples)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:09.611028Z","iopub.execute_input":"2026-09-19T13:56:09.611352Z","iopub.status.idle":"2026-09-19T13:56:09.643954Z","shell.execute_reply.started":"2026-09-19T13:56:09.611324Z","shell.execute_reply":"2026-09-19T13:56:09.643036Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — gold-labelled reports are multilingual and multi-label\n\nThe gold-labelled examples span eight languages, and most reports correspond to several positive targets rather than one isolated diagnosis. The report extractor must therefore identify multiple findings while distinguishing positive diagnoses from normal or explicitly negative statements elsewhere in the same report.\n\nFrench is absent because none of the 81 French reports has gold labels. Consequently, extraction quality for French cannot be measured directly using the provided labelled subset.","metadata":{}},{"cell_type":"markdown","source":"### Report length and diagnostic burden\n\nA report containing more abnormalities may require more detailed description. Within the 58 gold-labelled studies, we examine whether report length is associated with the number of positive targets.\n\nThis is exploratory: the gold subset is small and selected, so the relationship should not be assumed to represent all training studies.","metadata":{}},{"cell_type":"code","source":"# Compare report length with the number of positive targets\n\ngold_report_data[\"number_of_positive_targets\"] = (\n    gold_report_data[TARGET_COLUMNS]\n    .sum(axis=1)\n    .astype(int)\n)\n\nreport_burden_summary = (\n    gold_report_data\n    .groupby(\"number_of_positive_targets\", as_index=False)\n    .agg(\n        studies=(ID_COLUMN, \"size\"),\n        mean_words=(\"report_words\", \"mean\"),\n        median_words=(\"report_words\", \"median\"),\n    )\n)\n\ndisplay(report_burden_summary)\n\nspearman_correlation = (\n    gold_report_data[\"number_of_positive_targets\"]\n    .rank()\n    .corr(gold_report_data[\"report_words\"].rank())\n)\n\nprint(\n    \"Spearman correlation between positive-target count \"\n    f\"and report length: {spearman_correlation:.3f}\"\n)\n\nplt.figure(figsize=(9, 5))\n\nsns.regplot(\n    data=gold_report_data,\n    x=\"number_of_positive_targets\",\n    y=\"report_words\",\n    x_jitter=0.08,\n    ci=None,\n    scatter_kws={\n        \"alpha\": 0.7,\n        \"s\": 55,\n    },\n    line_kws={\n        \"color\": \"darkred\",\n        \"linewidth\": 2,\n    },\n)\n\nplt.title(\"Report length and number of positive targets\")\nplt.xlabel(\"Number of positive targets\")\nplt.ylabel(\"Number of words\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:09.648437Z","iopub.execute_input":"2026-09-19T13:56:09.648766Z","iopub.status.idle":"2026-09-19T13:56:09.862664Z","shell.execute_reply.started":"2026-09-19T13:56:09.648737Z","shell.execute_reply":"2026-09-19T13:56:09.861676Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — report length is a weak indicator of diagnostic burden\n\nReport length has a weak positive association with the number of positive targets, with a Spearman correlation of 0.249. Reports with five or six positive targets tend to be longer, but the pattern is not consistent across the full range.\n\nThe small group sizes also create considerable uncertainty—for example, only one study has eight positive targets and three have nine. Report length may provide some contextual information, but it is not a reliable proxy for the number of abnormalities.","metadata":{}},{"cell_type":"markdown","source":"### Report length by target\n\nSome abnormalities may require more detailed description than others. For each target, we compare the report lengths of positive and negative gold-labelled studies.\n\nBecause abnormalities frequently co-occur, any difference cannot be attributed solely to that target.","metadata":{}},{"cell_type":"code","source":"# Compare report length between positive and negative studies for each target\n\ntarget_report_length_rows = []\n\nfor target in TARGET_COLUMNS:\n    positive_lengths = gold_report_data.loc[\n        gold_report_data[target] == 1,\n        \"report_words\",\n    ]\n\n    negative_lengths = gold_report_data.loc[\n        gold_report_data[target] == 0,\n        \"report_words\",\n    ]\n\n    target_report_length_rows.append({\n        \"target\": target,\n        \"positive_studies\": len(positive_lengths),\n        \"median_words_positive\": positive_lengths.median(),\n        \"median_words_negative\": negative_lengths.median(),\n        \"median_difference\": (\n            positive_lengths.median()\n            - negative_lengths.median()\n        ),\n    })\n\ntarget_report_length_summary = (\n    pd.DataFrame(target_report_length_rows)\n    .sort_values(\"median_difference\", ascending=False)\n    .reset_index(drop=True)\n)\n\ndisplay(target_report_length_summary)\n\nplt.figure(figsize=(9, 6))\n\nax = sns.barplot(\n    data=target_report_length_summary,\n    x=\"median_difference\",\n    y=\"target\",\n    color=\"steelblue\",\n    errorbar=None,\n)\n\nax.axvline(0, color=\"black\", linewidth=1)\n\nplt.title(\"Difference in median report length by target\")\nplt.xlabel(\"Positive median minus negative median (words)\")\nplt.ylabel(\"Target\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:09.863736Z","iopub.execute_input":"2026-09-19T13:56:09.863988Z","iopub.status.idle":"2026-09-19T13:56:10.088832Z","shell.execute_reply.started":"2026-09-19T13:56:09.863964Z","shell.execute_reply":"2026-09-19T13:56:10.088054Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — report-length differences are target-specific\n\nReports with positive lateral osteoarthritis, medial osteoarthritis and lateral meniscus labels are notably longer than their negative counterparts. In contrast, fracture-positive reports have a median length 31 words shorter than fracture-negative reports.\n\nThese differences should not be interpreted as effects of the individual abnormalities. The gold subset is small, targets frequently co-occur, and report length also varies by language and reporting template. Length alone is therefore unlikely to provide reliable evidence for any particular target.","metadata":{}},{"cell_type":"markdown","source":"### Gold target coverage by report language\n\nThe gold labels are the only direct reference for evaluating labels extracted from the reports. We count the positive examples available for each target within each report language.\n\nThe heatmap shows raw counts rather than percentages because the language groups are small and vary greatly in size.","metadata":{}},{"cell_type":"code","source":"# Count gold-positive targets within each report language\n\nlanguage_sample_sizes = (\n    gold_report_data[\"language\"]\n    .value_counts()\n)\n\nlanguage_order = language_sample_sizes.index\n\ngold_target_language_counts = (\n    gold_report_data\n    .groupby(\"language\")[TARGET_COLUMNS]\n    .sum()\n    .astype(int)\n    .loc[language_order]\n)\n\ngold_target_language_counts.index = [\n    f\"{language} (n={language_sample_sizes[language]})\"\n    for language in gold_target_language_counts.index\n]\n\ngold_target_language_counts.index.name = \"report_language\"\n\ndisplay(gold_target_language_counts)\n\nplt.figure(figsize=(14, 6))\n\nsns.heatmap(\n    gold_target_language_counts,\n    annot=True,\n    fmt=\"d\",\n    cmap=\"Blues\",\n    linewidths=0.5,\n    cbar_kws={\"label\": \"Positive gold labels\"},\n)\n\nplt.title(\"Gold-positive target counts by report language\")\nplt.xlabel(\"Target\")\nplt.ylabel(\"Report language\")\nplt.xticks(rotation=45, ha=\"right\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:10.089873Z","iopub.execute_input":"2026-09-19T13:56:10.090265Z","iopub.status.idle":"2026-09-19T13:56:10.525611Z","shell.execute_reply.started":"2026-09-19T13:56:10.090237Z","shell.execute_reply":"2026-09-19T13:56:10.524787Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — per-language gold coverage is extremely limited\n\nEnglish is the only language with at least one positive example for every target, but even there the positive counts range from only 5 to 15. The non-English language groups contain between 2 and 10 gold-labelled reports, leaving many language–target combinations with no positive examples.\n\nA zero in this heatmap does not mean that an abnormality is absent from that language group generally; it only means that it is absent from the very small gold sample. Reliable target-specific evaluation within individual languages is therefore not possible, and French cannot be evaluated at all.","metadata":{}},{"cell_type":"markdown","source":"### Why keyword matching is insufficient: fracture\n\nA report containing the word *fracture* is not necessarily fracture-positive. For example, the report may state that there is “no fracture.”\n\nUsing the English gold-labelled reports, we compare literal fracture-term mentions with the actual fracture label and inspect several examples. This is a diagnostic illustration, not a label-extraction method.","metadata":{}},{"cell_type":"code","source":"# Audit literal fracture mentions in English gold-labelled reports\n\nimport re\n\nFRACTURE_PATTERN = r\"\\bfractur\\w*\"\n\nenglish_gold_reports = (\n    gold_report_data.loc[\n        gold_report_data[\"language\"] == \"English\"\n    ]\n    .copy()\n)\n\nenglish_gold_reports[\"fracture_term_present\"] = (\n    english_gold_reports[\"Report\"]\n    .str.contains(\n        FRACTURE_PATTERN,\n        case=False,\n        regex=True,\n        na=False,\n    )\n)\n\nfracture_mention_table = pd.crosstab(\n    english_gold_reports[\"Fracture\"].map({\n        0: \"Gold negative\",\n        1: \"Gold positive\",\n    }),\n    english_gold_reports[\"fracture_term_present\"].map({\n        False: \"Term absent\",\n        True: \"Term present\",\n    }),\n    margins=True,\n)\n\ndisplay(fracture_mention_table)\n\ndef extract_term_context(report, pattern, context_characters=120):\n    clean_report = \" \".join(str(report).split())\n    match = re.search(pattern, clean_report, flags=re.IGNORECASE)\n\n    if match is None:\n        return \"[term not present]\"\n\n    start = max(0, match.start() - context_characters)\n    end = min(\n        len(clean_report),\n        match.end() + context_characters,\n    )\n\n    prefix = \"...\" if start > 0 else \"\"\n    suffix = \"...\" if end < len(clean_report) else \"\"\n\n    return prefix + clean_report[start:end] + suffix\n\nenglish_gold_reports[\"fracture_context\"] = (\n    english_gold_reports[\"Report\"]\n    .apply(\n        lambda report: extract_term_context(\n            report,\n            FRACTURE_PATTERN,\n        )\n    )\n)\n\nfracture_examples = pd.concat([\n    english_gold_reports.loc[\n        (english_gold_reports[\"Fracture\"] == 1)\n        & english_gold_reports[\"fracture_term_present\"]\n    ].sort_values(ID_COLUMN).head(3),\n\n    english_gold_reports.loc[\n        (english_gold_reports[\"Fracture\"] == 0)\n        & english_gold_reports[\"fracture_term_present\"]\n    ].sort_values(ID_COLUMN).head(3),\n])\n\nfracture_examples[\"gold_fracture_label\"] = (\n    fracture_examples[\"Fracture\"]\n    .map({\n        0: \"Negative\",\n        1: \"Positive\",\n    })\n)\n\ndisplay(\n    fracture_examples[\n        [\n            \"gold_fracture_label\",\n            \"positive_targets\",\n            \"fracture_context\",\n        ]\n    ].reset_index(drop=True)\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:10.526958Z","iopub.execute_input":"2026-09-19T13:56:10.527391Z","iopub.status.idle":"2026-09-19T13:56:10.589824Z","shell.execute_reply.started":"2026-09-19T13:56:10.52735Z","shell.execute_reply":"2026-09-19T13:56:10.589136Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — mentioning fracture is not the same as diagnosing it\n\nThe fracture term appears in 7 of 8 fracture-positive English reports, but it also appears in 7 of 20 fracture-negative reports. If literal term presence were treated as a positive prediction, only half of the 14 predicted positives would be correct.\n\nOne gold-positive report also does not contain the searched fracture term, showing that keyword matching can miss alternative wording. A useful report extractor must therefore handle both negation and linguistic variation rather than relying on term presence alone.","metadata":{}},{"cell_type":"markdown","source":"### Literal concept mentions across targets\n\nFor each target, we define a small set of English terms associated with that anatomical structure or finding. We then compare how often the concept is mentioned in gold-positive and gold-negative English reports.\n\nFor the osteoarthritis targets, the patterns identify discussion of the relevant compartment rather than an explicit osteoarthritis diagnosis. The results measure textual mentions only; they do not account for negation, severity or diagnostic context.","metadata":{}},{"cell_type":"code","source":"# Compare literal concept mentions across English gold reports\n\nENGLISH_TARGET_PATTERNS = {\n    \"ACL\": (\n        r\"\\bACL\\b|anterior cruciate ligament\"\n    ),\n    \"MCL\": (\n        r\"\\bMCL\\b|medial collateral ligament\"\n    ),\n    \"Medial Meniscus\": (\n        r\"medial menisc\\w*\"\n    ),\n    \"Lateral Meniscus\": (\n        r\"lateral menisc\\w*\"\n    ),\n    \"Medial OA\": (\n        r\"medial compartment|medial femorotibial\"\n    ),\n    \"Lateral OA\": (\n        r\"lateral compartment|lateral femorotibial\"\n    ),\n    \"PF OA\": (\n        r\"patellofemoral|patello[- ]femoral\"\n    ),\n    \"Effusion\": (\n        r\"\\beffusion\\b\"\n    ),\n    \"Synovitis\": (\n        r\"\\bsynovit\\w*\"\n    ),\n    \"Baker's\": (\n        r\"\\bbaker(?:'s|s)?\\b|popliteal cyst\"\n    ),\n    \"Contusion\": (\n        r\"\\bcontusion\\b|bone bruise|\"\n        r\"bone marrow (?:edema|oedema)\"\n    ),\n    \"Fracture\": (\n        r\"\\bfractur\\w*\"\n    ),\n}\n\ntarget_term_rows = []\n\nfor target, pattern in ENGLISH_TARGET_PATTERNS.items():\n    term_present = (\n        english_gold_reports[\"Report\"]\n        .str.contains(\n            pattern,\n            case=False,\n            regex=True,\n            na=False,\n        )\n    )\n\n    positive_mask = english_gold_reports[target] == 1\n    negative_mask = english_gold_reports[target] == 0\n\n    positive_mention_percentage = (\n        term_present[positive_mask].mean() * 100\n    )\n\n    negative_mention_percentage = (\n        term_present[negative_mask].mean() * 100\n    )\n\n    target_term_rows.append({\n        \"target\": target,\n        \"positive_reports\": positive_mask.sum(),\n        \"positive_reports_with_term_percentage\":\n            positive_mention_percentage,\n        \"negative_reports\": negative_mask.sum(),\n        \"negative_reports_with_term_percentage\":\n            negative_mention_percentage,\n        \"percentage_point_difference\": (\n            positive_mention_percentage\n            - negative_mention_percentage\n        ),\n    })\n\ntarget_term_summary = (\n    pd.DataFrame(target_term_rows)\n    .sort_values(\n        \"percentage_point_difference\",\n        ascending=False,\n    )\n    .reset_index(drop=True)\n)\n\ndisplay(target_term_summary)\n\ntarget_term_plot_data = (\n    target_term_summary\n    .melt(\n        id_vars=\"target\",\n        value_vars=[\n            \"positive_reports_with_term_percentage\",\n            \"negative_reports_with_term_percentage\",\n        ],\n        var_name=\"gold_label_status\",\n        value_name=\"reports_with_term_percentage\",\n    )\n)\n\ntarget_term_plot_data[\"gold_label_status\"] = (\n    target_term_plot_data[\"gold_label_status\"]\n    .map({\n        \"positive_reports_with_term_percentage\":\n            \"Gold positive\",\n        \"negative_reports_with_term_percentage\":\n            \"Gold negative\",\n    })\n)\n\nplt.figure(figsize=(10, 7))\n\nsns.barplot(\n    data=target_term_plot_data,\n    x=\"reports_with_term_percentage\",\n    y=\"target\",\n    hue=\"gold_label_status\",\n    order=target_term_summary[\"target\"],\n    errorbar=None,\n)\n\nplt.xlim(0, 100)\nplt.title(\"Literal concept mentions in English gold reports\")\nplt.xlabel(\"Reports containing the target-related term (%)\")\nplt.ylabel(\"Target\")\nplt.legend(title=\"Gold label\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:10.590944Z","iopub.execute_input":"2026-09-19T13:56:10.591276Z","iopub.status.idle":"2026-09-19T13:56:10.900631Z","shell.execute_reply.started":"2026-09-19T13:56:10.591239Z","shell.execute_reply":"2026-09-19T13:56:10.899823Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — context carries more information than term presence\n\nTarget-related concepts are generally mentioned more often in positive reports, but term presence alone does not separate the labels reliably. Fracture, Baker's cyst and medial osteoarthritis show the largest differences, yet their terms also occur frequently in negative reports.\n\nThe clearest examples are ACL and effusion: their associated terms appear in every English gold report, regardless of the label. The diagnostic information must therefore come from surrounding language such as *intact*, *torn*, *no*, *small* or *large*, rather than from the target term itself.\n\nThe PF OA result also shows the limitations of manually chosen patterns. A broad patellofemoral concept pattern may identify discussion of that compartment without establishing osteoarthritis.","metadata":{}},{"cell_type":"markdown","source":"## Report analysis conclusion\n\nAll training studies have non-empty reports, so the full training set provides potential text-based information. The reports are multilingual and their structure varies by language, while repeated templates appear in a small subset.\n\nThe 58 gold-labelled reports cover only eight languages and provide limited target-specific examples, with no French-labelled studies. Reports can therefore support weak supervision, but report-based labels will require language-aware processing and careful validation against the small gold subset.","metadata":{}},{"cell_type":"markdown","source":"# MRI series metadata\n\nEach study may contain several MRI series acquired using different planes and sequence settings. We first examine how many series are available per study.","metadata":{}},{"cell_type":"code","source":"# 26. Summarise the number of MRI series per study\n\nseries_count_rows = []\n\nfor split_name, series_table in {\n    \"Train\": train_series,\n    \"Test\": test_series,\n}.items():\n    series_per_study = (\n        series_table\n        .groupby(ID_COLUMN)[\"SeriesInstanceUID\"]\n        .nunique()\n    )\n\n    series_count_rows.append({\n        \"split\": split_name,\n        \"studies\": series_per_study.size,\n        \"total_series\": series_per_study.sum(),\n        \"minimum\": series_per_study.min(),\n        \"q1\": series_per_study.quantile(0.25),\n        \"median\": series_per_study.median(),\n        \"mean\": series_per_study.mean(),\n        \"q3\": series_per_study.quantile(0.75),\n        \"maximum\": series_per_study.max(),\n    })\n\nseries_count_summary = pd.DataFrame(series_count_rows)\n\ndisplay(series_count_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:10.901652Z","iopub.execute_input":"2026-09-19T13:56:10.901902Z","iopub.status.idle":"2026-09-19T13:56:10.934337Z","shell.execute_reply.started":"2026-09-19T13:56:10.901877Z","shell.execute_reply":"2026-09-19T13:56:10.933542Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Distribution of series counts\n\nThe summary table describes the number of MRI series per study. A bar chart makes it easier to see whether most studies use the same number of series or whether the distribution is highly variable.","metadata":{}},{"cell_type":"code","source":"# 27. Visualise the number of series per training study\n\ntrain_series_per_study = (\n    train_series\n    .groupby(ID_COLUMN)[\"SeriesInstanceUID\"]\n    .nunique()\n)\n\nseries_count_distribution = (\n    train_series_per_study\n    .value_counts()\n    .sort_index()\n    .rename_axis(\"number_of_series\")\n    .reset_index(name=\"number_of_studies\")\n)\n\ndisplay(series_count_distribution)\n\nplt.figure(figsize=(9, 5))\n\nax = sns.barplot(\n    data=series_count_distribution,\n    x=\"number_of_series\",\n    y=\"number_of_studies\",\n    color=\"steelblue\",\n    errorbar=None,\n)\n\nax.bar_label(ax.containers[0], fmt=\"%d\", padding=3)\n\nplt.title(\"Distribution of MRI series per training study\")\nplt.xlabel(\"Number of MRI series\")\nplt.ylabel(\"Number of studies\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:10.935417Z","iopub.execute_input":"2026-09-19T13:56:10.935782Z","iopub.status.idle":"2026-09-19T13:56:11.171673Z","shell.execute_reply.started":"2026-09-19T13:56:10.935746Z","shell.execute_reply":"2026-09-19T13:56:11.170898Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — five series is the dominant study structure\n\nFive MRI series is the most common configuration, occurring in 2,299 studies, or approximately 52% of the training set. Most studies contain between four and six series, but the number ranges from three to fourteen.\n\nThis variation suggests that some studies contain additional sequences or acquisitions. A future model should therefore handle a variable number of series rather than assuming that every study has exactly five.","metadata":{}},{"cell_type":"markdown","source":"## Anatomical plane\n\nMRI sequences can be acquired in different anatomical planes. These planes provide different views of the knee and may contain complementary information for abnormality detection.","metadata":{}},{"cell_type":"code","source":"# 28. Distribution of anatomical planes\n\nplane_distribution = (\n    train_series[\"Anatomical_Plane\"]\n    .value_counts(dropna=False)\n    .rename_axis(\"anatomical_plane\")\n    .reset_index(name=\"number_of_series\")\n)\n\nplane_distribution[\"percentage\"] = (\n    plane_distribution[\"number_of_series\"]\n    / len(train_series)\n    * 100\n)\n\ndisplay(plane_distribution)\n\nplt.figure(figsize=(8, 5))\n\nax = sns.barplot(\n    data=plane_distribution,\n    x=\"anatomical_plane\",\n    y=\"number_of_series\",\n    color=\"steelblue\",\n    errorbar=None,\n)\n\nax.bar_label(ax.containers[0], fmt=\"%d\", padding=3)\n\nplt.title(\"Distribution of anatomical planes\")\nplt.xlabel(\"Anatomical plane\")\nplt.ylabel(\"Number of MRI series\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:11.17291Z","iopub.execute_input":"2026-09-19T13:56:11.173713Z","iopub.status.idle":"2026-09-19T13:56:11.340089Z","shell.execute_reply.started":"2026-09-19T13:56:11.173684Z","shell.execute_reply":"2026-09-19T13:56:11.339194Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — sagittal and coronal series are most common\n\nSagittal series account for approximately 40% of all training series, followed by coronal series at 35% and axial series at 24%. The dataset therefore provides multiple complementary views of the knee, although axial imaging is somewhat less frequent.\n\nThese are series-level percentages, so studies containing more series contribute more heavily to the totals. We should also examine the plane distribution at the study level before drawing conclusions about individual examinations.","metadata":{}},{"cell_type":"markdown","source":"### Anatomical planes per study\n\nA series-level distribution can be influenced by studies with many MRI series. Here we calculate the percentage of studies containing at least one series from each anatomical plane.","metadata":{}},{"cell_type":"code","source":"# 29. Anatomical plane coverage across studies\n\nstudy_plane_presence = pd.crosstab(\n    train_series[ID_COLUMN],\n    train_series[\"Anatomical_Plane\"],\n).gt(0)\n\nstudy_plane_presence = study_plane_presence.reindex(\n    columns=[\"Sagittal\", \"Coronal\", \"Axial\"],\n    fill_value=False,\n)\n\nplane_study_summary = (\n    study_plane_presence\n    .sum()\n    .rename(\"number_of_studies\")\n    .reset_index()\n    .rename(columns={\"Anatomical_Plane\": \"anatomical_plane\"})\n)\n\nplane_study_summary[\"percentage_of_studies\"] = (\n    plane_study_summary[\"number_of_studies\"]\n    / len(study_plane_presence)\n    * 100\n)\n\ndisplay(plane_study_summary)\n\nplt.figure(figsize=(8, 5))\n\nax = sns.barplot(\n    data=plane_study_summary,\n    x=\"anatomical_plane\",\n    y=\"percentage_of_studies\",\n    color=\"steelblue\",\n    errorbar=None,\n)\n\nax.bar_label(ax.containers[0], fmt=\"%.1f%%\", padding=3)\n\nplt.ylim(0, 105)\nplt.title(\"Studies containing each anatomical plane\")\nplt.xlabel(\"Anatomical plane\")\nplt.ylabel(\"Percentage of studies\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:11.341292Z","iopub.execute_input":"2026-09-19T13:56:11.341843Z","iopub.status.idle":"2026-09-19T13:56:11.66891Z","shell.execute_reply.started":"2026-09-19T13:56:11.341805Z","shell.execute_reply":"2026-09-19T13:56:11.668021Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — every study contains all three anatomical planes\n\nEvery training study contains at least one sagittal, coronal and axial series. Therefore, the lower number of axial series does not mean that axial imaging is missing from some studies; it means that studies generally contain fewer axial series than sagittal or coronal series.","metadata":{}},{"cell_type":"markdown","source":"## Fluid Sensitivity & Fat Suppression \n\nThe series table also records whether each series is fluid-sensitive and whether fat suppression was used. These properties describe how the MRI sequence was acquired and may help distinguish the most informative series for each abnormality.","metadata":{}},{"cell_type":"code","source":"# 30. Distribution of sequence properties\n\nsequence_properties = [\"Fluid_Sensitive\", \"Fat_Suppression\"]\n\nsequence_property_rows = []\n\nfor property_name in sequence_properties:\n    value_counts = (\n        train_series[property_name]\n        .value_counts(dropna=False)\n        .sort_index()\n    )\n\n    for value, count in value_counts.items():\n        sequence_property_rows.append({\n            \"property\": property_name,\n            \"value\": value,\n            \"number_of_series\": count,\n            \"percentage\": count / len(train_series) * 100,\n        })\n\nsequence_property_summary = pd.DataFrame(sequence_property_rows)\n\ndisplay(sequence_property_summary)\n\nfig, axes = plt.subplots(1, 2, figsize=(12, 5))\n\nfor ax, property_name in zip(axes, sequence_properties):\n    plot_data = sequence_property_summary[\n        sequence_property_summary[\"property\"] == property_name\n    ]\n\n    sns.barplot(\n        data=plot_data,\n        x=\"value\",\n        y=\"number_of_series\",\n        color=\"steelblue\",\n        errorbar=None,\n        ax=ax,\n    )\n\n    ax.bar_label(ax.containers[0], fmt=\"%d\", padding=3)\n    ax.set_title(property_name.replace(\"_\", \" \"))\n    ax.set_xlabel(\"Recorded value\")\n    ax.set_ylabel(\"Number of series\")\n\nplt.suptitle(\"Distribution of sequence properties\", y=1.02)\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:11.670025Z","iopub.execute_input":"2026-09-19T13:56:11.670363Z","iopub.status.idle":"2026-09-19T13:56:11.969711Z","shell.execute_reply.started":"2026-09-19T13:56:11.670332Z","shell.execute_reply":"2026-09-19T13:56:11.968972Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — the two sequence flags have identical distributions\n\nBoth `Fluid_Sensitive` and `Fat_Suppression` have exactly the same distribution: approximately 57.5% of series have value `1` and 42.5% have value `0`.\n\nThis may indicate that the two properties are perfectly aligned in this dataset. If so, they would be redundant features, although we should verify their row-by-row relationship before drawing that conclusion.","metadata":{}},{"cell_type":"markdown","source":"### Relationship between the two sequence flags\n\nIdentical marginal distributions do not necessarily mean that two variables are identical row by row. We therefore compare the two flags directly.","metadata":{}},{"cell_type":"code","source":"# 31. Compare Fluid_Sensitive and Fat_Suppression row by row\n\nflag_crosstab = pd.crosstab(\n    train_series[\"Fluid_Sensitive\"],\n    train_series[\"Fat_Suppression\"],\n    margins=True,\n)\n\nflag_crosstab.index.name = \"Fluid_Sensitive\"\nflag_crosstab.columns.name = \"Fat_Suppression\"\n\ndisplay(flag_crosstab)\n\nnumber_of_mismatches = (\n    train_series[\"Fluid_Sensitive\"]\n    != train_series[\"Fat_Suppression\"]\n).sum()\n\nagreement_percentage = (\n    1 - number_of_mismatches / len(train_series)\n) * 100\n\nprint(f\"Number of mismatched series: {number_of_mismatches:,}\")\nprint(f\"Percentage of series in agreement: {agreement_percentage:.2f}%\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:11.970736Z","iopub.execute_input":"2026-09-19T13:56:11.971061Z","iopub.status.idle":"2026-09-19T13:56:12.004898Z","shell.execute_reply.started":"2026-09-19T13:56:11.971025Z","shell.execute_reply":"2026-09-19T13:56:12.004056Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — the sequence flags are duplicates\n\n`Fluid_Sensitive` and `Fat_Suppression` agree for all 24,371 training series. Every series with value `0` for one flag also has value `0` for the other, and the same is true for value `1`.\n\nThese columns therefore contain no independent information. For modelling, one of them can be removed without losing information. We will retain both in the raw-data description but should avoid treating them as two separate predictors.","metadata":{}},{"cell_type":"markdown","source":"### Sequence properties by anatomical plane\n\nThe usefulness of a sequence depends on both its anatomical plane and its acquisition properties. We therefore examine how the fluid-sensitive flag is distributed within each plane.","metadata":{}},{"cell_type":"code","source":"# 32. Compare fluid sensitivity across anatomical planes\n\nplane_sequence_summary = (\n    train_series\n    .groupby(\n        [\"Anatomical_Plane\", \"Fluid_Sensitive\"],\n        as_index=False,\n    )\n    .size()\n    .rename(columns={\"size\": \"number_of_series\"})\n)\n\nplane_sequence_summary[\"percentage_within_plane\"] = (\n    plane_sequence_summary[\"number_of_series\"]\n    / plane_sequence_summary.groupby(\"Anatomical_Plane\")[\n        \"number_of_series\"\n    ].transform(\"sum\")\n    * 100\n)\n\nplane_sequence_summary[\"sequence_type\"] = (\n    plane_sequence_summary[\"Fluid_Sensitive\"]\n    .map({\n        0: \"Not fluid-sensitive\",\n        1: \"Fluid-sensitive\",\n    })\n)\n\ndisplay(plane_sequence_summary)\n\nplt.figure(figsize=(9, 5))\n\nax = sns.barplot(\n    data=plane_sequence_summary,\n    x=\"Anatomical_Plane\",\n    y=\"percentage_within_plane\",\n    hue=\"sequence_type\",\n    errorbar=None,\n)\n\nfor container in ax.containers:\n    ax.bar_label(container, fmt=\"%.1f%%\", padding=3)\n\nplt.ylim(0, 100)\nplt.title(\"Fluid sensitivity within each anatomical plane\")\nplt.xlabel(\"Anatomical plane\")\nplt.ylabel(\"Percentage of series within plane\")\nplt.legend(title=\"Sequence type\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:12.006102Z","iopub.execute_input":"2026-09-19T13:56:12.006474Z","iopub.status.idle":"2026-09-19T13:56:12.253843Z","shell.execute_reply.started":"2026-09-19T13:56:12.006415Z","shell.execute_reply":"2026-09-19T13:56:12.253009Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — sequence properties vary by anatomical plane\n\nFluid-sensitive sequences are especially common in the axial plane, where they account for 80% of series. The split is much more even for coronal and sagittal series.\n\nAnatomical plane and sequence type are therefore related rather than independent metadata fields. A future imaging model may benefit from distinguishing these plane–sequence combinations when selecting or combining series.","metadata":{}},{"cell_type":"markdown","source":"# Study–series data integrity\n\nThe competition contains separate study-level and series-level tables. Before examining image files, we verify that their identifiers are unique, that every series belongs to the expected split, and that training and test identifiers do not overlap.\n\nThese checks are important because accidental identifier duplication or overlap could cause incorrect joins or data leakage.","metadata":{}},{"cell_type":"code","source":"# 33. Validate study and series identifiers\n\ndef add_integrity_check(check_name, value, expected=0):\n    return {\n        \"check\": check_name,\n        \"value\": int(value),\n        \"status\": \"PASS\" if int(value) == expected else \"CHECK\",\n    }\n\ntrain_study_ids = set(train[ID_COLUMN].dropna())\ntest_study_ids = set(test[ID_COLUMN].dropna())\n\ntrain_series_study_ids = set(\n    train_series[ID_COLUMN].dropna()\n)\ntest_series_study_ids = set(\n    test_series[ID_COLUMN].dropna()\n)\n\ntrain_series_ids = set(\n    train_series[\"SeriesInstanceUID\"].dropna()\n)\ntest_series_ids = set(\n    test_series[\"SeriesInstanceUID\"].dropna()\n)\n\nseries_to_study_counts = (\n    train_series\n    .groupby(\"SeriesInstanceUID\")[ID_COLUMN]\n    .nunique()\n)\n\nintegrity_checks = [\n    add_integrity_check(\n        \"Duplicate training study IDs\",\n        train[ID_COLUMN].duplicated().sum(),\n    ),\n    add_integrity_check(\n        \"Duplicate test study IDs\",\n        test[ID_COLUMN].duplicated().sum(),\n    ),\n    add_integrity_check(\n        \"Duplicate training series IDs\",\n        train_series[\"SeriesInstanceUID\"].duplicated().sum(),\n    ),\n    add_integrity_check(\n        \"Duplicate test series IDs\",\n        test_series[\"SeriesInstanceUID\"].duplicated().sum(),\n    ),\n    add_integrity_check(\n        \"Training studies missing from train_series\",\n        len(train_study_ids - train_series_study_ids),\n    ),\n    add_integrity_check(\n        \"Unknown studies in train_series\",\n        len(train_series_study_ids - train_study_ids),\n    ),\n    add_integrity_check(\n        \"Test studies missing from test_series\",\n        len(test_study_ids - test_series_study_ids),\n    ),\n    add_integrity_check(\n        \"Unknown studies in test_series\",\n        len(test_series_study_ids - test_study_ids),\n    ),\n    add_integrity_check(\n        \"Study IDs shared by train and test\",\n        len(train_study_ids & test_study_ids),\n    ),\n    add_integrity_check(\n        \"Series IDs shared by train and test\",\n        len(train_series_ids & test_series_ids),\n    ),\n    add_integrity_check(\n        \"Series IDs assigned to multiple studies\",\n        (series_to_study_counts > 1).sum(),\n    ),\n]\n\nintegrity_summary = pd.DataFrame(integrity_checks)\n\ndisplay(integrity_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:12.255264Z","iopub.execute_input":"2026-09-19T13:56:12.255622Z","iopub.status.idle":"2026-09-19T13:56:12.320634Z","shell.execute_reply.started":"2026-09-19T13:56:12.255593Z","shell.execute_reply":"2026-09-19T13:56:12.319892Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — the study and series tables join cleanly\n\nAll identifier checks passed:\n\n- every study has a unique identifier;\n- every series has a unique identifier;\n- every study in the training and test tables has corresponding series metadata;\n- no series belongs to more than one study; and\n- there is no study or series identifier overlap between training and test.\n\nThis gives us a clean relational structure for joining study-level labels, series-level metadata and image files. These checks do not rule out visually similar or duplicated examinations, so possible related studies will need separate investigation later.","metadata":{}},{"cell_type":"markdown","source":"# Image file structure\n\nThe metadata tables describe the studies and series, but the MRI images are stored separately on disk. We first inspect a small sample of the directory hierarchy to confirm how studies, series and image files are connected.","metadata":{}},{"cell_type":"code","source":"# 34. Inspect a sample of the image directory hierarchy\n\ndef sample_image_tree(\n    root,\n    split_name,\n    maximum_studies=2,\n    maximum_series_per_study=3,\n):\n    rows = []\n\n    if not root.exists():\n        print(f\"{split_name} image directory not found: {root}\")\n        return pd.DataFrame()\n\n    study_entries = sorted(\n        [\n            entry\n            for entry in os.scandir(root)\n            if entry.is_dir(follow_symlinks=False)\n        ],\n        key=lambda entry: entry.name,\n    )\n\n    print(\n        f\"{split_name}: {len(study_entries):,} \"\n        f\"study directories found\"\n    )\n\n    for study_entry in study_entries[:maximum_studies]:\n        series_entries = sorted(\n            [\n                entry\n                for entry in os.scandir(study_entry.path)\n                if entry.is_dir(follow_symlinks=False)\n            ],\n            key=lambda entry: entry.name,\n        )\n\n        for series_entry in series_entries[:maximum_series_per_study]:\n            file_entries = [\n                entry\n                for entry in os.scandir(series_entry.path)\n                if entry.is_file(follow_symlinks=False)\n            ]\n\n            extension_counts = Counter(\n                Path(entry.name).suffix.lower() or \"[no extension]\"\n                for entry in file_entries\n            )\n\n            rows.append({\n                \"split\": split_name,\n                \"study_directory\": study_entry.name,\n                \"series_directory\": series_entry.name,\n                \"number_of_files\": len(file_entries),\n                \"file_types\": dict(extension_counts),\n            })\n\n    return pd.DataFrame(rows)\n\n\nTRAIN_IMAGES_ROOT = COMPETITION_ROOT / \"train_series\"\nTEST_IMAGES_ROOT = COMPETITION_ROOT / \"test_series\"\n\nimage_tree_sample = pd.concat([\n    sample_image_tree(\n        TRAIN_IMAGES_ROOT,\n        \"Train\",\n    ),\n    sample_image_tree(\n        TEST_IMAGES_ROOT,\n        \"Test\",\n    ),\n], ignore_index=True)\n\ndisplay(image_tree_sample)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:12.321619Z","iopub.execute_input":"2026-09-19T13:56:12.321917Z","iopub.status.idle":"2026-09-19T13:56:14.061394Z","shell.execute_reply.started":"2026-09-19T13:56:12.321881Z","shell.execute_reply":"2026-09-19T13:56:14.060593Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — each series is stored as a DICOM image stack\n\nThe sampled directories follow a clear structure:\n\n`study directory → series directory → DICOM files`\n\nAll sampled files use the `.dcm` extension, and the number of images varies between series. The examples contain between 16 and 38 DICOM files, so a series represents a stack of slices rather than a single image.\n\nThis was checked on a sample, so we should now count the image files across all series to identify missing, empty or non-DICOM directories and to understand the full slice-count distribution.","metadata":{}},{"cell_type":"markdown","source":"## Number of DICOM images per series\n\nThe number of files in a series corresponds approximately to the number of MRI slices available for that acquisition. We scan only the immediate files inside each known series directory, avoiding a recursive walk through the entire dataset.","metadata":{}},{"cell_type":"markdown","source":"### Sampling series sizes\n\nA complete scan of every image directory is unnecessarily expensive at this stage. We first sample series from each anatomical plane and count their files. This gives us an initial view of slice counts while keeping the EDA responsive.","metadata":{}},{"cell_type":"code","source":"# 35. Sample the number of image files per series\n\nsampled_train_series = pd.concat([\n    group.sample(\n        n=min(50, len(group)),\n        random_state=RANDOM_STATE,\n    )\n    for _, group in train_series.groupby(\"Anatomical_Plane\")\n])\n\nsampled_series_tables = [\n    (\"Train\", sampled_train_series, TRAIN_IMAGES_ROOT),\n    (\"Test\", test_series, TEST_IMAGES_ROOT),\n]\n\nsampled_file_rows = []\n\nfor split_name, series_table, image_root in sampled_series_tables:\n    for row in series_table[\n        [ID_COLUMN, \"SeriesInstanceUID\", \"Anatomical_Plane\"]\n    ].itertuples(index=False):\n        study_uid = getattr(row, ID_COLUMN)\n        series_uid = row.SeriesInstanceUID\n        series_path = image_root / study_uid / series_uid\n\n        if not series_path.exists():\n            sampled_file_rows.append({\n                \"split\": split_name,\n                \"anatomical_plane\": row.Anatomical_Plane,\n                \"number_of_files\": np.nan,\n                \"directory_exists\": False,\n            })\n            continue\n\n        number_of_files = sum(\n            entry.is_file(follow_symlinks=False)\n            for entry in os.scandir(series_path)\n        )\n\n        sampled_file_rows.append({\n            \"split\": split_name,\n            \"anatomical_plane\": row.Anatomical_Plane,\n            \"number_of_files\": number_of_files,\n            \"directory_exists\": True,\n        })\n\nsampled_file_inventory = pd.DataFrame(sampled_file_rows)\n\nsampled_file_summary = (\n    sampled_file_inventory\n    .groupby(\"split\")\n    .agg(\n        sampled_series=(\"number_of_files\", \"size\"),\n        missing_directories=(\n            \"directory_exists\",\n            lambda values: (~values).sum(),\n        ),\n        minimum_files=(\"number_of_files\", \"min\"),\n        median_files=(\"number_of_files\", \"median\"),\n        mean_files=(\"number_of_files\", \"mean\"),\n        maximum_files=(\"number_of_files\", \"max\"),\n    )\n    .reset_index()\n)\n\ndisplay(sampled_file_summary)\n\nplt.figure(figsize=(9, 5))\n\nsns.boxplot(\n    data=sampled_file_inventory.loc[\n        sampled_file_inventory[\"split\"] == \"Train\"\n    ],\n    x=\"anatomical_plane\",\n    y=\"number_of_files\",\n    color=\"steelblue\",\n)\n\nplt.title(\"Sampled number of files per training series\")\nplt.xlabel(\"Anatomical plane\")\nplt.ylabel(\"Number of image files\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:14.062572Z","iopub.execute_input":"2026-09-19T13:56:14.062878Z","iopub.status.idle":"2026-09-19T13:56:16.191823Z","shell.execute_reply.started":"2026-09-19T13:56:14.062852Z","shell.execute_reply":"2026-09-19T13:56:16.191041Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — slice counts vary across series\n\nThe sampled series contain a median of 30 image files. The training sample ranges from 12 to 160 files, while the visible test series range from 24 to 160.\n\nCoronal series have the tightest distribution in this sample. Axial and sagittal series show more variation and several high-count outliers. These counts describe files in a series, but they do not yet tell us whether image dimensions, spacing or other DICOM properties also vary.\n\nThe test summary is based on only 15 series, so it should not be treated as representative of the full hidden test set.","metadata":{}},{"cell_type":"markdown","source":"## DICOM header metadata\n\nDICOM files contain more information than the CSV metadata, including image dimensions, pixel spacing, slice thickness and acquisition identifiers.\n\nWe inspect one representative image from a small sample of series to check these properties and confirm that the identifiers stored inside the DICOM files agree with the directory structure.","metadata":{}},{"cell_type":"code","source":"# 36. Inspect representative DICOM headers\n\nimport pydicom\n\nheader_train_series = pd.concat([\n    group.sample(\n        n=min(3, len(group)),\n        random_state=RANDOM_STATE,\n    )\n    for _, group in train_series.groupby(\"Anatomical_Plane\")\n])\n\nheader_sample_tables = [\n    (\"Train\", header_train_series, TRAIN_IMAGES_ROOT),\n    (\"Test\", test_series, TEST_IMAGES_ROOT),\n]\n\ndicom_header_rows = []\n\nfor split_name, series_table, image_root in header_sample_tables:\n    for row in series_table[\n        [ID_COLUMN, \"SeriesInstanceUID\", \"Anatomical_Plane\"]\n    ].itertuples(index=False):\n\n        study_uid = getattr(row, ID_COLUMN)\n        series_uid = row.SeriesInstanceUID\n        series_path = image_root / study_uid / series_uid\n\n        if not series_path.exists():\n            continue\n\n        dicom_entries = sorted(\n            [\n                entry\n                for entry in os.scandir(series_path)\n                if (\n                    entry.is_file(follow_symlinks=False)\n                    and Path(entry.name).suffix.lower() == \".dcm\"\n                )\n            ],\n            key=lambda entry: entry.name,\n        )\n\n        if not dicom_entries:\n            continue\n\n        dicom_path = Path(dicom_entries[0].path)\n\n        dataset = pydicom.dcmread(\n            dicom_path,\n            stop_before_pixels=True,\n        )\n\n        dicom_header_rows.append({\n            \"split\": split_name,\n            ID_COLUMN: study_uid,\n            \"SeriesInstanceUID\": series_uid,\n            \"file_name\": dicom_path.name,\n            \"anatomical_plane\": row.Anatomical_Plane,\n            \"Modality\": dataset.get(\"Modality\"),\n            \"Rows\": dataset.get(\"Rows\"),\n            \"Columns\": dataset.get(\"Columns\"),\n            \"PixelSpacing\": str(dataset.get(\"PixelSpacing\")),\n            \"SliceThickness\": dataset.get(\"SliceThickness\"),\n            \"SpacingBetweenSlices\": dataset.get(\n                \"SpacingBetweenSlices\"\n            ),\n            \"InstanceNumber\": dataset.get(\"InstanceNumber\"),\n            \"PhotometricInterpretation\": dataset.get(\n                \"PhotometricInterpretation\"\n            ),\n            \"BitsAllocated\": dataset.get(\"BitsAllocated\"),\n            \"study_uid_matches_directory\": (\n                dataset.get(\"StudyInstanceUID\") == study_uid\n            ),\n            \"series_uid_matches_directory\": (\n                dataset.get(\"SeriesInstanceUID\") == series_uid\n            ),\n        })\n\ndicom_header_sample = pd.DataFrame(dicom_header_rows)\n\nprint(f\"pydicom version: {pydicom.__version__}\")\n\ndisplay(dicom_header_sample)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:16.192937Z","iopub.execute_input":"2026-09-19T13:56:16.193305Z","iopub.status.idle":"2026-09-19T13:56:17.629661Z","shell.execute_reply.started":"2026-09-19T13:56:16.193259Z","shell.execute_reply":"2026-09-19T13:56:17.628811Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — DICOM identifiers are consistent, but image geometry varies\n\nAll sampled files are genuine MR DICOM images with `MONOCHROME2` photometric interpretation and 16-bit pixel storage. The `StudyInstanceUID` and `SeriesInstanceUID` values inside the files agree with the directory structure, which provides an additional integrity check.\n\nThe image geometry is not uniform. In the sample, image dimensions range from 320×320 to 960×960, with one test series measuring 640×1,280. Pixel spacing ranges from approximately 0.156 to 0.500 mm, while slice thickness ranges from 0.6 to 5.0 mm.\n\nThis variation will matter for preprocessing. A model may need consistent resizing and possibly spatial resampling, and slices should be ordered using DICOM metadata rather than filenames. The missing `SpacingBetweenSlices` values in some test files are not automatically an error because this tag is optional; image positions may provide a better way to determine slice spacing.","metadata":{}},{"cell_type":"markdown","source":"## Visual MRI sanity check\n\nWe now display one representative series from each anatomical plane for a single training study. The slices are ordered using their DICOM `InstanceNumber` values rather than relying on file-name order.\n\nThe contrast adjustment is used only to make the images easier to view. It is not a final preprocessing method.","metadata":{}},{"cell_type":"code","source":"# 37. Display one ordered middle slice per anatomical plane\n\ndef get_ordered_dicom_paths(study_uid, series_uid, image_root):\n    series_path = image_root / study_uid / series_uid\n\n    dicom_paths = sorted(\n        [\n            Path(entry.path)\n            for entry in os.scandir(series_path)\n            if (\n                entry.is_file(follow_symlinks=False)\n                and Path(entry.name).suffix.lower() == \".dcm\"\n            )\n        ],\n        key=lambda path: path.name,\n    )\n\n    ordered_paths = []\n\n    for path in dicom_paths:\n        header = pydicom.dcmread(\n            path,\n            stop_before_pixels=True,\n            specific_tags=[\"InstanceNumber\"],\n        )\n\n        instance_number = header.get(\"InstanceNumber\")\n\n        try:\n            instance_number = float(instance_number)\n        except (TypeError, ValueError):\n            instance_number = np.nan\n\n        ordered_paths.append({\n            \"path\": path,\n            \"instance_number\": instance_number,\n        })\n\n    return sorted(\n        ordered_paths,\n        key=lambda item: (\n            pd.isna(item[\"instance_number\"]),\n            item[\"instance_number\"]\n            if not pd.isna(item[\"instance_number\"])\n            else 0,\n            item[\"path\"].name,\n        ),\n    )\n\n\ndef load_display_image(path):\n    dataset = pydicom.dcmread(path)\n    image = dataset.pixel_array.astype(np.float32)\n\n    slope = float(dataset.get(\"RescaleSlope\", 1.0))\n    intercept = float(dataset.get(\"RescaleIntercept\", 0.0))\n\n    image = image * slope + intercept\n\n    lower, upper = np.percentile(image, [1, 99])\n\n    if upper > lower:\n        image = np.clip(image, lower, upper)\n        image = (image - lower) / (upper - lower)\n\n    return image, dataset\n\n\nsample_study_uid = train_series[ID_COLUMN].iloc[0]\nstudy_series = train_series.loc[\n    train_series[ID_COLUMN] == sample_study_uid\n]\n\nplane_order = [\"Axial\", \"Coronal\", \"Sagittal\"]\n\nfig, axes = plt.subplots(\n    1,\n    len(plane_order),\n    figsize=(15, 5),\n)\n\nfor ax, plane in zip(axes, plane_order):\n    plane_rows = study_series.loc[\n        study_series[\"Anatomical_Plane\"] == plane\n    ]\n\n    if plane_rows.empty:\n        ax.set_title(f\"{plane}\\nnot available\")\n        ax.axis(\"off\")\n        continue\n\n    series_uid = plane_rows[\"SeriesInstanceUID\"].iloc[0]\n\n    ordered_files = get_ordered_dicom_paths(\n        sample_study_uid,\n        series_uid,\n        TRAIN_IMAGES_ROOT,\n    )\n\n    middle_index = len(ordered_files) // 2\n    middle_file = ordered_files[middle_index]\n\n    image, dataset = load_display_image(\n        middle_file[\"path\"]\n    )\n\n    ax.imshow(image, cmap=\"gray\")\n    ax.set_title(\n        f\"{plane}\\n\"\n        f\"{len(ordered_files)} slices, \"\n        f\"shape {image.shape}\"\n    )\n    ax.axis(\"off\")\n\nfig.suptitle(\n    f\"Study {sample_study_uid}\",\n    fontsize=12,\n)\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:17.630724Z","iopub.execute_input":"2026-09-19T13:56:17.631117Z","iopub.status.idle":"2026-09-19T13:56:19.26005Z","shell.execute_reply.started":"2026-09-19T13:56:17.631089Z","shell.execute_reply":"2026-09-19T13:56:19.258961Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — the image data contain plausible multi-plane knee MRI\n\nThe three displayed series have the expected anatomical orientations and clearly show knee anatomy. The number of slices and image dimensions differ between planes, confirming that the series are not identical copies of one another.\n\nThis is a visual sanity check rather than a diagnostic interpretation. Image contrast also varies between sequences, so the display window used here is only for visualisation and should not yet be treated as the final model preprocessing method.","metadata":{}},{"cell_type":"markdown","source":"## Geometry variation across series\n\nMRI series may differ in image dimensions, pixel spacing and slice thickness. These differences determine whether images can be resized directly or whether spatial resampling will be needed before modelling.\n\nWe inspect a sample of series rather than reading every DICOM file.","metadata":{}},{"cell_type":"code","source":"# 38. Audit DICOM geometry across a sample of series\n\nGEOMETRY_SAMPLE_SIZE = 200\n\ngeometry_series_sample = train_series.sample(\n    n=min(GEOMETRY_SAMPLE_SIZE, len(train_series)),\n    random_state=RANDOM_STATE,\n)\n\ngeometry_rows = []\n\nfor row in geometry_series_sample[\n    [ID_COLUMN, \"SeriesInstanceUID\", \"Anatomical_Plane\"]\n].itertuples(index=False):\n\n    study_uid = getattr(row, ID_COLUMN)\n    series_uid = row.SeriesInstanceUID\n    series_path = TRAIN_IMAGES_ROOT / study_uid / series_uid\n\n    dicom_paths = sorted(\n        [\n            Path(entry.path)\n            for entry in os.scandir(series_path)\n            if (\n                entry.is_file(follow_symlinks=False)\n                and Path(entry.name).suffix.lower() == \".dcm\"\n            )\n        ],\n        key=lambda path: path.name,\n    )\n\n    if not dicom_paths:\n        geometry_rows.append({\n            \"StudyInstanceUID\": study_uid,\n            \"SeriesInstanceUID\": series_uid,\n            \"anatomical_plane\": row.Anatomical_Plane,\n            \"error\": \"No DICOM files found\",\n        })\n        continue\n\n    try:\n        dataset = pydicom.dcmread(\n            dicom_paths[0],\n            stop_before_pixels=True,\n        )\n\n        pixel_spacing = dataset.get(\n            \"PixelSpacing\",\n            [np.nan, np.nan],\n        )\n\n        try:\n            pixel_spacing_row = float(pixel_spacing[0])\n            pixel_spacing_col = float(pixel_spacing[1])\n        except (TypeError, ValueError, IndexError):\n            pixel_spacing_row = np.nan\n            pixel_spacing_col = np.nan\n\n        geometry_rows.append({\n            \"StudyInstanceUID\": study_uid,\n            \"SeriesInstanceUID\": series_uid,\n            \"anatomical_plane\": row.Anatomical_Plane,\n            \"rows\": dataset.get(\"Rows\", np.nan),\n            \"columns\": dataset.get(\"Columns\", np.nan),\n            \"pixel_spacing_row_mm\": pixel_spacing_row,\n            \"pixel_spacing_col_mm\": pixel_spacing_col,\n            \"slice_thickness_mm\": dataset.get(\n                \"SliceThickness\",\n                np.nan,\n            ),\n            \"spacing_between_slices_mm\": dataset.get(\n                \"SpacingBetweenSlices\",\n                np.nan,\n            ),\n            \"bits_stored\": dataset.get(\n                \"BitsStored\",\n                np.nan,\n            ),\n            \"num_slices_in_series\": len(dicom_paths),\n        })\n\n    except Exception as error:\n        geometry_rows.append({\n            \"StudyInstanceUID\": study_uid,\n            \"SeriesInstanceUID\": series_uid,\n            \"anatomical_plane\": row.Anatomical_Plane,\n            \"error\": str(error),\n        })\n\ngeometry_audit = pd.DataFrame(geometry_rows)\n\nif \"error\" in geometry_audit.columns:\n    error_mask = geometry_audit[\"error\"].notna()\nelse:\n    error_mask = pd.Series(\n        False,\n        index=geometry_audit.index,\n    )\n\nvalid_geometry = geometry_audit.loc[\n    ~error_mask\n].copy()\n\nnumber_of_errors = int(error_mask.sum())\n\nnumeric_geometry_columns = [\n    \"rows\",\n    \"columns\",\n    \"pixel_spacing_row_mm\",\n    \"pixel_spacing_col_mm\",\n    \"slice_thickness_mm\",\n    \"spacing_between_slices_mm\",\n    \"bits_stored\",\n    \"num_slices_in_series\",\n]\n\nfor column in numeric_geometry_columns:\n    if column in valid_geometry.columns:\n        valid_geometry[column] = pd.to_numeric(\n            valid_geometry[column],\n            errors=\"coerce\",\n        )\n\nprint(\n    f\"Successfully inspected \"\n    f\"{len(valid_geometry)} \"\n    f\"of {len(geometry_audit)} sampled series.\"\n)\nprint(f\"Series with header-reading errors: {number_of_errors}\")\n\ngeometry_summary = (\n    valid_geometry\n    .groupby(\"anatomical_plane\")[\n        [\n            \"rows\",\n            \"columns\",\n            \"pixel_spacing_row_mm\",\n            \"pixel_spacing_col_mm\",\n            \"slice_thickness_mm\",\n            \"num_slices_in_series\",\n        ]\n    ]\n    .agg([\"median\", \"min\", \"max\"])\n)\n\ndisplay(geometry_summary)\n\nfig, axes = plt.subplots(1, 2, figsize=(12, 5))\n\nsns.boxplot(\n    data=valid_geometry,\n    x=\"anatomical_plane\",\n    y=\"pixel_spacing_row_mm\",\n    color=\"steelblue\",\n    ax=axes[0],\n)\n\naxes[0].set_title(\"Pixel spacing by anatomical plane\")\naxes[0].set_xlabel(\"Anatomical plane\")\naxes[0].set_ylabel(\"Pixel spacing (mm)\")\n\nsns.boxplot(\n    data=valid_geometry,\n    x=\"anatomical_plane\",\n    y=\"slice_thickness_mm\",\n    color=\"steelblue\",\n    ax=axes[1],\n)\n\naxes[1].set_title(\"Slice thickness by anatomical plane\")\naxes[1].set_xlabel(\"Anatomical plane\")\naxes[1].set_ylabel(\"Slice thickness (mm)\")\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:19.261222Z","iopub.execute_input":"2026-09-19T13:56:19.261581Z","iopub.status.idle":"2026-09-19T13:56:24.373097Z","shell.execute_reply.started":"2026-09-19T13:56:19.261545Z","shell.execute_reply":"2026-09-19T13:56:24.372232Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — MRI geometry varies substantially\n\nAll 200 sampled series were read successfully, with no DICOM header errors.\n\nThe image geometry varies considerably:\n\n- image dimensions range from 256 pixels to as many as 1,280 pixels in an axis;\n- median images are approximately 512×512 pixels;\n- pixel spacing ranges from about 0.137 mm to 0.781 mm;\n- median slice thickness is approximately 3 mm, but ranges from roughly 0.6 mm to 5 mm;\n- sagittal series show a particularly long slice-count tail, reaching 320 slices in the sample.\n\nThis means that resizing every image to the same pixel dimensions will not necessarily make the physical anatomy comparable. Any future preprocessing should consider pixel spacing and slice geometry, rather than treating the raw pixel grid as physically uniform.\n\nThese results come from a sample of 200 series, so the exact minimum and maximum values should not be treated as full-dataset limits.","metadata":{}},{"cell_type":"markdown","source":"## Gold-labelled visual example\n\nWe now display one English-language gold-labelled study with its positive targets, report excerpt and representative MRI views.\n\nThis is a qualitative illustration only. The labels apply to the entire study, and a single middle slice from each series will not necessarily show every abnormality.","metadata":{}},{"cell_type":"code","source":"# 39. Display a gold-labelled study with targets and MRI views\n\nTARGET_TO_ILLUSTRATE = \"Effusion\"\n\ngold_english_candidates = gold_report_data.loc[\n    (gold_report_data[\"language\"] == \"English\")\n    & (gold_report_data[TARGET_TO_ILLUSTRATE] == 1),\n    ID_COLUMN,\n]\n\nexample_study_uid = gold_english_candidates.sample(\n    n=1,\n    random_state=RANDOM_STATE,\n).iloc[0]\n\nexample_row = gold_report_data.loc[\n    gold_report_data[ID_COLUMN] == example_study_uid\n].iloc[0]\n\nprint(f\"Study: {example_study_uid}\")\nprint(f\"Language: {example_row['language']}\")\nprint(f\"Positive targets: {example_row['positive_targets']}\")\nprint()\nprint(\"Report excerpt:\")\nprint(create_report_excerpt(example_row[\"Report\"], 600))\n\nstudy_series = train_series.loc[\n    train_series[ID_COLUMN] == example_study_uid\n]\n\nfig, axes = plt.subplots(\n    1,\n    len(plane_order),\n    figsize=(15, 5),\n)\n\nfor ax, plane in zip(axes, plane_order):\n    plane_rows = study_series.loc[\n        study_series[\"Anatomical_Plane\"] == plane\n    ]\n\n    if plane_rows.empty:\n        ax.set_title(f\"{plane}\\nnot available\")\n        ax.axis(\"off\")\n        continue\n\n    series_uid = plane_rows[\"SeriesInstanceUID\"].iloc[0]\n\n    ordered_files = get_ordered_dicom_paths(\n        example_study_uid,\n        series_uid,\n        TRAIN_IMAGES_ROOT,\n    )\n\n    middle_index = len(ordered_files) // 2\n    middle_file = ordered_files[middle_index]\n\n    image, dataset = load_display_image(\n        middle_file[\"path\"]\n    )\n\n    ax.imshow(image, cmap=\"gray\")\n    ax.set_title(\n        f\"{plane}\\n\"\n        f\"{len(ordered_files)} slices, \"\n        f\"shape {image.shape}\"\n    )\n    ax.axis(\"off\")\n\nfig.suptitle(\n    f\"Gold-labelled study: {example_study_uid}\",\n    fontsize=12,\n)\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:24.374226Z","iopub.execute_input":"2026-09-19T13:56:24.374545Z","iopub.status.idle":"2026-09-19T13:56:26.708069Z","shell.execute_reply.started":"2026-09-19T13:56:24.374516Z","shell.execute_reply":"2026-09-19T13:56:26.707143Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — a single study can contain many co-existing abnormalities\n\nThis gold-labelled study has six positive targets. The report explicitly describes a fracture, meniscal injuries and a large joint effusion, alongside additional ligament and soft-tissue findings.\n\nThe three MRI views show a plausible knee examination with different anatomical perspectives and acquisition resolutions. However, these are representative middle slices, so they cannot independently confirm every study-level label. The example mainly demonstrates that the task is multi-label and study-level: abnormalities may be distributed across different slices and planes.","metadata":{}},{"cell_type":"markdown","source":"## Slice-order and image-identifier integrity\n\nDICOM series contain several identifiers that help order and distinguish individual slices. We check a sample of series for missing or duplicated `InstanceNumber`, `SOPInstanceUID` and spatial-position values.\n\nMissing `ImagePositionPatient` values are not automatically an error because that DICOM field may be optional.","metadata":{}},{"cell_type":"code","source":"# 40. Audit slice identifiers in a sample of series\n\nSLICE_AUDIT_SAMPLE_SIZE = 50\n\nslice_audit_sample = train_series.sample(\n    n=min(SLICE_AUDIT_SAMPLE_SIZE, len(train_series)),\n    random_state=RANDOM_STATE,\n)\n\nslice_audit_rows = []\n\nfor row in slice_audit_sample[\n    [ID_COLUMN, \"SeriesInstanceUID\", \"Anatomical_Plane\"]\n].itertuples(index=False):\n\n    study_uid = getattr(row, ID_COLUMN)\n    series_uid = row.SeriesInstanceUID\n    series_path = TRAIN_IMAGES_ROOT / study_uid / series_uid\n\n    dicom_paths = sorted(\n        [\n            Path(entry.path)\n            for entry in os.scandir(series_path)\n            if (\n                entry.is_file(follow_symlinks=False)\n                and Path(entry.name).suffix.lower() == \".dcm\"\n            )\n        ],\n        key=lambda path: path.name,\n    )\n\n    instance_numbers = []\n    sop_instance_uids = []\n    image_positions = []\n\n    for dicom_path in dicom_paths:\n        dataset = pydicom.dcmread(\n            dicom_path,\n            stop_before_pixels=True,\n            specific_tags=[\n                \"InstanceNumber\",\n                \"SOPInstanceUID\",\n                \"ImagePositionPatient\",\n            ],\n        )\n\n        instance_number = dataset.get(\"InstanceNumber\")\n        sop_instance_uid = dataset.get(\"SOPInstanceUID\")\n        image_position = dataset.get(\"ImagePositionPatient\")\n\n        if instance_number is not None:\n            instance_numbers.append(str(instance_number))\n\n        if sop_instance_uid is not None:\n            sop_instance_uids.append(str(sop_instance_uid))\n\n        if image_position is not None:\n            try:\n                image_positions.append(\n                    tuple(\n                        round(float(value), 5)\n                        for value in image_position\n                    )\n                )\n            except (TypeError, ValueError):\n                pass\n\n    slice_audit_rows.append({\n        \"StudyInstanceUID\": study_uid,\n        \"SeriesInstanceUID\": series_uid,\n        \"anatomical_plane\": row.Anatomical_Plane,\n        \"number_of_files\": len(dicom_paths),\n        \"missing_instance_numbers\": (\n            len(instance_numbers) < len(dicom_paths)\n        ),\n        \"duplicate_instance_numbers\": (\n            len(instance_numbers) != len(set(instance_numbers))\n        ),\n        \"missing_sop_instance_uids\": (\n            len(sop_instance_uids) < len(dicom_paths)\n        ),\n        \"duplicate_sop_instance_uids\": (\n            len(sop_instance_uids) != len(set(sop_instance_uids))\n        ),\n        \"number_of_position_values\": len(image_positions),\n        \"duplicate_image_positions\": (\n            len(image_positions) != len(set(image_positions))\n            if image_positions\n            else False\n        ),\n    })\n\nslice_audit = pd.DataFrame(slice_audit_rows)\n\nslice_audit_summary = pd.DataFrame({\n    \"check\": [\n        \"Series with missing InstanceNumber values\",\n        \"Series with duplicate InstanceNumber values\",\n        \"Series with missing SOPInstanceUID values\",\n        \"Series with duplicate SOPInstanceUID values\",\n        \"Series with duplicate image positions\",\n    ],\n    \"series_affected\": [\n        slice_audit[\"missing_instance_numbers\"].sum(),\n        slice_audit[\"duplicate_instance_numbers\"].sum(),\n        slice_audit[\"missing_sop_instance_uids\"].sum(),\n        slice_audit[\"duplicate_sop_instance_uids\"].sum(),\n        slice_audit[\"duplicate_image_positions\"].sum(),\n    ],\n})\n\ndisplay(slice_audit_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:26.709295Z","iopub.execute_input":"2026-09-19T13:56:26.709601Z","iopub.status.idle":"2026-09-19T13:56:43.694369Z","shell.execute_reply.started":"2026-09-19T13:56:26.709574Z","shell.execute_reply":"2026-09-19T13:56:43.693514Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — sampled slice identifiers are valid\n\nThe sampled series contain no missing or duplicate `InstanceNumber` or `SOPInstanceUID` values, and no duplicate image positions were detected. These results support using DICOM metadata when ordering slices in the sampled data.","metadata":{}},{"cell_type":"markdown","source":"## Pixel data integrity","metadata":{}},{"cell_type":"code","source":"# 41. Check representative pixel decoding\n\nPIXEL_AUDIT_SAMPLE_SIZE = 100\n\npixel_audit_sample = train_series.sample(\n    n=min(PIXEL_AUDIT_SAMPLE_SIZE, len(train_series)),\n    random_state=RANDOM_STATE + 1,\n)\n\npixel_audit_rows = []\n\nfor row in pixel_audit_sample[\n    [ID_COLUMN, \"SeriesInstanceUID\", \"Anatomical_Plane\"]\n].itertuples(index=False):\n\n    study_uid = getattr(row, ID_COLUMN)\n    series_uid = row.SeriesInstanceUID\n    series_path = TRAIN_IMAGES_ROOT / study_uid / series_uid\n\n    try:\n        dicom_paths = sorted(\n            [\n                Path(entry.path)\n                for entry in os.scandir(series_path)\n                if (\n                    entry.is_file(follow_symlinks=False)\n                    and Path(entry.name).suffix.lower() == \".dcm\"\n                )\n            ],\n            key=lambda path: path.name,\n        )\n\n        if not dicom_paths:\n            raise FileNotFoundError(\"No DICOM files found\")\n\n        representative_path = dicom_paths[len(dicom_paths) // 2]\n\n        dataset = pydicom.dcmread(representative_path)\n        pixel_array = dataset.pixel_array\n\n        expected_shape = (\n            int(dataset.Rows),\n            int(dataset.Columns),\n        )\n        actual_shape = tuple(pixel_array.shape[-2:])\n\n        pixel_audit_rows.append({\n            \"StudyInstanceUID\": study_uid,\n            \"SeriesInstanceUID\": series_uid,\n            \"anatomical_plane\": row.Anatomical_Plane,\n            \"expected_shape\": expected_shape,\n            \"actual_shape\": actual_shape,\n            \"shape_matches\": expected_shape == actual_shape,\n            \"has_non_finite_values\": not np.isfinite(pixel_array).all(),\n            \"is_constant_image\": np.ptp(pixel_array) == 0,\n            \"error\": None,\n        })\n\n    except Exception as error:\n        pixel_audit_rows.append({\n            \"StudyInstanceUID\": study_uid,\n            \"SeriesInstanceUID\": series_uid,\n            \"anatomical_plane\": row.Anatomical_Plane,\n            \"expected_shape\": None,\n            \"actual_shape\": None,\n            \"shape_matches\": False,\n            \"has_non_finite_values\": False,\n            \"is_constant_image\": False,\n            \"error\": f\"{type(error).__name__}: {error}\",\n        })\n\npixel_audit = pd.DataFrame(pixel_audit_rows)\n\npixel_audit_summary = pd.DataFrame({\n    \"check\": [\n        \"Sampled series\",\n        \"Representative pixel arrays decoded\",\n        \"Header/image shape mismatches\",\n        \"Images with non-finite values\",\n        \"Constant representative images\",\n        \"Series with decoding errors\",\n    ],\n    \"series_affected\": [\n        len(pixel_audit),\n        pixel_audit[\"error\"].isna().sum(),\n        (~pixel_audit[\"shape_matches\"]).sum(),\n        pixel_audit[\"has_non_finite_values\"].sum(),\n        pixel_audit[\"is_constant_image\"].sum(),\n        pixel_audit[\"error\"].notna().sum(),\n    ],\n})\n\ndisplay(pixel_audit_summary)\n\nif pixel_audit[\"error\"].notna().any():\n    display(\n        pixel_audit.loc[\n            pixel_audit[\"error\"].notna(),\n            [\n                \"StudyInstanceUID\",\n                \"SeriesInstanceUID\",\n                \"anatomical_plane\",\n                \"error\",\n            ],\n        ]\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:43.69536Z","iopub.execute_input":"2026-09-19T13:56:43.695662Z","iopub.status.idle":"2026-09-19T13:56:45.851782Z","shell.execute_reply.started":"2026-09-19T13:56:43.695627Z","shell.execute_reply":"2026-09-19T13:56:45.85095Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — sampled pixel data are readable\n\nAll 100 sampled series successfully decoded representative pixel arrays. The decoded image shapes matched the recorded DICOM dimensions, with no non-finite or constant images detected.\n\nThis indicates that the dataset has no immediate pixel-decoding or image-shape problems in the sample. Pixel intensities may still vary between scanners and sequences, so normalization should be examined before modelling.","metadata":{}},{"cell_type":"markdown","source":"## Pixel intensity variation","metadata":{}},{"cell_type":"code","source":"# 42. Audit representative pixel intensity ranges\n\nintensity_audit_rows = []\n\nfor row in pixel_audit_sample[\n    [ID_COLUMN, \"SeriesInstanceUID\", \"Anatomical_Plane\"]\n].itertuples(index=False):\n\n    study_uid = getattr(row, ID_COLUMN)\n    series_uid = row.SeriesInstanceUID\n    series_path = TRAIN_IMAGES_ROOT / study_uid / series_uid\n\n    try:\n        dicom_paths = sorted(\n            [\n                Path(entry.path)\n                for entry in os.scandir(series_path)\n                if (\n                    entry.is_file(follow_symlinks=False)\n                    and Path(entry.name).suffix.lower() == \".dcm\"\n                )\n            ],\n            key=lambda path: path.name,\n        )\n\n        representative_path = dicom_paths[len(dicom_paths) // 2]\n        dataset = pydicom.dcmread(representative_path)\n\n        pixel_array = dataset.pixel_array.astype(np.float32)\n\n        rescale_slope = float(dataset.get(\"RescaleSlope\", 1.0))\n        rescale_intercept = float(dataset.get(\"RescaleIntercept\", 0.0))\n\n        intensity_array = (\n            pixel_array * rescale_slope\n            + rescale_intercept\n        )\n\n        p01, median, p99 = np.percentile(\n            intensity_array,\n            [1, 50, 99],\n        )\n\n        intensity_audit_rows.append({\n            \"StudyInstanceUID\": study_uid,\n            \"SeriesInstanceUID\": series_uid,\n            \"anatomical_plane\": row.Anatomical_Plane,\n            \"p01\": p01,\n            \"median\": median,\n            \"p99\": p99,\n            \"display_range\": p99 - p01,\n            \"rescale_slope\": rescale_slope,\n            \"rescale_intercept\": rescale_intercept,\n            \"error\": None,\n        })\n\n    except Exception as error:\n        intensity_audit_rows.append({\n            \"StudyInstanceUID\": study_uid,\n            \"SeriesInstanceUID\": series_uid,\n            \"anatomical_plane\": row.Anatomical_Plane,\n            \"p01\": np.nan,\n            \"median\": np.nan,\n            \"p99\": np.nan,\n            \"display_range\": np.nan,\n            \"rescale_slope\": np.nan,\n            \"rescale_intercept\": np.nan,\n            \"error\": f\"{type(error).__name__}: {error}\",\n        })\n\nintensity_audit = pd.DataFrame(intensity_audit_rows)\n\nvalid_intensity_audit = intensity_audit[\n    intensity_audit[\"error\"].isna()\n].copy()\n\nintensity_summary = (\n    valid_intensity_audit\n    .groupby(\"anatomical_plane\")\n    .agg(\n        series_sampled=(\"SeriesInstanceUID\", \"count\"),\n        median_p01=(\"p01\", \"median\"),\n        median_pixel_value=(\"median\", \"median\"),\n        median_p99=(\"p99\", \"median\"),\n        median_display_range=(\"display_range\", \"median\"),\n        minimum_display_range=(\"display_range\", \"min\"),\n        maximum_display_range=(\"display_range\", \"max\"),\n    )\n    .round(3)\n    .reset_index()\n)\n\ndisplay(intensity_summary)\n\nif intensity_audit[\"error\"].notna().any():\n    display(\n        intensity_audit.loc[\n            intensity_audit[\"error\"].notna(),\n            [\n                \"StudyInstanceUID\",\n                \"SeriesInstanceUID\",\n                \"anatomical_plane\",\n                \"error\",\n            ],\n        ]\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:45.85291Z","iopub.execute_input":"2026-09-19T13:56:45.853232Z","iopub.status.idle":"2026-09-19T13:56:46.703084Z","shell.execute_reply.started":"2026-09-19T13:56:45.853199Z","shell.execute_reply":"2026-09-19T13:56:46.702301Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — pixel intensity ranges vary substantially\n\nThe representative images have widely varying intensity ranges, even within the same anatomical plane. Their lower percentile is usually zero because of the background, while the upper percentile and overall display range vary considerably between series.\n\nThis suggests that raw pixel values should not be treated as directly comparable across the dataset. Robust clipping and normalization will likely be needed during preprocessing. We will next check whether the recorded sequence type explains part of this variation.","metadata":{}},{"cell_type":"code","source":"# 43. Compare intensity ranges by sequence type\n\nintensity_with_metadata = intensity_audit.merge(\n    train_series[\n        [\"SeriesInstanceUID\", \"Fluid_Sensitive\"]\n    ],\n    on=\"SeriesInstanceUID\",\n    how=\"left\",\n)\n\nintensity_with_metadata[\"sequence_type\"] = (\n    intensity_with_metadata[\"Fluid_Sensitive\"]\n    .map({\n        0: \"Not fluid-sensitive\",\n        1: \"Fluid-sensitive\",\n    })\n)\n\nvalid_intensity_metadata = intensity_with_metadata[\n    intensity_with_metadata[\"error\"].isna()\n].copy()\n\nintensity_by_sequence = (\n    valid_intensity_metadata\n    .groupby(\n        [\"anatomical_plane\", \"sequence_type\"],\n        dropna=False,\n    )\n    .agg(\n        series_sampled=(\"SeriesInstanceUID\", \"count\"),\n        median_p01=(\"p01\", \"median\"),\n        median_pixel_value=(\"median\", \"median\"),\n        median_p99=(\"p99\", \"median\"),\n        median_display_range=(\"display_range\", \"median\"),\n        minimum_display_range=(\"display_range\", \"min\"),\n        maximum_display_range=(\"display_range\", \"max\"),\n    )\n    .round(3)\n    .reset_index()\n)\n\ndisplay(intensity_by_sequence)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:46.704137Z","iopub.execute_input":"2026-09-19T13:56:46.704448Z","iopub.status.idle":"2026-09-19T13:56:46.738376Z","shell.execute_reply.started":"2026-09-19T13:56:46.704422Z","shell.execute_reply":"2026-09-19T13:56:46.737415Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — sequence type contributes to intensity variation\n\nIntensity ranges differ between fluid-sensitive and non-fluid-sensitive series, particularly for coronal and sagittal images. However, substantial variation remains within each sequence group, so sequence type does not fully explain the differences.\n\nThe safest preprocessing approach will therefore use robust image normalization, while retaining sequence type as potentially useful metadata.","metadata":{}},{"cell_type":"markdown","source":"## Label availability and imaging composition","metadata":{}},{"cell_type":"code","source":"# 44. Compare imaging composition by label availability\n\ngold_study_ids = set(\n    train.loc[\n        train[TARGET_COLUMNS].notna().all(axis=1),\n        ID_COLUMN,\n    ]\n)\n\nstudy_imaging_summary = (\n    train_series\n    .groupby(ID_COLUMN)\n    .agg(\n        number_of_series=(\"SeriesInstanceUID\", \"nunique\"),\n        fluid_sensitive_series=(\"Fluid_Sensitive\", \"sum\"),\n        fluid_sensitive_percent=(\"Fluid_Sensitive\", \"mean\"),\n    )\n    .reset_index()\n)\n\nstudy_imaging_summary[\"fluid_sensitive_percent\"] *= 100\n\nstudy_imaging_summary[\"label_status\"] = np.where(\n    study_imaging_summary[ID_COLUMN].isin(gold_study_ids),\n    \"Gold-labelled\",\n    \"Report-only\",\n)\n\nimaging_composition_summary = (\n    study_imaging_summary\n    .groupby(\"label_status\")\n    .agg(\n        studies=(ID_COLUMN, \"count\"),\n        median_series=(\"number_of_series\", \"median\"),\n        mean_series=(\"number_of_series\", \"mean\"),\n        minimum_series=(\"number_of_series\", \"min\"),\n        maximum_series=(\"number_of_series\", \"max\"),\n        median_fluid_sensitive_percent=(\n            \"fluid_sensitive_percent\",\n            \"median\",\n        ),\n        mean_fluid_sensitive_percent=(\n            \"fluid_sensitive_percent\",\n            \"mean\",\n        ),\n    )\n    .round(2)\n    .reset_index()\n)\n\ndisplay(imaging_composition_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:46.740282Z","iopub.execute_input":"2026-09-19T13:56:46.740585Z","iopub.status.idle":"2026-09-19T13:56:46.783924Z","shell.execute_reply.started":"2026-09-19T13:56:46.740559Z","shell.execute_reply":"2026-09-19T13:56:46.782984Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — gold-labelled studies have similar imaging composition\n\nThe gold-labelled studies broadly resemble the report-only studies in terms of the number of MRI series and the proportion of fluid-sensitive series. This is encouraging because it suggests that the labelled subset is not obviously biased by these coarse acquisition characteristics.\n\nThis does not guarantee that the groups are comparable in image dimensions or spatial resolution, so we will compare those properties next.","metadata":{}},{"cell_type":"code","source":"# 45. Compare sampled DICOM geometry by label status\n\nGEOMETRY_LABEL_SAMPLE_SIZE = 50\n\ngold_series_sample = train_series[\n    train_series[ID_COLUMN].isin(gold_study_ids)\n].sample(\n    n=min(\n        GEOMETRY_LABEL_SAMPLE_SIZE,\n        train_series[train_series[ID_COLUMN].isin(gold_study_ids)].shape[0],\n    ),\n    random_state=RANDOM_STATE,\n)\n\nreport_only_series_sample = train_series[\n    ~train_series[ID_COLUMN].isin(gold_study_ids)\n].sample(\n    n=min(\n        GEOMETRY_LABEL_SAMPLE_SIZE,\n        train_series[~train_series[ID_COLUMN].isin(gold_study_ids)].shape[0],\n    ),\n    random_state=RANDOM_STATE,\n)\n\ngeometry_label_sample = pd.concat([\n    gold_series_sample.assign(label_status=\"Gold-labelled\"),\n    report_only_series_sample.assign(label_status=\"Report-only\"),\n])\n\ngeometry_label_rows = []\n\nfor row in geometry_label_sample[\n    [ID_COLUMN, \"SeriesInstanceUID\", \"Anatomical_Plane\", \"label_status\"]\n].itertuples(index=False):\n\n    study_uid = getattr(row, ID_COLUMN)\n    series_uid = row.SeriesInstanceUID\n    series_path = TRAIN_IMAGES_ROOT / study_uid / series_uid\n\n    try:\n        dicom_paths = sorted(\n            [\n                Path(entry.path)\n                for entry in os.scandir(series_path)\n                if (\n                    entry.is_file(follow_symlinks=False)\n                    and Path(entry.name).suffix.lower() == \".dcm\"\n                )\n            ],\n            key=lambda path: path.name,\n        )\n\n        if not dicom_paths:\n            raise FileNotFoundError(\"No DICOM files found\")\n\n        dataset = pydicom.dcmread(\n            dicom_paths[0],\n            stop_before_pixels=True,\n            specific_tags=[\n                \"Rows\",\n                \"Columns\",\n                \"PixelSpacing\",\n                \"SliceThickness\",\n            ],\n        )\n\n        pixel_spacing = dataset.get(\"PixelSpacing\")\n\n        geometry_label_rows.append({\n            \"label_status\": row.label_status,\n            \"anatomical_plane\": row.Anatomical_Plane,\n            \"rows\": dataset.get(\"Rows\"),\n            \"columns\": dataset.get(\"Columns\"),\n            \"pixel_spacing_row_mm\": (\n                float(pixel_spacing[0])\n                if pixel_spacing is not None\n                else np.nan\n            ),\n            \"pixel_spacing_col_mm\": (\n                float(pixel_spacing[1])\n                if pixel_spacing is not None\n                else np.nan\n            ),\n            \"slice_thickness_mm\": (\n                float(dataset.get(\"SliceThickness\"))\n                if dataset.get(\"SliceThickness\") is not None\n                else np.nan\n            ),\n            \"error\": None,\n        })\n\n    except Exception as error:\n        geometry_label_rows.append({\n            \"label_status\": row.label_status,\n            \"anatomical_plane\": row.Anatomical_Plane,\n            \"rows\": np.nan,\n            \"columns\": np.nan,\n            \"pixel_spacing_row_mm\": np.nan,\n            \"pixel_spacing_col_mm\": np.nan,\n            \"slice_thickness_mm\": np.nan,\n            \"error\": f\"{type(error).__name__}: {error}\",\n        })\n\ngeometry_label_audit = pd.DataFrame(geometry_label_rows)\n\ngeometry_label_summary = (\n    geometry_label_audit\n    .groupby(\"label_status\")\n    .agg(\n        series_sampled=(\"label_status\", \"size\"),\n        header_errors=(\"error\", lambda values: values.notna().sum()),\n        median_rows=(\"rows\", \"median\"),\n        median_columns=(\"columns\", \"median\"),\n        median_pixel_spacing_row_mm=(\n            \"pixel_spacing_row_mm\",\n            \"median\",\n        ),\n        median_pixel_spacing_col_mm=(\n            \"pixel_spacing_col_mm\",\n            \"median\",\n        ),\n        median_slice_thickness_mm=(\n            \"slice_thickness_mm\",\n            \"median\",\n        ),\n    )\n    .round(3)\n    .reset_index()\n)\n\ndisplay(geometry_label_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:46.785096Z","iopub.execute_input":"2026-09-19T13:56:46.78544Z","iopub.status.idle":"2026-09-19T13:56:48.406948Z","shell.execute_reply.started":"2026-09-19T13:56:46.785403Z","shell.execute_reply":"2026-09-19T13:56:48.406105Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — basic geometry is similar\n\nThe sampled gold-labelled and report-only series have the same typical image dimensions and slice thickness, with no DICOM header errors.\n\nThe gold-labelled sample has slightly larger median pixel spacing. Since spacing varies between anatomical planes, this difference may reflect the composition of the samples rather than a genuine difference between the labelled and unlabelled studies.","metadata":{}},{"cell_type":"code","source":"# 46. Compare DICOM geometry by label status and anatomical plane\n\ngeometry_by_plane_status = (\n    geometry_label_audit\n    .groupby(\n        [\"anatomical_plane\", \"label_status\"],\n        dropna=False,\n    )\n    .agg(\n        series_sampled=(\"label_status\", \"size\"),\n        median_rows=(\"rows\", \"median\"),\n        median_columns=(\"columns\", \"median\"),\n        median_pixel_spacing_row_mm=(\n            \"pixel_spacing_row_mm\",\n            \"median\",\n        ),\n        median_pixel_spacing_col_mm=(\n            \"pixel_spacing_col_mm\",\n            \"median\",\n        ),\n        median_slice_thickness_mm=(\n            \"slice_thickness_mm\",\n            \"median\",\n        ),\n    )\n    .round(3)\n    .reset_index()\n    .sort_values(\n        [\"anatomical_plane\", \"label_status\"]\n    )\n    .reset_index(drop=True)\n)\n\ndisplay(geometry_by_plane_status)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:48.408017Z","iopub.execute_input":"2026-09-19T13:56:48.408359Z","iopub.status.idle":"2026-09-19T13:56:48.431165Z","shell.execute_reply.started":"2026-09-19T13:56:48.408326Z","shell.execute_reply":"2026-09-19T13:56:48.430406Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — acquisition geometry differs across labelled and report-only samples\n\nWithin anatomical planes, the sampled gold-labelled and report-only series do not always have the same image dimensions or pixel spacing. The largest difference appears in sagittal series.\n\nThese results are descriptive because the samples are small, but they confirm substantial acquisition variability. Models will need consistent image preprocessing, and study-level validation will be important because the labelled subset may not perfectly represent the full dataset.","metadata":{}},{"cell_type":"markdown","source":"## Test image geometry","metadata":{}},{"cell_type":"code","source":"# 47. Inspect test-study image geometry\n\ntest_geometry_rows = []\n\nfor row in test_series[\n    [\n        ID_COLUMN,\n        \"SeriesInstanceUID\",\n        \"Anatomical_Plane\",\n        \"Fluid_Sensitive\",\n    ]\n].itertuples(index=False):\n\n    study_uid = getattr(row, ID_COLUMN)\n    series_uid = row.SeriesInstanceUID\n    series_path = TEST_IMAGES_ROOT / study_uid / series_uid\n\n    try:\n        dicom_paths = sorted(\n            [\n                Path(entry.path)\n                for entry in os.scandir(series_path)\n                if (\n                    entry.is_file(follow_symlinks=False)\n                    and Path(entry.name).suffix.lower() == \".dcm\"\n                )\n            ],\n            key=lambda path: path.name,\n        )\n\n        if not dicom_paths:\n            raise FileNotFoundError(\"No DICOM files found\")\n\n        dataset = pydicom.dcmread(\n            dicom_paths[0],\n            stop_before_pixels=True,\n            specific_tags=[\n                \"Rows\",\n                \"Columns\",\n                \"PixelSpacing\",\n                \"SliceThickness\",\n            ],\n        )\n\n        pixel_spacing = dataset.get(\"PixelSpacing\")\n\n        test_geometry_rows.append({\n            \"StudyInstanceUID\": study_uid,\n            \"SeriesInstanceUID\": series_uid,\n            \"anatomical_plane\": row.Anatomical_Plane,\n            \"fluid_sensitive\": row.Fluid_Sensitive,\n            \"number_of_files\": len(dicom_paths),\n            \"rows\": dataset.get(\"Rows\"),\n            \"columns\": dataset.get(\"Columns\"),\n            \"pixel_spacing_row_mm\": (\n                float(pixel_spacing[0])\n                if pixel_spacing is not None\n                else np.nan\n            ),\n            \"pixel_spacing_col_mm\": (\n                float(pixel_spacing[1])\n                if pixel_spacing is not None\n                else np.nan\n            ),\n            \"slice_thickness_mm\": (\n                float(dataset.get(\"SliceThickness\"))\n                if dataset.get(\"SliceThickness\") is not None\n                else np.nan\n            ),\n            \"error\": None,\n        })\n\n    except Exception as error:\n        test_geometry_rows.append({\n            \"StudyInstanceUID\": study_uid,\n            \"SeriesInstanceUID\": series_uid,\n            \"anatomical_plane\": row.Anatomical_Plane,\n            \"fluid_sensitive\": row.Fluid_Sensitive,\n            \"number_of_files\": np.nan,\n            \"rows\": np.nan,\n            \"columns\": np.nan,\n            \"pixel_spacing_row_mm\": np.nan,\n            \"pixel_spacing_col_mm\": np.nan,\n            \"slice_thickness_mm\": np.nan,\n            \"error\": f\"{type(error).__name__}: {error}\",\n        })\n\ntest_geometry = (\n    pd.DataFrame(test_geometry_rows)\n    .sort_values(\n        [\"StudyInstanceUID\", \"anatomical_plane\"]\n    )\n    .reset_index(drop=True)\n)\n\ndisplay(test_geometry)\n\nif test_geometry[\"error\"].notna().any():\n    display(\n        test_geometry.loc[\n            test_geometry[\"error\"].notna(),\n            [\n                \"StudyInstanceUID\",\n                \"SeriesInstanceUID\",\n                \"anatomical_plane\",\n                \"error\",\n            ],\n        ]\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:48.432303Z","iopub.execute_input":"2026-09-19T13:56:48.432712Z","iopub.status.idle":"2026-09-19T13:56:48.57628Z","shell.execute_reply.started":"2026-09-19T13:56:48.432684Z","shell.execute_reply":"2026-09-19T13:56:48.575431Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — test images are valid but heterogeneous\n\nAll test DICOM series were readable and contain the expected anatomical information. However, the test studies vary substantially in image dimensions, pixel spacing, slice thickness, and number of slices.\n\nThis reinforces the need for consistent resizing, intensity normalization, and controlled handling of variable-length series during modelling.","metadata":{}},{"cell_type":"markdown","source":"# Study-level modelling table","metadata":{}},{"cell_type":"code","source":"# 48. Assemble the fully labelled study-level table\n\ngold_study_table = (\n    train.loc[\n        train[ID_COLUMN].isin(gold_study_ids),\n        [ID_COLUMN] + TARGET_COLUMNS,\n    ]\n    .merge(\n        study_imaging_summary.drop(columns=\"label_status\"),\n        on=ID_COLUMN,\n        how=\"left\",\n        validate=\"one_to_one\",\n    )\n)\n\nprint(f\"Gold-labelled studies: {len(gold_study_table):,}\")\nprint(f\"Columns in study-level table: {gold_study_table.shape[1]}\")\n\ndisplay(gold_study_table.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:48.577483Z","iopub.execute_input":"2026-09-19T13:56:48.577852Z","iopub.status.idle":"2026-09-19T13:56:48.603593Z","shell.execute_reply.started":"2026-09-19T13:56:48.577814Z","shell.execute_reply":"2026-09-19T13:56:48.602653Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — studies are the correct modelling unit\n\nThe fully labelled dataset now has one row per study, combining the target labels with study-level imaging metadata.\n\nThis confirms that future validation must be performed at the study level rather than the series or image level. Because only 58 studies are labelled, the number of positive cases available for each target will strongly influence the choice of cross-validation strategy.","metadata":{}},{"cell_type":"markdown","source":"# Validation strategy","metadata":{}},{"cell_type":"code","source":"# 49. Assess validation feasibility for each target\n\nvalidation_feasibility = pd.DataFrame({\n    \"target\": TARGET_COLUMNS,\n    \"positive_studies\": [\n        int(gold_study_table[target].sum())\n        for target in TARGET_COLUMNS\n    ],\n})\n\nvalidation_feasibility[\"negative_studies\"] = (\n    len(gold_study_table)\n    - validation_feasibility[\"positive_studies\"]\n)\n\nvalidation_feasibility[\"average_positives_per_3_fold\"] = (\n    validation_feasibility[\"positive_studies\"] / 3\n).round(2)\n\nvalidation_feasibility[\"average_positives_per_5_fold\"] = (\n    validation_feasibility[\"positive_studies\"] / 5\n).round(2)\n\nvalidation_feasibility = (\n    validation_feasibility\n    .sort_values(\"positive_studies\")\n    .reset_index(drop=True)\n)\n\ndisplay(validation_feasibility)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:48.604735Z","iopub.execute_input":"2026-09-19T13:56:48.605082Z","iopub.status.idle":"2026-09-19T13:56:48.620231Z","shell.execute_reply.started":"2026-09-19T13:56:48.605047Z","shell.execute_reply":"2026-09-19T13:56:48.619505Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — validation is limited by rare positive targets\n\nThe 58 gold-labelled studies are sufficient for an initial validation scheme, but the rare targets provide very few positive cases per fold. This is especially true for MCL, Lateral OA and Baker's.\n\nThree-fold validation should provide more stable estimates than five-fold validation, although even three-fold estimates will remain uncertain. Because each study has multiple targets, the folds should be created using multi-label stratification rather than balancing only one target or the total number of abnormalities.","metadata":{}},{"cell_type":"code","source":"# 50. Inspect multi-label validation fold balance\n\ntry:\n    from iterstrat.ml_stratifiers import MultilabelStratifiedKFold\nexcept ImportError:\n    %pip install -q iterative-stratification\n    from iterstrat.ml_stratifiers import MultilabelStratifiedKFold\n\ngold_target_matrix = (\n    gold_study_table[TARGET_COLUMNS]\n    .to_numpy(dtype=int)\n)\n\nvalidation_assignments = gold_study_table[\n    [ID_COLUMN]\n].copy()\n\nfold_balance_rows = []\n\nfor number_of_folds in [3, 5]:\n\n    splitter = MultilabelStratifiedKFold(\n        n_splits=number_of_folds,\n        shuffle=True,\n        random_state=RANDOM_STATE,\n    )\n\n    fold_numbers = np.full(\n        len(gold_study_table),\n        fill_value=-1,\n        dtype=int,\n    )\n\n    for fold, (_, validation_indices) in enumerate(\n        splitter.split(\n            np.zeros((len(gold_target_matrix), 1)),\n            gold_target_matrix,\n        )\n    ):\n        fold_numbers[validation_indices] = fold\n\n    validation_assignments[\n        f\"fold_{number_of_folds}\"\n    ] = fold_numbers\n\n    for target in TARGET_COLUMNS:\n        positive_counts = [\n            int(\n                gold_study_table.loc[\n                    fold_numbers == fold,\n                    target,\n                ].sum()\n            )\n            for fold in range(number_of_folds)\n        ]\n\n        negative_counts = [\n            int(\n                (\n                    gold_study_table.loc[\n                        fold_numbers == fold,\n                        target,\n                    ] == 0\n                ).sum()\n            )\n            for fold in range(number_of_folds)\n        ]\n\n        fold_balance_rows.append({\n            \"number_of_folds\": number_of_folds,\n            \"target\": target,\n            \"positive_cases_by_fold\": positive_counts,\n            \"minimum_positive_cases\": min(positive_counts),\n            \"maximum_positive_cases\": max(positive_counts),\n            \"negative_cases_by_fold\": negative_counts,\n        })\n\nfold_balance_summary = (\n    pd.DataFrame(fold_balance_rows)\n    .sort_values(\n        [\"number_of_folds\", \"minimum_positive_cases\", \"target\"]\n    )\n    .reset_index(drop=True)\n)\n\ndisplay(fold_balance_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:56:48.621366Z","iopub.execute_input":"2026-09-19T13:56:48.621726Z","iopub.status.idle":"2026-09-19T13:56:48.700144Z","shell.execute_reply.started":"2026-09-19T13:56:48.62169Z","shell.execute_reply":"2026-09-19T13:56:48.699352Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — three-fold validation is the safer primary split\n\nThe multi-label stratification produces well-balanced three-fold splits. Every target has positive and negative cases in each fold, including the rare MCL target.\n\nThe five-fold split is less suitable as the primary evaluation because some validation folds contain only one positive MCL case or two negative Effusion cases. Three-fold validation will therefore be used for the main experiments, with five-fold results treated only as a secondary sensitivity check.","metadata":{}},{"cell_type":"code","source":"# 51. Check duplicate reports across the proposed three-fold split\n\ngold_report_folds = (\n    gold_report_data[\n        [ID_COLUMN, \"normalized_report\", \"language\"]\n    ]\n    .merge(\n        validation_assignments[\n            [ID_COLUMN, \"fold_3\"]\n        ],\n        on=ID_COLUMN,\n        how=\"left\",\n        validate=\"one_to_one\",\n    )\n)\n\nduplicate_report_fold_summary = (\n    gold_report_folds\n    .groupby(\"normalized_report\", as_index=False)\n    .agg(\n        number_of_gold_studies=(ID_COLUMN, \"size\"),\n        number_of_folds=(\"fold_3\", \"nunique\"),\n        folds=(\"fold_3\", lambda values: sorted(values.unique())),\n        languages=(\"language\", lambda values: sorted(values.unique())),\n    )\n    .query(\"number_of_gold_studies > 1\")\n    .sort_values(\n        [\"number_of_gold_studies\", \"number_of_folds\"],\n        ascending=False,\n    )\n    .reset_index(drop=True)\n)\n\ncross_fold_duplicate_reports = duplicate_report_fold_summary[\n    duplicate_report_fold_summary[\"number_of_folds\"] > 1\n]\n\nprint(\n    \"Duplicate normalized-report groups among gold-labelled studies:\",\n    len(duplicate_report_fold_summary),\n)\n\nprint(\n    \"Duplicate groups split across validation folds:\",\n    len(cross_fold_duplicate_reports),\n)\n\ndisplay(duplicate_report_fold_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T13:59:32.70374Z","iopub.execute_input":"2026-09-19T13:59:32.704589Z","iopub.status.idle":"2026-09-19T13:59:32.739426Z","shell.execute_reply.started":"2026-09-19T13:59:32.704552Z","shell.execute_reply":"2026-09-19T13:59:32.73875Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — gold-labelled reports do not duplicate across folds\n\nNo normalized report appears more than once among the gold-labelled studies. Therefore, the proposed three-fold split has no direct duplicate-report leakage within the gold-labelled validation data.\n\nWe still need to check whether any gold-labelled reports are duplicated among the report-only studies, because those reports may later be used for weak supervision.","metadata":{}},{"cell_type":"code","source":"# 52. Check overlap between gold-labelled and report-only reports\n\nreport_status_groups = (\n    report_data\n    .groupby(\"normalized_report\", as_index=False)\n    .agg(\n        total_studies=(ID_COLUMN, \"size\"),\n        gold_labelled_studies=(\n            \"label_status\",\n            lambda values: (values == \"Gold-labelled\").sum(),\n        ),\n        report_only_studies=(\n            \"label_status\",\n            lambda values: (values == \"Report-only\").sum(),\n        ),\n        example_report=(\"Report\", \"first\"),\n    )\n)\n\ngold_report_only_overlaps = (\n    report_status_groups\n    .query(\n        \"gold_labelled_studies > 0 \"\n        \"and report_only_studies > 0\"\n    )\n    .sort_values(\n        [\"gold_labelled_studies\", \"report_only_studies\"],\n        ascending=False,\n    )\n    .reset_index(drop=True)\n)\n\nprint(\n    \"Report groups shared by gold-labelled and report-only studies:\",\n    len(gold_report_only_overlaps),\n)\n\ndisplay(\n    gold_report_only_overlaps[\n        [\n            \"total_studies\",\n            \"gold_labelled_studies\",\n            \"report_only_studies\",\n            \"example_report\",\n        ]\n    ].head(15)\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T14:00:29.198733Z","iopub.execute_input":"2026-09-19T14:00:29.199197Z","iopub.status.idle":"2026-09-19T14:00:29.96653Z","shell.execute_reply.started":"2026-09-19T14:00:29.199147Z","shell.execute_reply":"2026-09-19T14:00:29.965722Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — one report template overlaps the labelled and report-only data\n\nOne normalized report appears in both the gold-labelled and report-only subsets. It occurs once among the gold-labelled studies and four times among the report-only studies.\n\nThis appears to be a repeated reporting template. It does not affect image-only validation, but report-based models should keep identical normalized reports in the same validation group to avoid learning from duplicated text.","metadata":{}},{"cell_type":"code","source":"# 53. Inspect the shared gold/report-only report group\n\nshared_normalized_reports = set(\n    gold_report_only_overlaps[\"normalized_report\"]\n)\n\nshared_report_details = (\n    report_data.loc[\n        report_data[\"normalized_report\"].isin(\n            shared_normalized_reports\n        ),\n        [\n            ID_COLUMN,\n            \"label_status\",\n            \"language\",\n            \"Report\",\n        ],\n    ]\n    .merge(\n        validation_assignments[\n            [ID_COLUMN, \"fold_3\"]\n        ],\n        on=ID_COLUMN,\n        how=\"left\",\n        validate=\"one_to_one\",\n    )\n)\n\nshared_report_details[\"report_excerpt\"] = (\n    shared_report_details[\"Report\"]\n    .str.replace(r\"\\s+\", \" \", regex=True)\n    .str.slice(0, 300)\n)\n\ndisplay(\n    shared_report_details[\n        [\n            ID_COLUMN,\n            \"label_status\",\n            \"language\",\n            \"fold_3\",\n            \"report_excerpt\",\n        ]\n    ]\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T14:02:04.719628Z","iopub.execute_input":"2026-09-19T14:02:04.720003Z","iopub.status.idle":"2026-09-19T14:02:04.740385Z","shell.execute_reply.started":"2026-09-19T14:02:04.719971Z","shell.execute_reply":"2026-09-19T14:02:04.739602Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — the duplicated report belongs to validation fold 1\n\nThe gold-labelled version of the repeated report is assigned to fold 1, while four identical report-only copies are available elsewhere in the dataset.\n\nFor image-based modelling this has no effect. For report-based or weakly supervised modelling, those four copies must be excluded when fold 1 is used for validation. This keeps the repeated report template entirely within one validation group.","metadata":{}},{"cell_type":"code","source":"# 54. Create report-aware assignments for leakage-safe validation\n\ngold_report_fold_map = (\n    gold_report_folds[\n        [\"normalized_report\", \"fold_3\"]\n    ]\n    .drop_duplicates()\n)\n\nassert (\n    gold_report_fold_map\n    .groupby(\"normalized_report\")[\"fold_3\"]\n    .nunique()\n    .le(1)\n    .all()\n)\n\nreport_validation_assignments = (\n    report_data[\n        [\n            ID_COLUMN,\n            \"label_status\",\n            \"normalized_report\",\n        ]\n    ]\n    .merge(\n        gold_report_fold_map,\n        on=\"normalized_report\",\n        how=\"left\",\n        validate=\"many_to_one\",\n    )\n    .rename(columns={\"fold_3\": \"report_validation_fold\"})\n)\n\nreport_only_overlap_summary = (\n    report_validation_assignments\n    .loc[\n        (\n            report_validation_assignments[\"label_status\"]\n            == \"Report-only\"\n        )\n        & report_validation_assignments[\n            \"report_validation_fold\"\n        ].notna()\n    ]\n    .groupby(\"report_validation_fold\")\n    .agg(\n        report_only_rows_to_exclude=(ID_COLUMN, \"size\"),\n        duplicated_report_groups=(\n            \"normalized_report\",\n            \"nunique\",\n        ),\n    )\n    .reindex(range(3), fill_value=0)\n    .rename_axis(\"validation_fold\")\n    .reset_index()\n)\n\ndisplay(report_only_overlap_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T14:03:34.914615Z","iopub.execute_input":"2026-09-19T14:03:34.914957Z","iopub.status.idle":"2026-09-19T14:03:34.959239Z","shell.execute_reply.started":"2026-09-19T14:03:34.914926Z","shell.execute_reply":"2026-09-19T14:03:34.958401Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Validation design decision\n\nThe primary validation scheme will use three-fold multi-label stratification at the study level.\n\nThis provides at least three positive cases for every target in every validation fold, including MCL. The fold assignments will remain fixed so that different modelling approaches can be compared fairly.\n\nFor report-based experiments, duplicated report groups will also be kept within the same validation group.","metadata":{"execution":{"iopub.status.busy":"2026-09-19T14:04:41.033727Z","iopub.execute_input":"2026-09-19T14:04:41.034723Z","iopub.status.idle":"2026-09-19T14:04:41.040929Z","shell.execute_reply.started":"2026-09-19T14:04:41.034684Z","shell.execute_reply":"2026-09-19T14:04:41.039795Z"}}},{"cell_type":"code","source":"# 55. Freeze the primary three-fold validation split\n\nPRIMARY_N_FOLDS = 3\nPRIMARY_FOLD_COLUMN = \"fold_3\"\n\nprimary_fold_assignments = (\n    validation_assignments[\n        [ID_COLUMN, PRIMARY_FOLD_COLUMN]\n    ]\n    .rename(\n        columns={\n            PRIMARY_FOLD_COLUMN: \"validation_fold\"\n        }\n    )\n)\n\ngold_study_table = (\n    gold_study_table\n    .drop(columns=[\"validation_fold\"], errors=\"ignore\")\n    .merge(\n        primary_fold_assignments,\n        on=ID_COLUMN,\n        how=\"left\",\n        validate=\"one_to_one\",\n    )\n)\n\nassert len(gold_study_table) == 58\nassert gold_study_table[\"validation_fold\"].notna().all()\nassert gold_study_table[\"validation_fold\"].nunique() == 3\n\ngold_study_table[\"number_of_positive_targets\"] = (\n    gold_study_table[TARGET_COLUMNS]\n    .sum(axis=1)\n    .astype(int)\n)\n\nprimary_fold_summary = (\n    gold_study_table\n    .groupby(\"validation_fold\")\n    .agg(\n        studies=(ID_COLUMN, \"size\"),\n        total_positive_labels=(\n            \"number_of_positive_targets\",\n            \"sum\",\n        ),\n        mean_positive_labels=(\n            \"number_of_positive_targets\",\n            \"mean\",\n        ),\n        minimum_positive_labels=(\n            \"number_of_positive_targets\",\n            \"min\",\n        ),\n        maximum_positive_labels=(\n            \"number_of_positive_targets\",\n            \"max\",\n        ),\n    )\n    .round(2)\n    .reset_index()\n)\n\ndisplay(primary_fold_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T14:04:58.743083Z","iopub.execute_input":"2026-09-19T14:04:58.743441Z","iopub.status.idle":"2026-09-19T14:04:58.76921Z","shell.execute_reply.started":"2026-09-19T14:04:58.74341Z","shell.execute_reply":"2026-09-19T14:04:58.768446Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Observation — the primary validation folds are well balanced\n\nThe three validation folds contain similar numbers of studies and positive labels. Their average number of positive abnormalities per study is also comparable.\n\nThis provides a reasonable fixed study-level split for initial experiments. Because the labelled dataset is small, differences between folds will still contribute uncertainty to the evaluation, but the split is sufficiently balanced for a first baseline.","metadata":{}},{"cell_type":"code","source":"# 56. Create fixed study-level fold splits\n\nfold_splits = {}\n\nfor fold in range(PRIMARY_N_FOLDS):\n\n    validation_ids = set(\n        gold_study_table.loc[\n            gold_study_table[\"validation_fold\"] == fold,\n            ID_COLUMN,\n        ]\n    )\n\n    training_ids = set(\n        gold_study_table.loc[\n            gold_study_table[\"validation_fold\"] != fold,\n            ID_COLUMN,\n        ]\n    )\n\n    fold_splits[fold] = {\n        \"training_ids\": training_ids,\n        \"validation_ids\": validation_ids,\n    }\n\n    print(\n        f\"Fold {fold}: \"\n        f\"{len(training_ids)} training studies, \"\n        f\"{len(validation_ids)} validation studies\"\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-19T14:06:21.844406Z","iopub.execute_input":"2026-09-19T14:06:21.844819Z","iopub.status.idle":"2026-09-19T14:06:21.852567Z","shell.execute_reply.started":"2026-09-19T14:06:21.844787Z","shell.execute_reply":"2026-09-19T14:06:21.851704Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}