{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"a98088a1-9abb-48b9-b48a-928fe35f7185","cell_type":"markdown","source":"# RSNA Knee Abnormality Detection: Deep Data Understanding\n\nThis notebook is a data-first audit before modeling. It inspects the competition tables, official labels, report text, MRI series metadata, and optional DICOM metadata so later pseudo-labeling and image-model choices are grounded in the actual data shape.\n\nMain questions:\n- How sparse are the official labels, and how imbalanced is each abnormality?\n- What does the report corpus look like, especially for the unlabeled studies?\n- Which report phrases are likely useful for weak supervision?\n- How many MRI series and imaging planes are available per study?\n- Which DICOM metadata fields can help select usable series for image training?\n","metadata":{}},{"id":"2a31a8c8-4d38-450e-bcf7-bba24be7143d","cell_type":"code","source":"from pathlib import Path\nimport os\nimport re\nimport warnings\n\nos.environ.setdefault('MPLCONFIGDIR', str(Path('/tmp') / 'matplotlib'))\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\ntry:\n    import seaborn as sns\n    sns.set_theme(style='whitegrid')\nexcept Exception:\n    sns = None\n\npd.set_option('display.max_columns', 120)\npd.set_option('display.max_colwidth', 240)\nwarnings.filterwarnings('ignore')\n\nCANDIDATE_DATA_ROOTS = [\n    Path('/kaggle/input/competitions/rsna-knee-abnormality-detection'),\n    Path('/kaggle/input/rsna-knee-abnormality-detection'),\n    Path.cwd() / 'data',\n    Path.cwd(),\n]\n\nDATA = next((p for p in CANDIDATE_DATA_ROOTS if (p / 'train.csv').exists()), None)\nif DATA is None:\n    raise FileNotFoundError(\n        'Could not find train.csv. On Kaggle, attach the competition dataset. '\n        'Locally, place the CSVs under ./data or the project root.'\n    )\n\nWORK = Path('/kaggle/working') if Path('/kaggle/working').exists() else Path('outputs')\nWORK.mkdir(parents=True, exist_ok=True)\n\nprint('DATA =', DATA)\nprint('WORK =', WORK)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-09T09:36:51.286126Z","iopub.execute_input":"2026-08-09T09:36:51.286268Z","iopub.status.idle":"2026-08-09T09:36:53.932593Z","shell.execute_reply.started":"2026-08-09T09:36:51.286251Z","shell.execute_reply":"2026-08-09T09:36:53.931805Z"}},"outputs":[],"execution_count":null},{"id":"bf9740c1-b069-4806-95ba-64a9e2db5c8a","cell_type":"markdown","source":"## Load Tables\n\nStart by loading each CSV and checking the exact columns. The notebook avoids assuming that every optional file is present, so it can be run locally with partial data and on Kaggle with the full dataset.\n","metadata":{}},{"id":"ff06f74b-6beb-42f5-9263-67470350520f","cell_type":"code","source":"def read_csv_if_exists(path):\n    if path.exists():\n        return pd.read_csv(path)\n    print(f'Missing optional file: {path}')\n    return None\n\ntrain = pd.read_csv(DATA / 'train.csv')\ntrain_series = read_csv_if_exists(DATA / 'train_series.csv')\ntest = read_csv_if_exists(DATA / 'test.csv')\ntest_series = read_csv_if_exists(DATA / 'test_series.csv')\nsample_submission = read_csv_if_exists(DATA / 'sample_submission.csv')\n\ntables = {\n    'train': train,\n    'train_series': train_series,\n    'test': test,\n    'test_series': test_series,\n    'sample_submission': sample_submission,\n}\n\ntable_summary = []\nfor name, df in tables.items():\n    if df is None:\n        continue\n    table_summary.append({\n        'table': name,\n        'rows': len(df),\n        'columns': df.shape[1],\n        'memory_mb': df.memory_usage(deep=True).sum() / 1_000_000,\n    })\n\ndisplay(pd.DataFrame(table_summary))\nfor name, df in tables.items():\n    if df is not None:\n        print(f'\\n{name} columns:')\n        print(list(df.columns))\n        display(df.head())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-09T09:37:47.789372Z","iopub.execute_input":"2026-08-09T09:37:47.789651Z","iopub.status.idle":"2026-08-09T09:37:47.964885Z","shell.execute_reply.started":"2026-08-09T09:37:47.789629Z","shell.execute_reply":"2026-08-09T09:37:47.964338Z"}},"outputs":[],"execution_count":null},{"id":"91204f22-779d-46c8-8630-ae886076257e","cell_type":"markdown","source":"## Schema And Key Integrity\n\nThese checks catch duplicate study IDs, unmatched series rows, and obvious submission-shape issues early.\n","metadata":{}},{"id":"0006af33-09c5-4f81-8d5d-f3e446b1a86b","cell_type":"code","source":"study_col = 'StudyInstanceUID'\nreport_col = 'Report'\n\nlabel_cols = [\n    'ACL',\n    'MCL',\n    'Medial Meniscus',\n    'Lateral Meniscus',\n    'Medial OA',\n    'Lateral OA',\n    'PF OA',\n    'Effusion',\n    'Synovitis',\n    \"Baker's\",\n    'Contusion',\n    'Fracture',\n]\nlabel_cols = [c for c in label_cols if c in train.columns]\n\nprint('Detected label columns:', label_cols)\nprint('Duplicate train study IDs:', int(train[study_col].duplicated().sum()))\n\nif train_series is not None:\n    print('Duplicate train series rows:', int(train_series.duplicated().sum()))\n    print('Train studies without series rows:', int((~train[study_col].isin(train_series[study_col])).sum()))\n    print('Series rows without train study:', int((~train_series[study_col].isin(train[study_col])).sum()))\n\nif test is not None and test_series is not None:\n    print('Test studies without series rows:', int((~test[study_col].isin(test_series[study_col])).sum()))\n    print('Test series rows without test study:', int((~test_series[study_col].isin(test[study_col])).sum()))\n\nif sample_submission is not None:\n    display(sample_submission.head())\n    print('Sample submission shape:', sample_submission.shape)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-09T09:37:49.773407Z","iopub.execute_input":"2026-08-09T09:37:49.774171Z","iopub.status.idle":"2026-08-09T09:37:49.804662Z","shell.execute_reply.started":"2026-08-09T09:37:49.77415Z","shell.execute_reply":"2026-08-09T09:37:49.804037Z"}},"outputs":[],"execution_count":null},{"id":"dd3c642d-31d5-44b5-8ff0-824390c6cbe8","cell_type":"markdown","source":"## Official Label Coverage\n\nThe official labels define the supervised signal. This section quantifies how many labels are present, how imbalanced each target is, and whether the same 58 studies are labeled for every target.\n","metadata":{}},{"id":"39a2c22c-3cae-4d58-a7b9-5fa890957103","cell_type":"code","source":"label_summary = []\nfor label in label_cols:\n    labeled = train[label].notna()\n    positives = int(train.loc[labeled, label].sum()) if labeled.any() else 0\n    negatives = int(labeled.sum() - positives)\n    label_summary.append({\n        'label': label,\n        'labeled': int(labeled.sum()),\n        'missing': int((~labeled).sum()),\n        'positive': positives,\n        'negative': negatives,\n        'prevalence': positives / labeled.sum() if labeled.any() else np.nan,\n    })\n\nlabel_summary = pd.DataFrame(label_summary).sort_values('prevalence', ascending=False)\ndisplay(label_summary)\n\ntrain['n_official_labels'] = train[label_cols].notna().sum(axis=1)\ntrain['has_any_official_label'] = train['n_official_labels'] > 0\nprint('Studies with any official label:', int(train['has_any_official_label'].sum()))\ndisplay(train['n_official_labels'].value_counts().sort_index().rename('n_studies').to_frame())\n\nfig, axes = plt.subplots(1, 2, figsize=(14, 4))\nlabel_summary.sort_values('prevalence').plot.barh(x='label', y='prevalence', ax=axes[0], legend=False)\naxes[0].set_title('Official positive prevalence')\naxes[0].set_xlabel('Positive fraction among labeled studies')\n\nlabel_summary.sort_values('labeled').plot.barh(x='label', y=['positive', 'negative'], stacked=True, ax=axes[1])\naxes[1].set_title('Official labeled examples')\naxes[1].set_xlabel('Study count')\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-09T09:37:50.730669Z","iopub.execute_input":"2026-08-09T09:37:50.730922Z","iopub.status.idle":"2026-08-09T09:37:51.151629Z","shell.execute_reply.started":"2026-08-09T09:37:50.730903Z","shell.execute_reply":"2026-08-09T09:37:51.151127Z"}},"outputs":[],"execution_count":null},{"id":"a2deb1c3-2ff4-4507-aa28-33ff50c7e29f","cell_type":"markdown","source":"## Label Relationships\n\nWith only a small official subset, correlations are noisy. They are still useful for spotting common co-occurrence patterns and label combinations that may affect validation splits.\n","metadata":{}},{"id":"b3d9b188-a409-4ccb-8cd0-1de8fb876ffd","cell_type":"code","source":"labeled_rows = train[train[label_cols].notna().all(axis=1)].copy()\nprint('Rows with a complete official label vector:', len(labeled_rows))\n\nif len(labeled_rows):\n    labeled_rows['positive_label_count'] = labeled_rows[label_cols].sum(axis=1).astype(int)\n    display(labeled_rows['positive_label_count'].describe().to_frame())\n    display(labeled_rows['positive_label_count'].value_counts().sort_index().rename('n_studies').to_frame())\n\n    co_occurrence = labeled_rows[label_cols].T.dot(labeled_rows[label_cols])\n    display(co_occurrence)\n\n    if sns is not None:\n        plt.figure(figsize=(9, 7))\n        sns.heatmap(labeled_rows[label_cols].corr(), vmin=-1, vmax=1, cmap='coolwarm', square=True)\n        plt.title('Official label correlation, labeled subset only')\n        plt.tight_layout()\n        plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-09T09:37:51.766596Z","iopub.execute_input":"2026-08-09T09:37:51.766839Z","iopub.status.idle":"2026-08-09T09:37:52.085395Z","shell.execute_reply.started":"2026-08-09T09:37:51.766823Z","shell.execute_reply":"2026-08-09T09:37:52.084849Z"}},"outputs":[],"execution_count":null},{"id":"394c4c6c-9439-497a-9b18-be6d91b72559","cell_type":"markdown","source":"## Report Text Audit\n\nReports are the practical bridge from 58 official labels to thousands of weak labels. This section inspects coverage, length, non-ASCII content, and example reports from labeled and unlabeled studies.\n","metadata":{}},{"id":"7a626413-96ec-4886-afa7-44547c269b29","cell_type":"code","source":"if report_col not in train.columns:\n    raise KeyError('Expected a Report column in train.csv')\n\ntrain['report_text'] = train[report_col].fillna('').astype(str)\ntrain['has_report'] = train['report_text'].str.strip().ne('')\ntrain['report_chars'] = train['report_text'].str.len()\ntrain['report_words'] = train['report_text'].str.split().str.len()\ntrain['report_lines'] = train['report_text'].str.count('\\n') + train['has_report'].astype(int)\ntrain['report_non_ascii_chars'] = train['report_text'].apply(lambda s: sum(ord(ch) > 127 for ch in s))\ntrain['report_non_ascii_frac'] = train['report_non_ascii_chars'] / train['report_chars'].replace(0, np.nan)\n\nreport_summary = train[[\n    'has_report',\n    'report_chars',\n    'report_words',\n    'report_lines',\n    'report_non_ascii_frac',\n    'has_any_official_label',\n]].describe(include='all')\ndisplay(report_summary)\n\ndisplay(pd.crosstab(train['has_report'], train['has_any_official_label'], margins=True))\n\nfig, axes = plt.subplots(1, 2, figsize=(14, 4))\ntrain['report_words'].clip(upper=train['report_words'].quantile(0.99)).hist(bins=50, ax=axes[0])\naxes[0].set_title('Report word count, clipped at p99')\naxes[0].set_xlabel('Words')\n\ntrain.boxplot(column='report_words', by='has_any_official_label', ax=axes[1])\naxes[1].set_title('Report length by official label availability')\naxes[1].set_xlabel('Has official label')\naxes[1].set_ylabel('Words')\nplt.suptitle('')\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-09T09:37:52.736982Z","iopub.execute_input":"2026-08-09T09:37:52.737253Z","iopub.status.idle":"2026-08-09T09:37:53.309458Z","shell.execute_reply.started":"2026-08-09T09:37:52.737234Z","shell.execute_reply":"2026-08-09T09:37:53.30879Z"}},"outputs":[],"execution_count":null},{"id":"35152580-cef1-401d-ac74-8077099984f6","cell_type":"code","source":"def show_report_examples(df, title, n=3, max_chars=1600):\n    print('\\n' + '=' * 100)\n    print(title)\n    print('=' * 100)\n    sample_df = df[df['has_report']].sample(min(n, df[df['has_report']].shape[0]), random_state=42)\n    for _, row in sample_df.iterrows():\n        print('\\nStudyInstanceUID:', row[study_col])\n        print(row['report_text'][:max_chars])\n\nshow_report_examples(train[train['has_any_official_label']], 'Officially labeled report examples')\nshow_report_examples(train[~train['has_any_official_label']], 'Unlabeled report examples')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-09T09:37:53.31064Z","iopub.execute_input":"2026-08-09T09:37:53.310878Z","iopub.status.idle":"2026-08-09T09:37:53.324275Z","shell.execute_reply.started":"2026-08-09T09:37:53.310859Z","shell.execute_reply":"2026-08-09T09:37:53.323648Z"}},"outputs":[],"execution_count":null},{"id":"6a071187-6bcf-4f80-b14e-5bcd5a3bf980","cell_type":"markdown","source":"## Report Keyword Signals\n\nThis lightweight audit does not create final pseudo labels. It shows whether expected abnormality words appear in reports and how often they overlap with official positives/negatives. Use it to design transparent rules or sanity-check text-model pseudo labels.\n","metadata":{}},{"id":"d1725f01-271f-4ab4-ad93-6fc0f3b62d1a","cell_type":"code","source":"keyword_patterns = {\n    'ACL': r'\\b(?:acl|anterior cruciate)\\b',\n    'MCL': r'\\b(?:mcl|medial collateral)\\b',\n    'Medial Meniscus': r'\\bmedial menisc(?:us|al)\\b|\\bposterior horn medial\\b',\n    'Lateral Meniscus': r'\\blateral menisc(?:us|al)\\b|\\bposterior horn lateral\\b',\n    'Medial OA': r'\\bmedial (?:compartment )?(?:osteoarthritis|oa|chondrosis|cartilage loss)\\b',\n    'Lateral OA': r'\\blateral (?:compartment )?(?:osteoarthritis|oa|chondrosis|cartilage loss)\\b',\n    'PF OA': r'\\b(?:patellofemoral|pf) (?:osteoarthritis|oa|chondrosis|cartilage loss)\\b',\n    'Effusion': r'\\beffusion\\b|\\bjoint fluid\\b',\n    'Synovitis': r'\\bsynovitis\\b|\\bsynovial\\b',\n    \"Baker's\": r'\\bbaker(?:\\'s)? cyst\\b|\\bpopliteal cyst\\b',\n    'Contusion': r'\\bcontusion\\b|\\bbone bruise\\b|\\bmarrow edema\\b',\n    'Fracture': r'\\bfracture\\b|\\bfx\\b',\n}\n\nnegation_window = r'(?:no|not|without|absent|negative for|free of|intact|unremarkable)'\n\nkeyword_rows = []\nfor label in label_cols:\n    pattern = keyword_patterns.get(label)\n    if pattern is None:\n        continue\n    mention = train['report_text'].str.contains(pattern, flags=re.IGNORECASE, regex=True, na=False)\n    negated = train['report_text'].str.contains(\n        negation_window + r'[^.\\n;]{0,80}' + pattern + r'|' + pattern + r'[^.\\n;]{0,80}' + negation_window,\n        flags=re.IGNORECASE,\n        regex=True,\n        na=False,\n    )\n    official = train[label].notna()\n    keyword_rows.append({\n        'label': label,\n        'mentions_anywhere': int(mention.sum()),\n        'negated_near_keyword': int((mention & negated).sum()),\n        'mentions_in_official_positive': int((mention & official & (train[label] == 1)).sum()),\n        'mentions_in_official_negative': int((mention & official & (train[label] == 0)).sum()),\n        'mentions_in_unlabeled': int((mention & ~train['has_any_official_label']).sum()),\n    })\n\nkeyword_summary = pd.DataFrame(keyword_rows)\ndisplay(keyword_summary)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-09T09:37:54.677249Z","iopub.execute_input":"2026-08-09T09:37:54.677497Z","iopub.status.idle":"2026-08-09T09:38:01.726059Z","shell.execute_reply.started":"2026-08-09T09:37:54.67746Z","shell.execute_reply":"2026-08-09T09:38:01.725486Z"}},"outputs":[],"execution_count":null},{"id":"881356b3-47a6-4a07-a22f-4059aa7f7b7e","cell_type":"markdown","source":"## Series Metadata\n\nThe first image model should not ingest every slice blindly. This section summarizes imaging planes, fluid-sensitive series, fat suppression, and per-study series availability.\n","metadata":{}},{"id":"f11c506d-b432-4a09-9498-7833d2c55a5e","cell_type":"code","source":"if train_series is None:\n    print('train_series.csv is unavailable; skipping series metadata audit.')\nelse:\n    display(train_series.head())\n    display(train_series.describe(include='all'))\n\n    categorical_cols = [c for c in ['Anatomical_Plane', 'SeriesDescription', 'Fluid_Sensitive', 'Fat_Suppression'] if c in train_series.columns]\n    for col in categorical_cols:\n        print('\\n' + col)\n        display(train_series[col].value_counts(dropna=False).head(30).to_frame('n_rows'))\n\n    if {'Anatomical_Plane', 'Fluid_Sensitive', 'Fat_Suppression'}.issubset(train_series.columns):\n        display(pd.crosstab(\n            train_series['Anatomical_Plane'],\n            [train_series['Fluid_Sensitive'], train_series['Fat_Suppression']],\n            margins=True,\n        ))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-09T09:38:01.727257Z","iopub.execute_input":"2026-08-09T09:38:01.727505Z","iopub.status.idle":"2026-08-09T09:38:01.795582Z","shell.execute_reply.started":"2026-08-09T09:38:01.727453Z","shell.execute_reply":"2026-08-09T09:38:01.794867Z"}},"outputs":[],"execution_count":null},{"id":"3813460a-ad8f-4a3b-8590-144426b4a59a","cell_type":"code","source":"if train_series is not None:\n    study_series_features = train_series.groupby(study_col).agg(\n        n_series=('SeriesInstanceUID', 'count'),\n        n_planes=('Anatomical_Plane', 'nunique') if 'Anatomical_Plane' in train_series.columns else ('SeriesInstanceUID', 'count'),\n        n_fluid_sensitive=('Fluid_Sensitive', 'sum') if 'Fluid_Sensitive' in train_series.columns else ('SeriesInstanceUID', 'count'),\n        n_fat_suppressed=('Fat_Suppression', 'sum') if 'Fat_Suppression' in train_series.columns else ('SeriesInstanceUID', 'count'),\n    )\n\n    if 'Anatomical_Plane' in train_series.columns:\n        plane_counts = train_series.pivot_table(\n            index=study_col,\n            columns='Anatomical_Plane',\n            values='SeriesInstanceUID',\n            aggfunc='count',\n            fill_value=0,\n        ).add_prefix('n_plane_')\n        study_series_features = study_series_features.join(plane_counts, how='left').fillna(0)\n\n    study_series_features = study_series_features.reset_index()\n    study_series_features = train[[study_col, 'has_any_official_label']].merge(study_series_features, on=study_col, how='left')\n    display(study_series_features.describe(include='all'))\n\n    fig, axes = plt.subplots(1, 2, figsize=(14, 4))\n    study_series_features['n_series'].hist(bins=40, ax=axes[0])\n    axes[0].set_title('MRI series per train study')\n    axes[0].set_xlabel('Series count')\n\n    study_series_features.boxplot(column='n_series', by='has_any_official_label', ax=axes[1])\n    axes[1].set_title('Series count by official label availability')\n    axes[1].set_xlabel('Has official label')\n    plt.suptitle('')\n    plt.tight_layout()\n    plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-09T09:38:01.796396Z","iopub.execute_input":"2026-08-09T09:38:01.796664Z","iopub.status.idle":"2026-08-09T09:38:02.055256Z","shell.execute_reply.started":"2026-08-09T09:38:01.79664Z","shell.execute_reply":"2026-08-09T09:38:02.054662Z"}},"outputs":[],"execution_count":null},{"id":"ba3eed1d-970f-417f-bfbd-7a3bdec1be53","cell_type":"markdown","source":"## Optional DICOM Metadata Sampling\n\nThis scans a small number of DICOM files without loading full pixel arrays. It helps confirm slice counts, image dimensions, spacing, and protocol naming. Increase `MAX_STUDIES` only after the quick sample works.\n","metadata":{}},{"id":"b91c878c-d63d-4959-b850-f3a87e661852","cell_type":"code","source":"try:\n    import pydicom\nexcept Exception as exc:\n    pydicom = None\n    print('pydicom is unavailable:', exc)\n\nTRAIN_SERIES_DIR = DATA / 'train_series'\nMAX_STUDIES = 12\nMAX_SERIES_PER_STUDY = 4\nMAX_SLICES_PER_SERIES = 3\n\ndicom_rows = []\nif pydicom is not None and TRAIN_SERIES_DIR.exists():\n    study_dirs = sorted([p for p in TRAIN_SERIES_DIR.iterdir() if p.is_dir()])[:MAX_STUDIES]\n    for study_dir in study_dirs:\n        series_dirs = sorted([p for p in study_dir.iterdir() if p.is_dir()])[:MAX_SERIES_PER_STUDY]\n        for series_dir in series_dirs:\n            dcm_files = sorted(series_dir.glob('*.dcm'))\n            for dcm_path in dcm_files[:MAX_SLICES_PER_SERIES]:\n                ds = pydicom.dcmread(dcm_path, stop_before_pixels=True, force=True)\n                dicom_rows.append({\n                    'StudyInstanceUID': study_dir.name,\n                    'SeriesInstanceUID': series_dir.name,\n                    'SOPInstanceUID': dcm_path.stem,\n                    'SeriesDescription': getattr(ds, 'SeriesDescription', None),\n                    'ProtocolName': getattr(ds, 'ProtocolName', None),\n                    'Rows': getattr(ds, 'Rows', None),\n                    'Columns': getattr(ds, 'Columns', None),\n                    'PixelSpacing': str(getattr(ds, 'PixelSpacing', None)),\n                    'SliceThickness': getattr(ds, 'SliceThickness', None),\n                    'SpacingBetweenSlices': getattr(ds, 'SpacingBetweenSlices', None),\n                    'ImagePositionPatient': str(getattr(ds, 'ImagePositionPatient', None)),\n                })\n\ndicom_sample = pd.DataFrame(dicom_rows)\nprint('DICOM metadata rows sampled:', len(dicom_sample))\ndisplay(dicom_sample.head(20))\n\nif len(dicom_sample):\n    display(dicom_sample[['Rows', 'Columns', 'SliceThickness', 'SpacingBetweenSlices']].describe())\n    display(dicom_sample['SeriesDescription'].value_counts(dropna=False).head(30).to_frame('n_sampled_slices'))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-09T09:38:02.056478Z","iopub.execute_input":"2026-08-09T09:38:02.05672Z","iopub.status.idle":"2026-08-09T09:38:09.762564Z","shell.execute_reply.started":"2026-08-09T09:38:02.056696Z","shell.execute_reply":"2026-08-09T09:38:09.761884Z"}},"outputs":[],"execution_count":null},{"id":"a28eeb1b-1398-4d6e-bb94-41690d3d665d","cell_type":"markdown","source":"## Modeling Implications And Saved Artifacts\n\nSave compact EDA artifacts that downstream notebooks can reuse. The key output is a study-level table combining report length, official label availability, and series availability.\n","metadata":{}},{"id":"077ce791-d03a-4f25-a99b-956885672db6","cell_type":"code","source":"artifacts = {}\nartifacts['label_summary'] = label_summary\nartifacts['keyword_summary'] = keyword_summary if 'keyword_summary' in globals() else pd.DataFrame()\n\nif 'study_series_features' in globals():\n    study_understanding = train[[\n        study_col,\n        'has_report',\n        'report_chars',\n        'report_words',\n        'report_non_ascii_frac',\n        'has_any_official_label',\n        'n_official_labels',\n    ] + label_cols].merge(study_series_features.drop(columns=['has_any_official_label'], errors='ignore'), on=study_col, how='left')\nelse:\n    study_understanding = train[[\n        study_col,\n        'has_report',\n        'report_chars',\n        'report_words',\n        'report_non_ascii_frac',\n        'has_any_official_label',\n        'n_official_labels',\n    ] + label_cols].copy()\n\nartifacts['study_understanding'] = study_understanding\n\nfor name, df in artifacts.items():\n    out_path = WORK / f'{name}.csv'\n    df.to_csv(out_path, index=False)\n    print('Saved', out_path)\n\ndisplay(study_understanding.head())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-09T09:38:09.763424Z","iopub.execute_input":"2026-08-09T09:38:09.763673Z","iopub.status.idle":"2026-08-09T09:38:09.821475Z","shell.execute_reply.started":"2026-08-09T09:38:09.763652Z","shell.execute_reply":"2026-08-09T09:38:09.820638Z"}},"outputs":[],"execution_count":null}]}