{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"e8dfc852","cell_type":"markdown","source":"**Read the Reports: Radiology Text EDA**\n\n[`train.csv`](https://www.kaggle.com/competitions/rsna-knee-abnormality-detection/data?select=train.csv) contains a free-text radiology report for the training studies.\n\nThis notebook is only a simple text EDA. We will look at:\n\n- report availability\n- report length\n- a few examples\n- basic text cleaning\n- common words\n\nThe reports are part of the training data; they are not provided in the hidden test set.","metadata":{}},{"id":"817ffa94-8885-4037-9448-5af092616854","cell_type":"markdown","source":"### Previous notebooks\n\nThis notebook is part of a small RSNA Knee exploration series:\n\n1. [RSNA Knee -> Simple Starting Point (EDA)](https://www.kaggle.com/code/h17ann/1-rsna-knee-simple-starting-point-eda)\n2. [RSNA Knee -> Look Under the Hood: DICOM Metadata Explorer](https://www.kaggle.com/code/h17ann/2-rsna-knee-dicom-metadata-explorer)\n3. [RSNA Knee -> See the MRI: From DICOM to Slices & Studies](https://www.kaggle.com/code/h17ann/3-rsna-knee-see-the-mri-ax-cor-sag)\n4. [RSNA Knee -> Know the Targets: 12 Abnormalities Explained](https://www.kaggle.com/code/h17ann/4-rsna-knee-know-the-12-targets)","metadata":{}},{"id":"4d6f5f64","cell_type":"markdown","source":"## 1. Imports","metadata":{}},{"id":"6a1ba462","cell_type":"code","source":"from pathlib import Path\nfrom collections import Counter\n\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:01:34.925901Z","iopub.execute_input":"2026-08-25T20:01:34.926275Z","iopub.status.idle":"2026-08-25T20:01:34.975033Z","shell.execute_reply.started":"2026-08-25T20:01:34.926247Z","shell.execute_reply":"2026-08-25T20:01:34.972538Z"}},"outputs":[],"execution_count":null},{"id":"25a087b9","cell_type":"markdown","source":"## 2. Load the training table","metadata":{}},{"id":"5816f577","cell_type":"code","source":"DATA_DIR = Path(\"/kaggle/input/competitions/rsna-knee-abnormality-detection\")\n# Download the data if it is not already attached\nif not DATA_DIR.exists():\n    DATA_DIR = Path(\n        kagglehub.competition_download(\n            \"rsna-knee-abnormality-detection\"\n        )\n    )\ntrain = pd.read_csv(DATA_DIR / \"train.csv\")\n\nprint(\"Training studies:\", len(train))\n\ndisplay(\n    train[[\"StudyInstanceUID\", \"Report\"]].head()\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:01:34.980879Z","iopub.execute_input":"2026-08-25T20:01:34.981552Z","iopub.status.idle":"2026-08-25T20:01:35.239361Z","shell.execute_reply.started":"2026-08-25T20:01:34.981489Z","shell.execute_reply":"2026-08-25T20:01:35.237398Z"}},"outputs":[],"execution_count":null},{"id":"c4684b34","cell_type":"markdown","source":"## 3. Report availability","metadata":{}},{"id":"f84527a1","cell_type":"code","source":"print(\"Available reports:\", train[\"Report\"].notna().sum())\nprint(\"Missing reports  :\", train[\"Report\"].isna().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:19:35.626415Z","iopub.execute_input":"2026-08-25T20:19:35.627024Z","iopub.status.idle":"2026-08-25T20:19:35.639686Z","shell.execute_reply.started":"2026-08-25T20:19:35.626989Z","shell.execute_reply":"2026-08-25T20:19:35.637461Z"}},"outputs":[],"execution_count":null},{"id":"cc0a005f","cell_type":"markdown","source":"## 4. Report length in characters","metadata":{}},{"id":"7f385e52","cell_type":"code","source":"reports = train[\"Report\"].dropna().astype(str)\n\nchar_length = reports.str.len()\n\ndisplay(\n    char_length.describe().to_frame(\"characters\")\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:01:35.301634Z","iopub.execute_input":"2026-08-25T20:01:35.302056Z","iopub.status.idle":"2026-08-25T20:01:35.351021Z","shell.execute_reply.started":"2026-08-25T20:01:35.302023Z","shell.execute_reply":"2026-08-25T20:01:35.349109Z"}},"outputs":[],"execution_count":null},{"id":"ee5c5527","cell_type":"code","source":"plt.figure(figsize=(10, 5))\n\nsns.histplot(\n    char_length,\n    bins=40\n)\n\nplt.xlabel(\"Characters\")\nplt.ylabel(\"Number of reports\")\nplt.title(\"Report length in characters\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:01:35.353265Z","iopub.execute_input":"2026-08-25T20:01:35.354173Z","iopub.status.idle":"2026-08-25T20:01:35.988194Z","shell.execute_reply.started":"2026-08-25T20:01:35.354111Z","shell.execute_reply":"2026-08-25T20:01:35.983652Z"}},"outputs":[],"execution_count":null},{"id":"2f72cb79","cell_type":"markdown","source":"## 5. Report length in words","metadata":{}},{"id":"bac276ad","cell_type":"code","source":"word_length = reports.str.split().str.len()\n\ndisplay(\n    word_length.describe().to_frame(\"words\")\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:01:35.993334Z","iopub.execute_input":"2026-08-25T20:01:35.993947Z","iopub.status.idle":"2026-08-25T20:01:36.284838Z","shell.execute_reply.started":"2026-08-25T20:01:35.993896Z","shell.execute_reply":"2026-08-25T20:01:36.283308Z"}},"outputs":[],"execution_count":null},{"id":"e928b1bd","cell_type":"code","source":"plt.figure(figsize=(10, 5))\n\nsns.histplot(\n    word_length,\n    bins=40\n)\n\nplt.xlabel(\"Words\")\nplt.ylabel(\"Number of reports\")\nplt.title(\"Report length in words\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:01:36.287279Z","iopub.execute_input":"2026-08-25T20:01:36.287804Z","iopub.status.idle":"2026-08-25T20:01:36.736391Z","shell.execute_reply.started":"2026-08-25T20:01:36.287734Z","shell.execute_reply":"2026-08-25T20:01:36.734064Z"}},"outputs":[],"execution_count":null},{"id":"fedefd96","cell_type":"markdown","source":"## 6. A few examples","metadata":{}},{"id":"a80b6bcc","cell_type":"code","source":"sample_reports = train.loc[\n    train[\"Report\"].notna(),\n    [\"StudyInstanceUID\", \"Report\"]\n].sample(5, random_state=42)\n\nfor row in sample_reports.itertuples(index=False):\n    print(\"=\" * 80)\n    print(\"Study:\", row.StudyInstanceUID)\n    print(row.Report)\n    print()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:01:36.738283Z","iopub.execute_input":"2026-08-25T20:01:36.738728Z","iopub.status.idle":"2026-08-25T20:01:36.75817Z","shell.execute_reply.started":"2026-08-25T20:01:36.738695Z","shell.execute_reply":"2026-08-25T20:01:36.756279Z"}},"outputs":[],"execution_count":null},{"id":"fe710986","cell_type":"markdown","source":"## 7. Simple text cleaning\n\nThe reports can contain different languages, so I do not restrict the text to English characters.\n\nThis small cleaning step only:\n\n- converts text to lowercase\n- keeps alphabetic characters from any language\n- replaces numbers and most punctuation with spaces, while keeping apostrophes.","metadata":{}},{"id":"cb3dc5bd","cell_type":"code","source":"def clean_text(text):\n    text = text.lower()\n\n    cleaned = []\n\n    for char in text:\n        if char.isalpha() or char == \"'\":\n            cleaned.append(char)\n        else:\n            cleaned.append(\" \")\n\n    return \" \".join(\"\".join(cleaned).split())\n\n\nclean_reports = reports.map(clean_text)\n\ndisplay(clean_reports.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:01:36.760396Z","iopub.execute_input":"2026-08-25T20:01:36.760853Z","iopub.status.idle":"2026-08-25T20:01:37.434117Z","shell.execute_reply.started":"2026-08-25T20:01:36.76082Z","shell.execute_reply":"2026-08-25T20:01:37.432376Z"}},"outputs":[],"execution_count":null},{"id":"8471a9ed","cell_type":"markdown","source":"## 8. Most common words","metadata":{}},{"id":"330066fd","cell_type":"code","source":"words = []\n\nfor report in clean_reports:\n    words.extend(report.split())\n\nword_counts = Counter(words)\n\ncommon_words = pd.DataFrame(\n    word_counts.most_common(30),\n    columns=[\"word\", \"count\"]\n)\n\ndisplay(common_words)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:01:37.436176Z","iopub.execute_input":"2026-08-25T20:01:37.437937Z","iopub.status.idle":"2026-08-25T20:01:37.838974Z","shell.execute_reply.started":"2026-08-25T20:01:37.437896Z","shell.execute_reply":"2026-08-25T20:01:37.836879Z"}},"outputs":[],"execution_count":null},{"id":"bbbed084","cell_type":"code","source":"plt.figure(figsize=(10, 8))\n\nsns.barplot(\n    data=common_words,\n    x=\"count\",\n    y=\"word\"\n)\n\nplt.xlabel(\"Count\")\nplt.ylabel(\"\")\nplt.title(\"Most common words in the reports\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:01:37.841559Z","iopub.execute_input":"2026-08-25T20:01:37.842567Z","iopub.status.idle":"2026-08-25T20:01:38.645062Z","shell.execute_reply.started":"2026-08-25T20:01:37.842528Z","shell.execute_reply":"2026-08-25T20:01:38.643604Z"}},"outputs":[],"execution_count":null},{"id":"9a4d7d02","cell_type":"markdown","source":"The reports are multilingual, so this is only a descriptive word-frequency view.\n\nA language-aware NLP pipeline would need more careful tokenization and preprocessing.","metadata":{}},{"id":"3491c95a","cell_type":"markdown","source":"## 9. Vocabulary size","metadata":{}},{"id":"8352c654","cell_type":"code","source":"print(\"Total word tokens:\", len(words))\nprint(\"Unique words     :\", len(set(words)))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-25T20:01:38.64995Z","iopub.execute_input":"2026-08-25T20:01:38.650518Z","iopub.status.idle":"2026-08-25T20:01:38.73005Z","shell.execute_reply.started":"2026-08-25T20:01:38.650484Z","shell.execute_reply":"2026-08-25T20:01:38.728251Z"}},"outputs":[],"execution_count":null},{"id":"97b28de4","cell_type":"markdown","source":"## Takeaways\n\n- The radiology reports are stored in `train.csv`.\n- Report length varies across studies.\n- The reports are multilingual, so text cleaning should preserve non-English characters.\n- A simple word-frequency view gives a first look at the vocabulary.\n- This notebook stays at the descriptive EDA level and does not derive targets from the reports.\n\nThat completes this small public EDA series:\n\n1. **Simple Starting Point (EDA)**\n2. **Look Under the Hood: DICOM Metadata Explorer**\n3. **See the MRI: From DICOM to Slices & Studies**\n4. **Know the Targets: 12 Abnormalities Explained**\n5. **Read the Reports: Radiology Text EDA**","metadata":{}},{"id":"a6c0f544-32b9-41e1-aa01-73c8b92537cf","cell_type":"markdown","source":"More notebooks are coming **SOON** ... \nNext up: baseline models. **STAY TUNED!**  \ntbc ...","metadata":{}}]}