{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceType":"competition","sourceId":99552,"databundleVersionId":13441085},{"sourceType":"datasetVersion","sourceId":12486089,"datasetId":7878914,"databundleVersionId":13063697},{"sourceType":"datasetVersion","sourceId":12462836,"datasetId":7861961,"databundleVersionId":13036760},{"sourceType":"datasetVersion","sourceId":12462842,"datasetId":7861966,"databundleVersionId":13036767}],"dockerImageVersionId":31090,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🧠 RSNA Intracranial Aneurysm Detection","metadata":{}},{"cell_type":"code","source":"from IPython.display import Image, display\ndisplay(Image('/kaggle/input/photo-1/1.png'))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:02:32.9276Z","iopub.execute_input":"2025-08-29T05:02:32.927886Z","iopub.status.idle":"2025-08-29T05:02:32.964299Z","shell.execute_reply.started":"2025-08-29T05:02:32.927861Z","shell.execute_reply":"2025-08-29T05:02:32.963354Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## [1.1] What is an Intracranial Aneurysm?\n- An **intracranial aneurysm** is a balloon-like bulge in a blood vessel of the brain caused by a weakness in the vessel wall.  \n- Often **asymptomatic**, but rupture leads to **subarachnoid hemorrhage (SAH)**, a severe type of stroke.  \n- **Prevalence:** ~3% globally.  \n- **Mortality rate:** Up to **50% of ruptures** result in death or severe disability.\n","metadata":{"execution":{"iopub.status.busy":"2025-07-30T13:51:58.690059Z","iopub.execute_input":"2025-07-30T13:51:58.690321Z","iopub.status.idle":"2025-07-30T13:51:58.696021Z","shell.execute_reply.started":"2025-07-30T13:51:58.690301Z","shell.execute_reply":"2025-07-30T13:51:58.695157Z"}}},{"cell_type":"markdown","source":"## [1.2] Problem Statement\n- Brain aneurysms are **frequently missed in routine brain imaging**, especially when scans are done for unrelated reasons.  \n- Manual detection is:\n  - **Time-consuming** for radiologists.\n  - **Error-prone**, especially for **small aneurysms** or when imaging varies by modality/scanner.\n- Automated **AI-based detection systems** can:\n  - Serve as a **“second reader”** for radiologists.\n  - Enable **early detection and treatment**, reducing fatal outcomes.\n","metadata":{}},{"cell_type":"markdown","source":"## [1.3] Project Goals\n- **Analyze and understand the dataset** (labels, modalities, demographics).\n- Identify **patterns and challenges** (imbalances, modality-specific variations).\n- Prepare data for a **deep learning model**:\n  - Binary aneurysm detection.\n  - Multi-label classification across 13 vessel locations.\n- Lay the foundation for a **3D CNN pipeline** that can process volumetric medical scans.","metadata":{}},{"cell_type":"markdown","source":"## [1.4] What Are We Predicting?\nWe aim to predict:\n1. **Aneurysm Present (Main Target):**\n   - Binary label (0 = No, 1 = Yes).\n   \n2. **13 Anatomical Vessel Locations (Multi-label):**\n   - `Left/Right Infraclinoid Internal Carotid Artery`\n   - `Left/Right Supraclinoid Internal Carotid Artery`\n   - `Left/Right Middle Cerebral Artery`\n   - `Anterior Communicating Artery`\n   - `Left/Right Anterior Cerebral Artery`\n   - `Left/Right Posterior Communicating Artery`\n   - `Basilar Tip`\n   - `Other Posterior Circulation`","metadata":{}},{"cell_type":"markdown","source":"📌 **Final output:** A **14-label prediction** (13 vessels + aneurysm presence).","metadata":{}},{"cell_type":"code","source":"from IPython.display import Image, display\ndisplay(Image('/kaggle/input/photo-2/2.png'))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:02:32.965799Z","iopub.execute_input":"2025-08-29T05:02:32.966114Z","iopub.status.idle":"2025-08-29T05:02:32.99964Z","shell.execute_reply.started":"2025-08-29T05:02:32.966088Z","shell.execute_reply":"2025-08-29T05:02:32.998839Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## [2.1] Dataset Overview\n\n| **File Name**           | **Description**                                                                |\n|--------------------------|--------------------------------------------------------------------------------|\n| `train.csv`             | Series-level labels (13 locations + `Aneurysm Present` target).               |\n| `train_localizers.csv`   | Coordinates of aneurysm centers for localization tasks.                       |\n| `series/`               | Folders of brain scans in **DICOM format**, grouped by `SeriesInstanceUID`.   |\n| `segmentations/`        | NIfTI vessel segmentation masks (subset of cases).                            |\n\n### **Key Columns in `train.csv`**\n- `SeriesInstanceUID`: Unique identifier for each scan series.  \n- `Modality`: Imaging type (CTA, MRA, MRI).  \n- `PatientAge` & `PatientSex`: Demographics.  \n- **13 binary vessel location labels.**  \n- `Aneurysm Present`: Binary indicator (main target).","metadata":{}},{"cell_type":"markdown","source":"## [2.2] Why This Matters\n- Early detection allows **timely intervention** before rupture.  \n- AI can **reduce radiologist workload** by automatically flagging aneurysm cases.  \n- Improves access to **equitable healthcare**, especially in resource-limited settings.\n\n\n## [2.3] Workflow for This Notebook\nThis notebook focuses on **Data Analysis and Understanding**:  \n\n✅ **Step 1: Introduction (Current Section)**  \n✅ **Step 2: Load data & inspect structure (Next)**  \n✅ **Step 3: Prepare data summaries (EDA – Coming soon)**  \n✅ **Step 4: Draw insights for modeling pipeline preparation**  ","metadata":{}},{"cell_type":"markdown","source":"### Library Imports","metadata":{}},{"cell_type":"code","source":"import os\nimport pydicom\nimport matplotlib.pyplot as plt\nimport pandas as pd\nimport nibabel as nib\nimport numpy as np\nimport seaborn as sns\n\nsns.set(style=\"whitegrid\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:02:33.000497Z","iopub.execute_input":"2025-08-29T05:02:33.000766Z","iopub.status.idle":"2025-08-29T05:02:36.939619Z","shell.execute_reply.started":"2025-08-29T05:02:33.00074Z","shell.execute_reply":"2025-08-29T05:02:36.938649Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Step 1: Segmentation (.nii) Visualization  \nSegmentation masks are 3D NifTI volumes highlighting aneurysm regions.  \nWe will load and display the middle slice from one segmentation file.\n","metadata":{}},{"cell_type":"code","source":"### Segmentation Mask\nseg_path = \"/kaggle/input/rsna-intracranial-aneurysm-detection/segmentations/1.2.826.0.1.3680043.8.498.10759842474698331813589731619457567641_cowseg.nii\"\nseg = nib.load(seg_path)\nseg_data = seg.get_fdata()\n\nplt.imshow(seg_data[:, :, seg_data.shape[2] // 2], cmap='gray')\nplt.title(\"Middle slice of aneurysm mask (.nii)\")\nplt.axis('off')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:02:36.940442Z","iopub.execute_input":"2025-08-29T05:02:36.941104Z","iopub.status.idle":"2025-08-29T05:02:37.341819Z","shell.execute_reply.started":"2025-08-29T05:02:36.941078Z","shell.execute_reply":"2025-08-29T05:02:37.340926Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Step 2: DICOM Series Visualization  \nWe will now visualize a DICOM series: stack slices in correct order and display the middle slice.\n","metadata":{}},{"cell_type":"code","source":"study_uid = \"1.2.826.0.1.3680043.8.498.10009383108068795488741533244914370182\"\nfolder = f\"/kaggle/input/rsna-intracranial-aneurysm-detection/series/{study_uid}\"\n\nslices = [pydicom.dcmread(os.path.join(folder, f)) for f in os.listdir(folder)]\nslices.sort(key=lambda x: float(x.ImagePositionPatient[2]))\nvolume = np.stack([s.pixel_array for s in slices])\n\nplt.imshow(volume[len(volume)//2], cmap='gray')\nplt.title(\"Middle slice of DICOM series\")\nplt.axis('off')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:02:37.344424Z","iopub.execute_input":"2025-08-29T05:02:37.344675Z","iopub.status.idle":"2025-08-29T05:02:46.924236Z","shell.execute_reply.started":"2025-08-29T05:02:37.344653Z","shell.execute_reply":"2025-08-29T05:02:46.923255Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Step 3: Localizer Annotations  \n`train_localizers.csv` provides exact (x,y) coordinates of aneurysms within slices.  \nWe’ll annotate an aneurysm on its corresponding DICOM slice.","metadata":{}},{"cell_type":"code","source":"localizers_df = pd.read_csv(\"/kaggle/input/rsna-intracranial-aneurysm-detection/train_localizers.csv\")\nseries_id = \"1.2.826.0.1.3680043.8.498.10491885999343016971277789732392506995\"\nlocalizers_series = localizers_df[localizers_df['SeriesInstanceUID'] == series_id]\n\nif not localizers_series.empty:\n    first = localizers_series.iloc[0]\n    sop_uid = first['SOPInstanceUID']\n    coord = eval(first['coordinates'])\n\n    slices = [pydicom.dcmread(os.path.join(f\"/kaggle/input/rsna-intracranial-aneurysm-detection/series/{series_id}\", f)) \n              for f in os.listdir(f\"/kaggle/input/rsna-intracranial-aneurysm-detection/series/{series_id}\")]\n    target_slice = next((s for s in slices if s.SOPInstanceUID == sop_uid), None)\n\n    if target_slice:\n        img = target_slice.pixel_array\n        plt.imshow(img, cmap='gray')\n        plt.annotate('Aneurysm', xy=(coord['x'], coord['y']),\n                     xytext=(coord['x']+10, coord['y']-10),\n                     arrowprops=dict(color='red', lw=1))\n        plt.title(\"Annotated aneurysm location\")\n        plt.axis('off')\n        plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:02:46.925038Z","iopub.execute_input":"2025-08-29T05:02:46.925269Z","iopub.status.idle":"2025-08-29T05:02:49.24978Z","shell.execute_reply.started":"2025-08-29T05:02:46.925249Z","shell.execute_reply":"2025-08-29T05:02:49.248886Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Step 4: Tabular Data Exploration  \nThe `train.csv` file provides:\n- Patient demographics (Age, Sex)\n- Imaging modality (CTA, MRA, MRI)\n- 13 binary aneurysm location columns\n- Overall aneurysm presence\n","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(\"/kaggle/input/rsna-intracranial-aneurysm-detection/train.csv\")\nlocalizer_df = pd.read_csv(\"/kaggle/input/rsna-intracranial-aneurysm-detection/train_localizers.csv\")\n\nprint(\"Train CSV Shape:\", train_df.shape)\nprint(\"Localizer CSV Shape:\", localizer_df.shape)\n\n# Preview training labels\ntrain_df.head()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:02:49.250682Z","iopub.execute_input":"2025-08-29T05:02:49.250973Z","iopub.status.idle":"2025-08-29T05:02:49.302244Z","shell.execute_reply.started":"2025-08-29T05:02:49.250932Z","shell.execute_reply":"2025-08-29T05:02:49.301493Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"num_positive = train_df['Aneurysm Present'].sum()\nprint(f\"✅ Number of aneurysm-positive cases: {num_positive}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:02:49.303089Z","iopub.execute_input":"2025-08-29T05:02:49.303337Z","iopub.status.idle":"2025-08-29T05:02:49.308125Z","shell.execute_reply.started":"2025-08-29T05:02:49.30331Z","shell.execute_reply":"2025-08-29T05:02:49.307469Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from IPython.display import Image, display\ndisplay(Image('/kaggle/input/photo-3/3.png'))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:02:49.309359Z","iopub.execute_input":"2025-08-29T05:02:49.309655Z","iopub.status.idle":"2025-08-29T05:02:49.346498Z","shell.execute_reply.started":"2025-08-29T05:02:49.309629Z","shell.execute_reply":"2025-08-29T05:02:49.345593Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In this section, we will analyze:\n- **Target label distributions (Aneurysm Present)**\n- **Vessel-wise aneurysm frequencies (13 anatomical locations)**\n- **Patient demographics (Age & Sex)**\n- **Imaging modalities distribution (CTA, MRA, MRI)**\n- **Sample DICOM visualization with aneurysm overlay**\n","metadata":{}},{"cell_type":"markdown","source":"## [3.1] Target Label Distribution (Aneurysm Present)","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\nplt.figure(figsize=(6,4))\nsns.countplot(data=train_df, x='Aneurysm Present', palette='Reds')\nplt.title(\"Distribution of Aneurysm Presence\")\nplt.xlabel(\"Aneurysm Present (0 = No, 1 = Yes)\")\nplt.ylabel(\"Count\")\nplt.show()\n\n# Calculate percentage\naneurysm_rate = train_df['Aneurysm Present'].mean() * 100\nprint(f\"💉 Aneurysm present in {aneurysm_rate:.2f}% of cases.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:02:49.347529Z","iopub.execute_input":"2025-08-29T05:02:49.347813Z","iopub.status.idle":"2025-08-29T05:02:49.503464Z","shell.execute_reply.started":"2025-08-29T05:02:49.34779Z","shell.execute_reply":"2025-08-29T05:02:49.50263Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## [3.2] Vessel-Wise Aneurysm Frequency","metadata":{}},{"cell_type":"code","source":"anatomy_cols = [col for col in train_df.columns if col not in \n                ['SeriesInstanceUID','Modality','PatientAge','PatientSex','Aneurysm Present']]\n\nvessel_counts = train_df[anatomy_cols].sum().sort_values(ascending=False)\n\n# Plot vessel frequency\nplt.figure(figsize=(10,6))\nsns.barplot(x=vessel_counts.values, y=vessel_counts.index, palette=\"mako\")\nplt.title(\"Frequency of Aneurysms by Vessel Location\")\nplt.xlabel(\"Number of Cases\")\nplt.ylabel(\"Vessel Location\")\nplt.show()\n\npd.DataFrame({\"Vessel Location\": vessel_counts.index, \"Count\": vessel_counts.values})\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:02:49.504282Z","iopub.execute_input":"2025-08-29T05:02:49.504495Z","iopub.status.idle":"2025-08-29T05:02:49.846278Z","shell.execute_reply.started":"2025-08-29T05:02:49.504477Z","shell.execute_reply":"2025-08-29T05:02:49.845478Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## [3.3] Patient Demographics (Age & Sex)","metadata":{}},{"cell_type":"code","source":"# Age distribution\nplt.figure(figsize=(8,4))\nsns.histplot(train_df['PatientAge'], bins=30, kde=True, color=\"teal\")\nplt.title(\"Patient Age Distribution\")\nplt.xlabel(\"Age\")\nplt.ylabel(\"Count\")\nplt.show()\n\n# Sex distribution\nplt.figure(figsize=(6,4))\nsns.countplot(data=train_df, x='PatientSex', palette=\"pastel\")\nplt.title(\"Patient Sex Distribution\")\nplt.xlabel(\"Sex (M/F)\")\nplt.ylabel(\"Count\")\nplt.show()\n\n# Age vs Aneurysm Presence\nplt.figure(figsize=(8,4))\nsns.boxplot(data=train_df, x='Aneurysm Present', y='PatientAge', palette='coolwarm')\nplt.title(\"Age vs Aneurysm Presence\")\nplt.xlabel(\"Aneurysm Present (0 = No, 1 = Yes)\")\nplt.ylabel(\"Age\")\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:02:49.8472Z","iopub.execute_input":"2025-08-29T05:02:49.847542Z","iopub.status.idle":"2025-08-29T05:02:50.520784Z","shell.execute_reply.started":"2025-08-29T05:02:49.84752Z","shell.execute_reply":"2025-08-29T05:02:50.519923Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## [3.4] Imaging Modalities (CTA vs MRA vs MRI)","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(6,4))\nsns.countplot(data=train_df, x='Modality', palette='Blues')\nplt.title(\"Distribution of Imaging Modalities\")\nplt.xlabel(\"Imaging Modality\")\nplt.ylabel(\"Count\")\nplt.show()\n\nmodality_dist = train_df['Modality'].value_counts(normalize=True) * 100\nprint(\"Modality distribution (%):\")\ndisplay(modality_dist)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:02:50.521653Z","iopub.execute_input":"2025-08-29T05:02:50.521983Z","iopub.status.idle":"2025-08-29T05:02:50.704878Z","shell.execute_reply.started":"2025-08-29T05:02:50.521955Z","shell.execute_reply":"2025-08-29T05:02:50.70408Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## [3.5] Sample DICOM Visualization","metadata":{}},{"cell_type":"code","source":"import pydicom\nimport os\nimport random\n\naneurysm_scans = train_df[train_df['Aneurysm Present'] == 1]['SeriesInstanceUID'].tolist()\nexample_uid = random.choice(aneurysm_scans)\nexample_path = f\"/kaggle/input/rsna-intracranial-aneurysm-detection/series/{example_uid}\"\n\n# List files and load a mid-slice\ndcm_files = sorted(os.listdir(example_path))\ndcm = pydicom.dcmread(f\"{example_path}/{dcm_files[len(dcm_files)//2]}\")\nimg = dcm.pixel_array\n\nplt.figure(figsize=(6,6))\nplt.imshow(img, cmap='gray')\nplt.title(f\"Sample DICOM Scan\\nSeries: {example_uid} | Modality: {dcm.Modality}\")\nplt.axis(\"off\")\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:02:50.707691Z","iopub.execute_input":"2025-08-29T05:02:50.708399Z","iopub.status.idle":"2025-08-29T05:02:51.007563Z","shell.execute_reply.started":"2025-08-29T05:02:50.708377Z","shell.execute_reply":"2025-08-29T05:02:51.006689Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## [3.6] Aneurysm Localization Overlay","metadata":{}},{"cell_type":"code","source":"coords = localizer_df[localizer_df['SeriesInstanceUID'] == example_uid]\n\nprint(f\"🔴 Found {len(coords)} aneurysm localization points for this scan.\")\n\n# Overlay aneurysm points on the DICOM slice\nplt.figure(figsize=(6,6))\nplt.imshow(img, cmap='gray')\nfor _, row in coords.iterrows():\n    x, y = eval(row['coordinates']) \n    plt.scatter(x, y, color='red', s=40, label='Aneurysm')\nplt.title(\"Aneurysm Localization Overlay\")\nplt.axis(\"off\")\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:02:51.00849Z","iopub.execute_input":"2025-08-29T05:02:51.00875Z","iopub.status.idle":"2025-08-29T05:02:51.168112Z","shell.execute_reply.started":"2025-08-29T05:02:51.008723Z","shell.execute_reply":"2025-08-29T05:02:51.167343Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📊 EDA Findings → Modeling Actions\n\n| **EDA Finding**                                    | **Observation / Insight**                                                                                   | **Modeling Action**                                                                                   |\n|----------------------------------------------------|-------------------------------------------------------------------------------------------------------------|-------------------------------------------------------------------------------------------------------|\n| **Target Imbalance (Aneurysm Present)**           | ~15-20% positive cases (majority are negative scans).                                                      | Use **class-weighted loss** (BCE with weights) or **focal loss**. <br> - Stratified train/val split. |\n| **Vessel-wise Label Distribution**                | Some vessels (e.g., Anterior Communicating Artery, MCA) dominate; others are rare.                         | Treat as **multi-label classification** with **sigmoid outputs**.<br>- Consider **label smoothing**. |\n| **Multi-label nature (13 locations + binary)**     | One scan can have multiple aneurysms in different vessels.                                                 | Multi-label **BCE loss** (one output neuron per label).                                             |\n| **Demographics (Age & Sex)**                       | Older age groups more likely positive; slight female predominance noted clinically.                        | Include **Age & Sex** as auxiliary features or embeddings during training.                          |\n| **Imaging Modality (CTA > MRA/MRI)**              | CTA is dominant (~70-80%); MRI/MRA scans fewer but present.                                                | **Modality-specific normalization**.<br>- Optionally **train modality-specific CNNs** or use modality as input feature. |\n| **Image Quality (DICOM preview)**                 | Slices vary in count/intensity across scanners and protocols.                                               | Implement **intensity normalization** & **resampling to fixed voxel size** for 3D CNN input.        |\n| **Localization (train_localizers.csv)**           | Positive scans often have 1-2 aneurysm center points annotated.                                            | Use for **weak localization** (Grad-CAM, attention maps).<br>- Optional **auxiliary localization loss**. |\n| **Segmentation Masks (subset)**                   | Vessel segmentation NIfTI files available for a subset of cases.                                           | Can be used for **vessel-masked input preprocessing** to reduce background noise.                   |\n","metadata":{}},{"cell_type":"markdown","source":"These insights directly inform our **data preprocessing** and **modeling strategy** for the next phase (3D CNN pipeline).","metadata":{}},{"cell_type":"markdown","source":"# Model Traning","metadata":{}},{"cell_type":"code","source":"import os\nimport pandas as pd\n\nseries_dir = '/kaggle/input/rsna-intracranial-aneurysm-detection/series'\nseries_folders = [f for f in os.listdir(series_dir) if os.path.isdir(os.path.join(series_dir, f))]\n\nseries_counts = []\nfor folder in series_folders:\n    dcm_files = os.listdir(os.path.join(series_dir, folder))\n    series_counts.append({'SeriesInstanceUID': folder, 'NumSlices': len(dcm_files)})\n\ndf_series = pd.DataFrame(series_counts)\ndf_series.describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:02:51.169054Z","iopub.execute_input":"2025-08-29T05:02:51.169301Z","iopub.status.idle":"2025-08-29T05:06:22.98579Z","shell.execute_reply.started":"2025-08-29T05:02:51.169283Z","shell.execute_reply":"2025-08-29T05:06:22.984962Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pydicom\n\nsample_path = os.path.join(series_dir, series_folders[0])\nsample_file = os.listdir(sample_path)[0]\ndcm = pydicom.dcmread(os.path.join(sample_path, sample_file))\n\nprint(f\"Orientation: {dcm.ImageOrientationPatient}\")\nprint(f\"Voxel spacing: {dcm.PixelSpacing}\")\nprint(f\"Slice Thickness: {dcm.SliceThickness}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:06:22.98672Z","iopub.execute_input":"2025-08-29T05:06:22.987311Z","iopub.status.idle":"2025-08-29T05:06:23.005365Z","shell.execute_reply.started":"2025-08-29T05:06:22.987282Z","shell.execute_reply":"2025-08-29T05:06:23.00456Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport matplotlib.widgets as widgets\nimport numpy as np\n\ndef load_series(series_path):\n    files = sorted(os.listdir(series_path), key=lambda x: pydicom.dcmread(os.path.join(series_path, x)).InstanceNumber)\n    images = [pydicom.dcmread(os.path.join(series_path, f)).pixel_array for f in files]\n    return np.stack(images)\n\nvolume = load_series(os.path.join(series_dir, series_folders[0]))\n\n# Scrollable plot\nfrom ipywidgets import interact\n@interact(slice=(0, volume.shape[0]-1))\ndef show_slice(slice=0):\n    plt.imshow(volume[slice], cmap='gray')\n    plt.axis('off')\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:06:23.006445Z","iopub.execute_input":"2025-08-29T05:06:23.007065Z","iopub.status.idle":"2025-08-29T05:06:25.266551Z","shell.execute_reply.started":"2025-08-29T05:06:23.007037Z","shell.execute_reply":"2025-08-29T05:06:25.265945Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport pydicom\n\n# Count slices and shape per series\ndicom_dir = '/kaggle/input/rsna-intracranial-aneurysm-detection/series'\nseries_stats = []\n\nfor series_id in os.listdir(dicom_dir)[:50]:  # limit for speed\n    series_path = os.path.join(dicom_dir, series_id)\n    files = os.listdir(series_path)\n    num_slices = len(files)\n    sample_dcm = pydicom.dcmread(os.path.join(series_path, files[0]))\n    shape = (sample_dcm.Rows, sample_dcm.Columns)\n    series_stats.append({\"SeriesInstanceUID\": series_id, \"Slices\": num_slices, \"Shape\": shape})\n\npd.DataFrame(series_stats).value_counts(\"Shape\").plot(kind=\"barh\", title=\"Common Image Resolutions\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:06:25.267393Z","iopub.execute_input":"2025-08-29T05:06:25.267685Z","iopub.status.idle":"2025-08-29T05:06:26.421381Z","shell.execute_reply.started":"2025-08-29T05:06:25.267659Z","shell.execute_reply":"2025-08-29T05:06:26.420381Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"voxel_data = []\n\nfor series_id in os.listdir(dicom_dir)[:50]:\n    slices = []\n    for f in sorted(os.listdir(os.path.join(dicom_dir, series_id))):\n        path = os.path.join(dicom_dir, series_id, f)\n        dcm = pydicom.dcmread(path)\n        slices.append(dcm)\n\n    try:\n        spacing = slices[0].PixelSpacing\n        thickness = float(slices[0].SliceThickness)\n        voxel_data.append({\n            \"SeriesInstanceUID\": series_id,\n            \"PixelSpacingX\": spacing[0],\n            \"PixelSpacingY\": spacing[1],\n            \"SliceThickness\": thickness,\n            \"NumSlices\": len(slices)\n        })\n    except:\n        continue\n\nvoxel_df = pd.DataFrame(voxel_data)\nsns.histplot(voxel_df[\"SliceThickness\"], bins=20)\nplt.title(\"Slice Thickness Distribution\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:06:26.422354Z","iopub.execute_input":"2025-08-29T05:06:26.422649Z","iopub.status.idle":"2025-08-29T05:07:42.31614Z","shell.execute_reply.started":"2025-08-29T05:06:26.422621Z","shell.execute_reply":"2025-08-29T05:07:42.31474Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"seg_dir = \"/kaggle/input/rsna-intracranial-aneurysm-detection/segmentations\"\nseg_series = [f.replace(\".nii.gz\", \"\") for f in os.listdir(seg_dir)]\nprint(\"Segmented Series:\", len(seg_series))\n\n# % of series with segmentation\nseg_percent = len(seg_series) / len(os.listdir(dicom_dir)) * 100\nprint(f\"{seg_percent:.2f}% of series have vessel segmentation.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:07:42.317312Z","iopub.execute_input":"2025-08-29T05:07:42.317596Z","iopub.status.idle":"2025-08-29T05:07:42.412089Z","shell.execute_reply.started":"2025-08-29T05:07:42.317567Z","shell.execute_reply":"2025-08-29T05:07:42.411257Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport shutil\nimport gc\nfrom collections import defaultdict\nfrom typing import Tuple, List\nimport warnings\nwarnings.filterwarnings('ignore')\n\nimport numpy as np\nimport pandas as pd\nimport polars as pl\nimport pydicom\nfrom scipy import ndimage\nfrom sklearn.preprocessing import StandardScaler\n\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torch.utils.data import Dataset, DataLoader\nimport torch.optim as optim\n\nimport kaggle_evaluation.rsna_inference_server\n\n# Competition constants\nID_COL = 'SeriesInstanceUID'\nLABEL_COLS = [\n    'Left Infraclinoid Internal Carotid Artery',\n    'Right Infraclinoid Internal Carotid Artery',\n    'Left Supraclinoid Internal Carotid Artery',\n    'Right Supraclinoid Internal Carotid Artery',\n    'Left Middle Cerebral Artery',\n    'Right Middle Cerebral Artery',\n    'Anterior Communicating Artery',\n    'Left Anterior Cerebral Artery',\n    'Right Anterior Cerebral Artery',\n    'Left Posterior Communicating Artery',\n    'Right Posterior Communicating Artery',\n    'Basilar Tip',\n    'Other Posterior Circulation',\n    'Aneurysm Present',\n]\n\nDICOM_TAG_ALLOWLIST = [\n    'BitsAllocated', 'BitsStored', 'Columns', 'FrameOfReferenceUID', 'HighBit',\n    'ImageOrientationPatient', 'ImagePositionPatient', 'InstanceNumber', 'Modality',\n    'PatientID', 'PhotometricInterpretation', 'PixelRepresentation', 'PixelSpacing',\n    'PlanarConfiguration', 'RescaleIntercept', 'RescaleSlope', 'RescaleType', 'Rows',\n    'SOPClassUID', 'SOPInstanceUID', 'SamplesPerPixel', 'SliceThickness',\n    'SpacingBetweenSlices', 'StudyInstanceUID', 'TransferSyntaxUID',\n]\n\n# Model configuration\nTARGET_SIZE = (64, 64, 64)  # Reduced size for memory efficiency\nDEVICE = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\n\nclass DICOMProcessor:\n    \"\"\"Process DICOM series into normalized 3D volumes\"\"\"\n    \n    def __init__(self, target_size: Tuple[int, int, int] = TARGET_SIZE):\n        self.target_size = target_size\n        self.scaler = StandardScaler()\n    \n    def load_dicom_series(self, series_path: str) -> np.ndarray:\n        \"\"\"Load and process a DICOM series into a 3D volume\"\"\"\n        try:\n            # Get all DICOM files\n            dicom_files = []\n            for root, _, files in os.walk(series_path):\n                for file in files:\n                    if file.endswith('.dcm'):\n                        dicom_files.append(os.path.join(root, file))\n            \n            if not dicom_files:\n                raise ValueError(f\"No DICOM files found in {series_path}\")\n            \n            # Load DICOMs and sort by instance number\n            dicoms = []\n            for filepath in dicom_files:\n                try:\n                    ds = pydicom.dcmread(filepath, force=True)\n                    if hasattr(ds, 'PixelData'):\n                        dicoms.append((ds, filepath))\n                except Exception as e:\n                    print(f\"Error reading {filepath}: {e}\")\n                    continue\n            \n            if not dicoms:\n                raise ValueError(f\"No valid DICOM files with pixel data in {series_path}\")\n            \n            # Sort by instance number\n            dicoms.sort(key=lambda x: getattr(x[0], 'InstanceNumber', 0))\n            \n            # Extract volume\n            volume_slices = []\n            for ds, _ in dicoms:\n                try:\n                    # Get pixel array\n                    pixel_array = ds.pixel_array.astype(np.float32)\n                    \n                    # Apply rescale if available\n                    if hasattr(ds, 'RescaleSlope') and hasattr(ds, 'RescaleIntercept'):\n                        slope = float(ds.RescaleSlope)\n                        intercept = float(ds.RescaleIntercept)\n                        pixel_array = pixel_array * slope + intercept\n                    \n                    volume_slices.append(pixel_array)\n                except Exception as e:\n                    print(f\"Error processing slice: {e}\")\n                    continue\n            \n            if not volume_slices:\n                raise ValueError(\"No valid slices extracted\")\n            \n            # Stack into 3D volume\n            volume = np.stack(volume_slices, axis=0)  # Shape: (depth, height, width)\n            \n            # Normalize and resize\n            volume = self.preprocess_volume(volume)\n            \n            return volume\n            \n        except Exception as e:\n            print(f\"Error processing series {series_path}: {e}\")\n            # Return zeros if processing fails\n            return np.zeros(self.target_size, dtype=np.float32)\n    \n    def preprocess_volume(self, volume: np.ndarray) -> np.ndarray:\n        \"\"\"Preprocess 3D volume: normalize, clip, resize\"\"\"\n        # Handle potential issues\n        if volume.size == 0:\n            return np.zeros(self.target_size, dtype=np.float32)\n        \n        # Clip extreme values (robust to outliers)\n        p1, p99 = np.percentile(volume, [1, 99])\n        volume = np.clip(volume, p1, p99)\n        \n        # Normalize to [0, 1]\n        volume_min, volume_max = volume.min(), volume.max()\n        if volume_max > volume_min:\n            volume = (volume - volume_min) / (volume_max - volume_min)\n        \n        # Resize to target size\n        if volume.shape != self.target_size:\n            zoom_factors = [\n                self.target_size[i] / volume.shape[i] for i in range(3)\n            ]\n            volume = ndimage.zoom(volume, zoom_factors, order=1)\n        \n        return volume.astype(np.float32)\n\nclass Simple3DCNN(nn.Module):\n    \"\"\"Lightweight 3D CNN for aneurysm detection\"\"\"\n    \n    def __init__(self, num_classes: int = len(LABEL_COLS)):\n        super(Simple3DCNN, self).__init__()\n        \n        # 3D Convolutional layers\n        self.conv1 = nn.Conv3d(1, 16, kernel_size=3, padding=1)\n        self.pool1 = nn.MaxPool3d(2)\n        self.conv2 = nn.Conv3d(16, 32, kernel_size=3, padding=1)\n        self.pool2 = nn.MaxPool3d(2)\n        self.conv3 = nn.Conv3d(32, 64, kernel_size=3, padding=1)\n        self.pool3 = nn.MaxPool3d(2)\n        self.conv4 = nn.Conv3d(64, 128, kernel_size=3, padding=1)\n        self.pool4 = nn.MaxPool3d(2)\n        \n        # Adaptive pooling to handle variable sizes\n        self.adaptive_pool = nn.AdaptiveAvgPool3d((2, 2, 2))\n        \n        # Fully connected layers\n        self.fc1 = nn.Linear(128 * 2 * 2 * 2, 256)\n        self.dropout1 = nn.Dropout(0.5)\n        self.fc2 = nn.Linear(256, 128)\n        self.dropout2 = nn.Dropout(0.3)\n        self.fc3 = nn.Linear(128, num_classes)\n        \n        # Batch normalization\n        self.bn1 = nn.BatchNorm3d(16)\n        self.bn2 = nn.BatchNorm3d(32)\n        self.bn3 = nn.BatchNorm3d(64)\n        self.bn4 = nn.BatchNorm3d(128)\n        \n    def forward(self, x):\n        # Input shape: (batch_size, 1, depth, height, width)\n        x = self.pool1(F.relu(self.bn1(self.conv1(x))))\n        x = self.pool2(F.relu(self.bn2(self.conv2(x))))\n        x = self.pool3(F.relu(self.bn3(self.conv3(x))))\n        x = self.pool4(F.relu(self.bn4(self.conv4(x))))\n        \n        # Adaptive pooling\n        x = self.adaptive_pool(x)\n        \n        # Flatten\n        x = x.view(x.size(0), -1)\n        \n        # Fully connected layers\n        x = F.relu(self.fc1(x))\n        x = self.dropout1(x)\n        x = F.relu(self.fc2(x))\n        x = self.dropout2(x)\n        x = self.fc3(x)\n        \n        return torch.sigmoid(x)\n\nclass AneurysmDataset(Dataset):\n    \"\"\"Dataset for loading training data\"\"\"\n    \n    def __init__(self, data_df: pd.DataFrame, series_dir: str, processor: DICOMProcessor):\n        self.data_df = data_df\n        self.series_dir = series_dir\n        self.processor = processor\n        \n    def __len__(self):\n        return len(self.data_df)\n    \n    def __getitem__(self, idx):\n        row = self.data_df.iloc[idx]\n        series_id = row[ID_COL]\n        \n        # Load volume\n        series_path = os.path.join(self.series_dir, series_id)\n        volume = self.processor.load_dicom_series(series_path)\n        \n        # Get labels\n        labels = row[LABEL_COLS].values.astype(np.float32)\n        \n        # Convert to tensor and add channel dimension\n        volume_tensor = torch.from_numpy(volume).unsqueeze(0)  # Add channel dim\n        labels_tensor = torch.from_numpy(labels)\n        \n        return volume_tensor, labels_tensor\n\n# Global model and processor\nmodel = None\nprocessor = None\n\ndef initialize_model():\n    \"\"\"Initialize model and processor (called once)\"\"\"\n    global model, processor\n    \n    if model is not None:\n        return\n    \n    print(\"Initializing model...\")\n    processor = DICOMProcessor(TARGET_SIZE)\n    model = Simple3DCNN(num_classes=len(LABEL_COLS))\n    \n    # Load pre-trained weights if available\n    try:\n        if os.path.exists('/kaggle/input/model_weights.pth'):\n            model.load_state_dict(torch.load('/kaggle/input/model_weights.pth', map_location='cpu'))\n            print(\"Loaded pre-trained weights\")\n        else:\n            print(\"No pre-trained weights found, using random initialization\")\n    except Exception as e:\n        print(f\"Error loading weights: {e}\")\n    \n    model.to(DEVICE)\n    model.eval()\n    print(f\"Model initialized on {DEVICE}\")\n\ndef predict(series_path: str) -> pl.DataFrame:\n    \"\"\"Make prediction for a single series\"\"\"\n    \n    # Initialize model on first call\n    initialize_model()\n    \n    series_id = os.path.basename(series_path)\n    \n    try:\n        # Process the DICOM series\n        volume = processor.load_dicom_series(series_path)\n        \n        # Convert to tensor and add batch dimension\n        volume_tensor = torch.from_numpy(volume).unsqueeze(0).unsqueeze(0)  # (1, 1, D, H, W)\n        volume_tensor = volume_tensor.to(DEVICE)\n        \n        # Make prediction\n        with torch.no_grad():\n            predictions = model(volume_tensor)\n            predictions = predictions.cpu().numpy().flatten()\n        \n        # Create result DataFrame\n        result_data = [[series_id] + predictions.tolist()]\n        result_df = pl.DataFrame(\n            data=result_data,\n            schema=[ID_COL] + LABEL_COLS,\n            orient='row'\n        )\n        \n        # Clean up memory\n        del volume_tensor\n        torch.cuda.empty_cache()\n        gc.collect()\n        \n    except Exception as e:\n        print(f\"Error predicting for series {series_id}: {e}\")\n        # Return baseline predictions (0.5 for all classes)\n        result_data = [[series_id] + [0.5] * len(LABEL_COLS)]\n        result_df = pl.DataFrame(\n            data=result_data,\n            schema=[ID_COL] + LABEL_COLS,\n            orient='row'\n        )\n    \n    # Mandatory cleanup\n    shutil.rmtree('/kaggle/shared', ignore_errors=True)\n    \n    return result_df.drop(ID_COL)\n\ndef train_model(train_df_path: str, series_dir: \"/kaggle/input/rsna-intracranial-aneurysm-detection/series\", num_epochs: int = 50, batch_size: int = 32):\n    \"\"\"Training function (for reference - would be run separately)\"\"\"\n    \n    # Load training data\n    train_df = pd.read_csv(\"/kaggle/input/rsna-intracranial-aneurysm-detection/train.csv\")\n    \n    # Initialize components\n    processor = DICOMProcessor(TARGET_SIZE)\n    model = Simple3DCNN(num_classes=len(LABEL_COLS))\n    model.to(DEVICE)\n    \n    # Create dataset and dataloader\n    dataset = AneurysmDataset(train_df, series_dir, processor)\n    dataloader = DataLoader(dataset, batch_size=batch_size, shuffle=True, num_workers=0)\n    \n    # Loss and optimizer\n    criterion = nn.BCELoss()\n    optimizer = optim.Adam(model.parameters(), lr=0.001, weight_decay=1e-5)\n    scheduler = optim.lr_scheduler.ReduceLROnPlateau(optimizer, mode='min', patience=2)\n    \n    # Training loop\n    model.train()\n    for epoch in range(num_epochs):\n        total_loss = 0\n        num_batches = 0\n        \n        for batch_idx, (volumes, labels) in enumerate(dataloader):\n            volumes, labels = volumes.to(DEVICE), labels.to(DEVICE)\n            \n            optimizer.zero_grad()\n            outputs = model(volumes)\n            loss = criterion(outputs, labels)\n            loss.backward()\n            optimizer.step()\n            \n            total_loss += loss.item()\n            num_batches += 1\n            \n            if batch_idx % 10 == 0:\n                print(f'Epoch {epoch+1}/{num_epochs}, Batch {batch_idx+1}, Loss: {loss.item():.4f}')\n        \n        avg_loss = total_loss / num_batches\n        scheduler.step(avg_loss)\n        print(f'Epoch {epoch+1}/{num_epochs} completed, Average Loss: {avg_loss:.4f}')\n    \n    # Save model\n    torch.save(model.state_dict(), 'model_weights.pth')\n    print(\"Model saved to model_weights.pth\")\n\n# Competition server setup\ninference_server = kaggle_evaluation.rsna_inference_server.RSNAInferenceServer(predict)\n\nif os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n    inference_server.serve()\nelse:\n    inference_server.run_local_gateway()\n    display(pl.read_parquet('/kaggle/working/submission.parquet'))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-29T05:07:42.412986Z","iopub.execute_input":"2025-08-29T05:07:42.413241Z","iopub.status.idle":"2025-08-29T05:08:11.295929Z","shell.execute_reply.started":"2025-08-29T05:07:42.413222Z","shell.execute_reply":"2025-08-29T05:08:11.295115Z"}},"outputs":[],"execution_count":null}]}