{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":99552,"databundleVersionId":13190393,"sourceType":"competition"}],"dockerImageVersionId":31090,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Importing Libraries","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport os\nimport glob\nfrom pathlib import Path\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport ast\nimport json\nfrom collections import Counter\n\nimport pydicom\nimport nibabel as nib\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom matplotlib.patches import Circle\nimport warnings\nwarnings.filterwarnings('ignore')\n\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torch.utils.data import Dataset, DataLoader\nimport torchvision.transforms as transforms\nfrom sklearn.model_selection import train_test_split\nimport cv2\n\nimport torch.optim as optim\nfrom torch.optim.lr_scheduler import StepLR\nimport time\nfrom tqdm import tqdm","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:22:48.981989Z","iopub.execute_input":"2025-08-11T15:22:48.982283Z","iopub.status.idle":"2025-08-11T15:22:48.987627Z","shell.execute_reply.started":"2025-08-11T15:22:48.982263Z","shell.execute_reply":"2025-08-11T15:22:48.987048Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n**An intracranial aneurysm is a bulging or ballooning of a blood vessel in the brain. If undetected, it can rupture — leading to life-threatening bleeding. Early detection is crucial. The challenge? These aneurysms are often subtle, hidden in complex anatomy, and differ widely between patients.That’s where machine learning can step in — augmenting radiologist workflows by detecting patterns across thousands of image slices in a fraction of the time.**\n","metadata":{}},{"cell_type":"markdown","source":"**What Are We Predicting?**\nTwo key tasks define this competition:\n* **Classification** — Does this scan contain an aneurysm?\n* **Localization** — If yes, where is it located? (3D coordinates + 1 of 13 brain regions)\n\n**Dataset Overview**\n* train.csv — Aneurysm presence and location labels\n* train_localizers.csv — 3D coordinates for aneurysms (subset)\n* DICOM folders — Raw scan data (one folder per patient)\n* segmentations — Optional vessel masks in .nii.gz format","metadata":{}},{"cell_type":"markdown","source":"# Data Exploration and Setup\n- Sets up file paths for the Kaggle dataset\n- Loads train.csv and train_localizers.csv files\n- Explores the directory structure\n- Counts DICOM (.dcm) and NIfTI (.nii) files\n- Provides overview of data organization\n\n**Key outputs:**\n\n- Dataset dimensions and column names\n- Number of segmentation folders and files\n- Total DICOM files across all series\n- Sample of folder structures with file counts","metadata":{}},{"cell_type":"markdown","source":"## Analysis of train","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/rsna-intracranial-aneurysm-detection/train.csv')\ntrain_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:18:55.609219Z","iopub.execute_input":"2025-08-11T15:18:55.6095Z","iopub.status.idle":"2025-08-11T15:18:55.661583Z","shell.execute_reply.started":"2025-08-11T15:18:55.609481Z","shell.execute_reply":"2025-08-11T15:18:55.660807Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Aneurysm Counts per Location\n\nWhere do aneurysms tend to occur? This bar chart answers that — using the 13 location_* columns.","metadata":{}},{"cell_type":"code","source":"# Filter and count data\nlocation_cols = [col for col in train_df.columns if col not in ['SeriesInstanceUID', 'PatientAge', 'PatientSex', 'Modality', 'Aneurysm Present']]\nlocation_counts = train_df[location_cols].sum().sort_values(ascending=False)\n\n# Set style (dark theme similar to plotly_dark)\nplt.style.use('dark_background')\nsns.set_palette(sns.color_palette([\"#00BFC4\", \"#C77CFF\"], n_colors=len(location_counts)))\n\n# Create the plot\nplt.figure(figsize=(10, 6))\nbars = plt.bar(location_counts.index, location_counts.values)\n\n# Add gradient coloring (manual approximation)\nfor i, bar in enumerate(bars):\n    # Interpolate between #00BFC4 (start) and #C77CFF (end)\n    ratio = i / len(bars)\n    r = int(0x00 * (1 - ratio) + 0xC7 * ratio)\n    g = int(0xBF * (1 - ratio) + 0x7C * ratio)\n    b = int(0xC4 * (1 - ratio) + 0xFF * ratio)\n    bar.set_color(f'#{r:02x}{g:02x}{b:02x}')\n\n# Add labels and title\nplt.title('Aneurysm Count by Location', fontsize=14, pad=20)\nplt.xlabel('Location', fontsize=12)\nplt.ylabel('Count', fontsize=12)\nplt.xticks(rotation=45, ha='right')\n\n# Adjust layout\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:23:24.259752Z","iopub.execute_input":"2025-08-11T15:23:24.26008Z","iopub.status.idle":"2025-08-11T15:23:24.520055Z","shell.execute_reply.started":"2025-08-11T15:23:24.260063Z","shell.execute_reply":"2025-08-11T15:23:24.519305Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Imaging Modality Distribution","metadata":{}},{"cell_type":"code","source":"# Get modality counts\nmodality_counts = train_df['Modality'].value_counts()\n\n# Set dark theme (similar to plotly_dark)\nplt.style.use('dark_background')\n\n# Create pie chart\nfig, ax = plt.subplots(figsize=(8, 6))\nwedges, texts, autotexts = ax.pie(\n    modality_counts.values,\n    labels=modality_counts.index,\n    colors=['#00BFC4', '#C77CFF'],  # Custom colors\n    autopct='%1.1f%%',              # Show percentages\n    startangle=90,                  # Start at top\n    wedgeprops={'linewidth': 1, 'edgecolor': 'black'},  # Edge styling\n    pctdistance=0.8,                # Move percentages inward\n    textprops={'color': 'white'},   # Text color\n)\n\n# Add a hole (donut chart)\ncentre_circle = plt.Circle((0, 0), 0.4, fc='#1f1f1f')  # Dark background for hole\nax.add_artist(centre_circle)\n\n# Equal aspect ratio ensures pie is drawn as a circle\nax.axis('equal')\n\n# Add title\nplt.title('Imaging Modality Distribution', pad=20, fontsize=14)\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:24:47.407921Z","iopub.execute_input":"2025-08-11T15:24:47.408239Z","iopub.status.idle":"2025-08-11T15:24:47.538858Z","shell.execute_reply.started":"2025-08-11T15:24:47.408205Z","shell.execute_reply":"2025-08-11T15:24:47.53807Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Age & Sex Demographics","metadata":{}},{"cell_type":"code","source":"plt.style.use('dark_background')\nplt.figure(figsize=(10,6))\nsns.histplot(\n    data=train_df,\n    x='PatientAge',\n    hue='Aneurysm Present',\n    bins=30,\n    kde=True,\n    palette={0: '#00BFC4', 1: '#C77CFF'}\n)\nplt.title(\"Age Distribution by Aneurysm Presence\", fontsize=14)\nplt.xlabel(\"Age\")\nplt.ylabel(\"Frequency\")\nplt.grid(alpha=0.2)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:25:22.788996Z","iopub.execute_input":"2025-08-11T15:25:22.789724Z","iopub.status.idle":"2025-08-11T15:25:23.155791Z","shell.execute_reply.started":"2025-08-11T15:25:22.789701Z","shell.execute_reply":"2025-08-11T15:25:23.154945Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pd.crosstab(train_df['PatientSex'], train_df['Aneurysm Present'], normalize='index') * 100","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:25:59.501043Z","iopub.execute_input":"2025-08-11T15:25:59.501639Z","iopub.status.idle":"2025-08-11T15:25:59.52878Z","shell.execute_reply.started":"2025-08-11T15:25:59.501614Z","shell.execute_reply":"2025-08-11T15:25:59.528215Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Class Imbalance: Do Most Patients Have Aneurysms?","metadata":{}},{"cell_type":"code","source":"# Set dark theme\nplt.style.use('dark_background')\n\n# Get counts\ncounts = train_df['Aneurysm Present'].value_counts().sort_index()\n\n# Define colors\ncolors = ['#00BFC4', '#C77CFF']\nlabels = [\"No Aneurysm\", \"Aneurysm\"]\n\n# Create figure\nfig, ax = plt.subplots(figsize=(8, 6))\n\n# Plot bars\nbars = ax.bar(labels, counts, color=colors, edgecolor='white', linewidth=1)\n\n# Add count labels on top of bars\nfor bar in bars:\n    height = bar.get_height()\n    ax.text(bar.get_x() + bar.get_width()/2., height,\n            f'{int(height)}',\n            ha='center', va='bottom', color='white', fontsize=12)\n\n# Customize axes and title\nax.set_title('Class Imbalance: Any Aneurysm Present', pad=20, fontsize=14)\nax.set_ylabel('Count', fontsize=12)\nax.grid(axis='y', alpha=0.3)\n\n# Remove top/right spines\nfor spine in ['top', 'right']:\n    ax.spines[spine].set_visible(False)\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:27:38.423939Z","iopub.execute_input":"2025-08-11T15:27:38.424238Z","iopub.status.idle":"2025-08-11T15:27:38.562052Z","shell.execute_reply.started":"2025-08-11T15:27:38.424216Z","shell.execute_reply":"2025-08-11T15:27:38.561461Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Analysis on Train Localizer","metadata":{}},{"cell_type":"code","source":"localizers_df = pd.read_csv('/kaggle/input/rsna-intracranial-aneurysm-detection/train_localizers.csv')\n\n# Convert coordinate strings to dicts\nlocalizers_df['coords'] = localizers_df['coordinates'].apply(ast.literal_eval)\nlocalizers_df['x'] = localizers_df['coords'].apply(lambda d: d['x'])\nlocalizers_df['y'] = localizers_df['coords'].apply(lambda d: d['y'])\n\nlocalizers_df.drop(columns=['coordinates', 'coords'], inplace=True)\nlocalizers_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:29:07.791958Z","iopub.execute_input":"2025-08-11T15:29:07.792266Z","iopub.status.idle":"2025-08-11T15:29:07.860907Z","shell.execute_reply.started":"2025-08-11T15:29:07.792244Z","shell.execute_reply":"2025-08-11T15:29:07.860311Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Heatmap of Aneurysm Locations in Image Space","metadata":{}},{"cell_type":"code","source":"plt.style.use('dark_background')\nsns.set_style(\"dark\")\n\n# Create figure\nplt.figure(figsize=(10, 8))\n\n# Create 2D histogram (heatmap)\nheatmap, xedges, yedges = np.histogram2d(\n    localizers_df['x'],\n    localizers_df['y'],\n    bins=(50, 50)  # Matching nbinsx=50, nbinsy=50\n)\n\n# Plot with Seaborn\nax = sns.heatmap(\n    heatmap.T,  # Transpose to match orientation\n    cmap='turbo',  # Matches 'Turbo' colorscale\n    square=True,\n    cbar_kws={'label': 'Density'}\n)\n\n# Reverse y-axis to match autorange=\"reversed\"\nax.invert_yaxis()\n\n# Add labels and title\nax.set_title('🧠 Heatmap of Aneurysm Locations in Image Space', pad=20, fontsize=14)\nax.set_xlabel('x')\nax.set_ylabel('y')\n\n# Adjust layout\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:31:06.355768Z","iopub.execute_input":"2025-08-11T15:31:06.356056Z","iopub.status.idle":"2025-08-11T15:31:06.962971Z","shell.execute_reply.started":"2025-08-11T15:31:06.356036Z","shell.execute_reply":"2025-08-11T15:31:06.962088Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 2D Scatter of Aneurysm Coordinates by Location","metadata":{}},{"cell_type":"code","source":"# Set dark theme\nplt.style.use('dark_background')\npalette = sns.color_palette(\"husl\", n_colors=len(localizers_df['location'].unique()))\n\n# Create figure\nplt.figure(figsize=(10, 8))\n\n# Create scatter plot\nscatter = sns.scatterplot(\n    data=localizers_df,\n    x='x',\n    y='y',\n    hue='location',\n    palette=palette,\n    s=60,  # Marker size\n    edgecolor='white',\n    linewidth=0.3\n)\n\n# Reverse y-axis\nplt.gca().invert_yaxis()\n\n# Add title and labels\nplt.title('2D Scatter of Aneurysm Coordinates by Location', pad=20, fontsize=14)\nplt.xlabel('x')\nplt.ylabel('y')\n\n# Customize legend\nplt.legend(\n    title='Location',\n    bbox_to_anchor=(1.05, 1),\n    loc='upper left',\n    borderaxespad=0\n)\n\n# Adjust grid\nplt.grid(alpha=0.2)\n\n# Equal aspect ratio\nplt.gca().set_aspect('equal')\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:47:25.665217Z","iopub.execute_input":"2025-08-11T15:47:25.665498Z","iopub.status.idle":"2025-08-11T15:47:26.286293Z","shell.execute_reply.started":"2025-08-11T15:47:25.665477Z","shell.execute_reply":"2025-08-11T15:47:26.285543Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Aneurysm Co-occurrence Across Brain Regions","metadata":{}},{"cell_type":"markdown","source":"Each patient in the dataset can have multiple aneurysms in different brain arteries. But are certain aneurysm sites more likely to occur together? To investigate this, we treat the location columns as binary labels and compute a co-occurrence matrix — it tells us how often, for instance, an aneurysm in the Left MCA is also accompanied by one in the ACoA, and so on.","metadata":{}},{"cell_type":"code","source":"df = train_df\n# Identify location columns correctly\nlocation_cols = df.columns[4:-1]  # skip UID, Age, Sex, Modality, and skip final label\nlocation_df = df[location_cols].astype(int)  # just in case they're still object type\n\n# Co-occurrence matrix\nco_matrix = location_df.T.dot(location_df)\n\n# Plot heatmap\nplt.figure(figsize=(12, 10))\nsns.heatmap(co_matrix, cmap=\"magma\", annot=True, fmt=\".0f\", linewidths=0.5)\nplt.title(\"Aneurysm Co-occurrence Matrix\", fontsize=16, color='white')\nplt.xticks(rotation=45, ha='right', fontsize=9)\nplt.yticks(rotation=0, fontsize=9)\nplt.gca().set_facecolor('black')\nplt.gcf().set_facecolor('#111111')\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:41:03.3785Z","iopub.execute_input":"2025-08-11T15:41:03.378774Z","iopub.status.idle":"2025-08-11T15:41:04.12229Z","shell.execute_reply.started":"2025-08-11T15:41:03.378754Z","shell.execute_reply":"2025-08-11T15:41:04.121576Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.style.use('dark_background')\n# Filter only rows where aneurysm is present\ndf_pos = df[df[\"Aneurysm Present\"] == 1]\n\n# List of all aneurysm location columns\nlocation_cols = df_pos.columns[4:-1]  # From 'Left Infraclinoid...' to 'Other Posterior Circulation'\n\n# Group by PatientSex and sum each location\nsex_location = df_pos.groupby(\"PatientSex\")[location_cols].sum().T\n\n# Reset index for plotting\nsex_location = sex_location.reset_index().melt(id_vars=\"index\", var_name=\"Sex\", value_name=\"Count\")\nsex_location = sex_location.rename(columns={\"index\": \"Location\"})\n\n# Plot\nplt.figure(figsize=(16, 6))\nsns.barplot(data=sex_location, x=\"Location\", y=\"Count\", hue=\"Sex\")\nplt.xticks(rotation=45, ha=\"right\")\nplt.title(\"Aneurysm Location Frequency by Sex\")\nplt.ylabel(\"Aneurysm Count\")\nplt.xlabel(\"Location\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:47:16.165934Z","iopub.execute_input":"2025-08-11T15:47:16.166457Z","iopub.status.idle":"2025-08-11T15:47:16.583475Z","shell.execute_reply.started":"2025-08-11T15:47:16.166432Z","shell.execute_reply":"2025-08-11T15:47:16.582727Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# List of location columns\nlocation_cols = df.columns[4:-1]\n\n# Filter only positive cases\ndf_pos = df[df[\"Aneurysm Present\"] == 1]\n\n# Prepare a DataFrame: for each location, collect patients with that location = 1\nage_location = []\n\nfor loc in location_cols:\n    subset = df_pos[df_pos[loc] == 1][[\"PatientAge\"]].copy()\n    subset[\"Location\"] = loc\n    age_location.append(subset)\n\n# Combine all into one DataFrame\nage_location_df = pd.concat(age_location)\n\n# Plot\nplt.figure(figsize=(16, 6))\nsns.boxplot(data=age_location_df, x=\"Location\", y=\"PatientAge\", palette=\"crest\")\nplt.xticks(rotation=45, ha=\"right\")\nplt.title(\"Patient Age Distribution per Aneurysm Location\")\nplt.ylabel(\"Age\")\nplt.xlabel(\"Location\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:48:06.939653Z","iopub.execute_input":"2025-08-11T15:48:06.940127Z","iopub.status.idle":"2025-08-11T15:48:07.354449Z","shell.execute_reply.started":"2025-08-11T15:48:06.940103Z","shell.execute_reply":"2025-08-11T15:48:07.353659Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Filter rows with aneurysm present\ndf_pos = df[df[\"Aneurysm Present\"] == 1].copy()\n\n# Initialize a dictionary to store counts\nmodality_counts = {}\n\nfor loc in location_cols:\n    subset = df_pos[df_pos[loc] == 1]\n    modality_distribution = subset[\"Modality\"].value_counts()\n    modality_counts[loc] = modality_distribution\n\n# Convert to DataFrame and fill missing values with 0\nmodality_df = pd.DataFrame(modality_counts).T.fillna(0).astype(int)\nmodality_df = modality_df[[\"CTA\", \"MRA\"]] if \"CTA\" in modality_df.columns and \"MRA\" in modality_df.columns else modality_df\n\n# Plot heatmap\nplt.figure(figsize=(12, 6))\nsns.heatmap(modality_df.T, annot=True, cmap=\"crest\", fmt=\"d\")\nplt.title(\"Modality Preference per Aneurysm Location\")\nplt.xlabel(\"Location\")\nplt.ylabel(\"Modality\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:48:43.970867Z","iopub.execute_input":"2025-08-11T15:48:43.971707Z","iopub.status.idle":"2025-08-11T15:48:44.405477Z","shell.execute_reply.started":"2025-08-11T15:48:43.971681Z","shell.execute_reply":"2025-08-11T15:48:44.404662Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_localizations = pd.read_csv(\"/kaggle/input/rsna-intracranial-aneurysm-detection/train_localizers.csv\")\n# Parse coordinates\ntrain_localizations[\"coord_dict\"] = train_localizations[\"coordinates\"].apply(ast.literal_eval)\ntrain_localizations[\"x\"] = train_localizations[\"coord_dict\"].apply(lambda d: d[\"x\"])\ntrain_localizations[\"y\"] = train_localizations[\"coord_dict\"].apply(lambda d: d[\"y\"])\n\n# Top 4-5 most frequent locations\ntop_locations = train_localizations[\"location\"].value_counts().head(5).index\n\n# Plot heatmap per location\nfor loc in top_locations:\n    subset = train_localizations[train_localizations[\"location\"] == loc]\n    plt.figure(figsize=(6, 5))\n    sns.kdeplot(data=subset, x=\"x\", y=\"y\", fill=True, cmap=\"crest\")\n    plt.title(f\"Spatial Heatmap for {loc}\")\n    plt.xlabel(\"X Coordinate\")\n    plt.ylabel(\"Y Coordinate\")\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:49:33.178908Z","iopub.execute_input":"2025-08-11T15:49:33.179459Z","iopub.status.idle":"2025-08-11T15:49:35.60014Z","shell.execute_reply.started":"2025-08-11T15:49:33.179433Z","shell.execute_reply":"2025-08-11T15:49:35.59947Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Dataset Exploration and Setup\n# Set up paths\nBASE_PATH = \"/kaggle/input/rsna-intracranial-aneurysm-detection\"\nTRAIN_CSV = f\"{BASE_PATH}/train.csv\"\nTRAIN_LOCALIZERS_CSV = f\"{BASE_PATH}/train_localizers.csv\"\nSEGMENTATIONS_PATH = f\"{BASE_PATH}/segmentations\"\nSERIES_PATH = f\"{BASE_PATH}/series\"\n\n# Load CSV files\nprint(\"Loading CSV files...\")\n\ntrain_df = pd.read_csv(TRAIN_CSV)\nprint(f\"Train.csv shape: {train_df.shape}\")\nprint(f\"Train.csv columns: {train_df.columns.tolist()}\")\nprint(\"\\nTrain.csv head:\")\ntrain_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:50:52.593727Z","iopub.execute_input":"2025-08-11T15:50:52.594005Z","iopub.status.idle":"2025-08-11T15:50:52.619855Z","shell.execute_reply.started":"2025-08-11T15:50:52.593985Z","shell.execute_reply":"2025-08-11T15:50:52.619283Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_localizers_df = pd.read_csv(TRAIN_LOCALIZERS_CSV)\nprint(f\"\\nTrain_localizers.csv shape: {train_localizers_df.shape}\")\nprint(f\"Train_localizers.csv columns: {train_localizers_df.columns.tolist()}\")\nprint(\"\\nTrain_localizers.csv head:\")\ntrain_localizers_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:50:57.945559Z","iopub.execute_input":"2025-08-11T15:50:57.946209Z","iopub.status.idle":"2025-08-11T15:50:57.967544Z","shell.execute_reply.started":"2025-08-11T15:50:57.946185Z","shell.execute_reply":"2025-08-11T15:50:57.966923Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Explore directory structure\nprint(\"\\n\" + \"=\"*50)\nprint(\"Segmentation Directory STRUCTURE ANALYSIS\")\nprint(\"=\"*50)\n\n# Count segmentation files\nseg_folders = glob.glob(f\"{SEGMENTATIONS_PATH}/*\")\nprint(f\"Number of segmentation folders: {len(seg_folders)}\")\n\nnii_files = glob.glob(f\"{SEGMENTATIONS_PATH}/*/*.nii\")\ncowseg_files = glob.glob(f\"{SEGMENTATIONS_PATH}/*/*_cowseg.nii\")\nregular_nii = [f for f in nii_files if not f.endswith('_cowseg.nii')]\n\nprint(f\"Total .nii files: {len(nii_files)}\")\nprint(f\"Cowseg files: {len(cowseg_files)}\")\nprint(f\"Regular .nii files: {len(regular_nii)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:51:02.007997Z","iopub.execute_input":"2025-08-11T15:51:02.008607Z","iopub.status.idle":"2025-08-11T15:51:03.235481Z","shell.execute_reply.started":"2025-08-11T15:51:02.008582Z","shell.execute_reply":"2025-08-11T15:51:03.23471Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"\\n\" + \"=\"*50)\nprint(\"Series Directory STRUCTURE ANALYSIS\")\nprint(\"=\"*50)\n# Count DICOM series\nseries_folders = glob.glob(f\"{SERIES_PATH}/*\")\nprint(f\"Number of series folders: {len(series_folders)}\")\n\ndcm_files = glob.glob(f\"{SERIES_PATH}/*/*.dcm\")\nprint(f\"Total .dcm files: {len(dcm_files)}\")\n\nif len(series_folders) > 0:\n    # Sample a few folders to see DICOM count distribution\n    sample_folders = series_folders[:5]\n    for folder in sample_folders:\n        dcm_count = len(glob.glob(f\"{folder}/*.dcm\"))\n        folder_name = os.path.basename(folder)\n        print(f\"  {folder_name}: {dcm_count} DICOM files\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:51:08.20793Z","iopub.execute_input":"2025-08-11T15:51:08.208643Z","iopub.status.idle":"2025-08-11T15:54:50.197587Z","shell.execute_reply.started":"2025-08-11T15:51:08.208618Z","shell.execute_reply":"2025-08-11T15:54:50.196714Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data Analysis\n- Parses coordinate strings from train_localizers.csv\n- Analyzes aneurysm location distribution\n- Calculates coordinate statistics (x,y ranges)\n- Creates visualizations of location frequency\n- Shows spatial distribution of aneurysm coordinates\n\n**Key insights:**\n\n- Which anatomical locations have most/least aneurysms\n- Coordinate ranges (helps determine image dimensions)\n- Data balance across different aneurysm types\n- Scatter plot showing where aneurysms typically appear","metadata":{}},{"cell_type":"code","source":"#Detailed Data Analysis\nprint(\"DETAILED DATA ANALYSIS\")\nprint(\"=\"*50)\n\n# Analyze train_localizers.csv in detail\nif 'train_localizers_df' in locals():\n    print(\"Analyzing localizer annotations...\")\n    \n    # Parse coordinates\n    def parse_coordinates(coord_str):\n        try:\n            if isinstance(coord_str, str):\n                coord_dict = ast.literal_eval(coord_str)\n                return coord_dict['x'], coord_dict['y']\n            return None, None\n        except:\n            return None, None\n    \n    train_localizers_df['x_coord'] = train_localizers_df['coordinates'].apply(lambda x: parse_coordinates(x)[0])\n    train_localizers_df['y_coord'] = train_localizers_df['coordinates'].apply(lambda x: parse_coordinates(x)[1])\n    \n    # Analyze locations\n    location_counts = train_localizers_df['location'].value_counts()\n    print(f\"\\nAneurysm locations distribution:\")\n    print(location_counts)\n    \n    # Coordinate statistics\n    print(f\"\\nCoordinate statistics:\")\n    print(f\"X coordinates - Min: {train_localizers_df['x_coord'].min():.2f}, Max: {train_localizers_df['x_coord'].max():.2f}\")\n    print(f\"Y coordinates - Min: {train_localizers_df['y_coord'].min():.2f}, Max: {train_localizers_df['y_coord'].max():.2f}\")\n    \n    # Plot location distribution\n    plt.figure(figsize=(12, 6))\n    plt.subplot(1, 2, 1)\n    location_counts.plot(kind='bar', rot=45)\n    plt.title('Aneurysm Location Distribution')\n    plt.tight_layout()\n    \n    plt.subplot(1, 2, 2)\n    plt.scatter(train_localizers_df['x_coord'], train_localizers_df['y_coord'], alpha=0.6)\n    plt.xlabel('X Coordinate')\n    plt.ylabel('Y Coordinate')\n    plt.title('Aneurysm Coordinate Distribution')\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:54:50.198974Z","iopub.execute_input":"2025-08-11T15:54:50.199795Z","iopub.status.idle":"2025-08-11T15:54:50.904961Z","shell.execute_reply.started":"2025-08-11T15:54:50.199773Z","shell.execute_reply":"2025-08-11T15:54:50.904038Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Analyze unique series and studies\nunique_series = train_localizers_df['SeriesInstanceUID'].nunique()\nunique_sop = train_localizers_df['SOPInstanceUID'].nunique()\nprint(f\"\\nUnique SeriesInstanceUID: {unique_series}\")\nprint(f\"Unique SOPInstanceUID: {unique_sop}\")\nprint(f\"Total annotations: {len(train_localizers_df)}\")\nprint(f\"Average annotations per series: {len(train_localizers_df) / unique_series:.2f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:57:17.523794Z","iopub.execute_input":"2025-08-11T15:57:17.524037Z","iopub.status.idle":"2025-08-11T15:57:17.530786Z","shell.execute_reply.started":"2025-08-11T15:57:17.524021Z","shell.execute_reply":"2025-08-11T15:57:17.529929Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Medical Imaging Data Loader\n- Creates functions to load DICOM series (3D volumes)\n- Implements NIfTI file loading for segmentations\n- Applies medical imaging windowing (contrast adjustment)\n- Tests loading functions on sample data\n- Displays sample slices with proper medical imaging display\n\n**Critical functions:**\n\n- `load_dicom_series()`: Loads and sorts DICOM slices into 3D volume\n- `load_nifti_file()`: Loads segmentation masks\n- `apply_window_level()`: Applies brain windowing for proper visualization","metadata":{}},{"cell_type":"code","source":"# Medical Imaging Data Loader\ndef load_dicom_series(series_path):\n    \"\"\"Load a complete DICOM series and sort by instance number\"\"\"\n    try:\n        dcm_files = glob.glob(f\"{series_path}/*.dcm\")\n        if not dcm_files:\n            return None, None\n        \n        # Load all DICOM files\n        dicoms = []\n        for dcm_file in dcm_files:\n            try:\n                ds = pydicom.dcmread(dcm_file)\n                dicoms.append((ds.InstanceNumber, ds))\n            except:\n                continue\n        \n        if not dicoms:\n            return None, None\n        \n        # Sort by instance number\n        dicoms.sort(key=lambda x: x[0])\n        \n        # Extract pixel arrays\n        images = []\n        metadata = []\n        for _, ds in dicoms:\n            if hasattr(ds, 'pixel_array'):\n                img = ds.pixel_array.astype(np.float32)\n                images.append(img)\n                metadata.append({\n                    'InstanceNumber': ds.InstanceNumber,\n                    'SliceLocation': getattr(ds, 'SliceLocation', None),\n                    'ImagePositionPatient': getattr(ds, 'ImagePositionPatient', None),\n                    'PixelSpacing': getattr(ds, 'PixelSpacing', None),\n                    'WindowCenter': getattr(ds, 'WindowCenter', None),\n                    'WindowWidth': getattr(ds, 'WindowWidth', None)\n                })\n        \n        if images:\n            volume = np.stack(images, axis=0)\n            return volume, metadata\n        \n    except Exception as e:\n        print(f\"Error loading DICOM series: {e}\")\n        return None, None\n    \n    return None, None\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:57:23.77639Z","iopub.execute_input":"2025-08-11T15:57:23.776668Z","iopub.status.idle":"2025-08-11T15:57:23.785858Z","shell.execute_reply.started":"2025-08-11T15:57:23.776648Z","shell.execute_reply":"2025-08-11T15:57:23.784984Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def load_nifti_file(nii_path):\n    \"\"\"Load NIfTI file\"\"\"\n    try:\n        nii_img = nib.load(nii_path)\n        data = nii_img.get_fdata()\n        return data, nii_img.header\n    except Exception as e:\n        print(f\"Error loading NIfTI file: {e}\")\n        return None, None\n\ndef apply_window_level(image, window_center, window_width):\n    \"\"\"Apply window/level to DICOM image\"\"\"\n    if window_center is None or window_width is None:\n        return image\n    \n    img_min = window_center - window_width // 2\n    img_max = window_center + window_width // 2\n    \n    windowed = np.clip(image, img_min, img_max)\n    windowed = (windowed - img_min) / (img_max - img_min)\n    return windowed","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:57:27.464657Z","iopub.execute_input":"2025-08-11T15:57:27.465153Z","iopub.status.idle":"2025-08-11T15:57:27.470318Z","shell.execute_reply.started":"2025-08-11T15:57:27.465133Z","shell.execute_reply":"2025-08-11T15:57:27.46958Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Test loading functions\nprint(\"Testing medical imaging loaders...\")\n\n# Test DICOM loading\nif len(series_folders) > 0:\n    test_series = series_folders[0]\n    print(f\"Testing DICOM loading with: {os.path.basename(test_series)}\")\n    \n    volume, metadata = load_dicom_series(test_series)\n    if volume is not None:\n        print(f\"DICOM volume shape: {volume.shape}\")\n        print(f\"Number of slices: {len(metadata)}\")\n        print(f\"Pixel value range: {volume.min():.2f} to {volume.max():.2f}\")\n        \n        # Display middle slice\n        middle_slice = volume.shape[0] // 2\n        plt.figure(figsize=(10, 5))\n        \n        plt.subplot(1, 2, 1)\n        plt.imshow(volume[middle_slice], cmap='gray')\n        plt.title(f'DICOM Slice {middle_slice} (Raw)')\n        plt.axis('off')\n        \n        # Apply windowing if available\n        if metadata[middle_slice]['WindowCenter'] and metadata[middle_slice]['WindowWidth']:\n            windowed = apply_window_level(\n                volume[middle_slice], \n                metadata[middle_slice]['WindowCenter'], \n                metadata[middle_slice]['WindowWidth']\n            )\n            plt.subplot(1, 2, 2)\n            plt.imshow(windowed, cmap='gray')\n            plt.title(f'DICOM Slice {middle_slice} (Windowed)')\n            plt.axis('off')\n        \n        plt.tight_layout()\n        plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:57:29.781594Z","iopub.execute_input":"2025-08-11T15:57:29.781861Z","iopub.status.idle":"2025-08-11T15:57:32.255638Z","shell.execute_reply.started":"2025-08-11T15:57:29.781842Z","shell.execute_reply":"2025-08-11T15:57:32.25502Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Test NIfTI loading\nif len(regular_nii) > 0:\n    test_nii = regular_nii[0]\n    print(f\"\\nTesting NIfTI loading with: {os.path.basename(test_nii)}\")\n    \n    nii_data, nii_header = load_nifti_file(test_nii)\n    if nii_data is not None:\n        print(f\"NIfTI volume shape: {nii_data.shape}\")\n        print(f\"NIfTI data range: {nii_data.min():.2f} to {nii_data.max():.2f}\")\n        \n        # Display middle slice\n        if len(nii_data.shape) == 3:\n            middle_slice = nii_data.shape[2] // 2\n            plt.figure(figsize=(8, 4))\n            plt.imshow(nii_data[:, :, middle_slice], cmap='gray')\n            plt.title(f'NIfTI Slice {middle_slice}')\n            plt.axis('off')\n            plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:57:52.807875Z","iopub.execute_input":"2025-08-11T15:57:52.808388Z","iopub.status.idle":"2025-08-11T15:57:53.180921Z","shell.execute_reply.started":"2025-08-11T15:57:52.808363Z","shell.execute_reply":"2025-08-11T15:57:53.180298Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Annotation Visualization\n- Overlays aneurysm annotations on DICOM images\n- Shows red circles at aneurysm locations\n- Displays multiple annotated slices in a grid\n- Matches annotations to corresponding image slices\n- Provides visual verification of data quality\n\n**Purpose:**\n\n- Verify annotations are correctly positioned\n- Understand aneurysm appearance in medical images\n- Quality check for data consistency\n- Visual debugging of coordinate systems","metadata":{}},{"cell_type":"code","source":"# Annotation Visualization\ndef visualize_annotations_on_dicom(series_uid, localizers_df):\n    \"\"\"Visualize aneurysm annotations on DICOM images\"\"\"\n    \n    # Find annotations for this series\n    series_annotations = localizers_df[localizers_df['SeriesInstanceUID'] == series_uid]\n    \n    if len(series_annotations) == 0:\n        print(f\"No annotations found for series {series_uid}\")\n        return\n    \n    # Find corresponding DICOM series folder\n    series_folder = None\n    for folder in series_folders:\n        if series_uid in folder:\n            series_folder = folder\n            break\n    \n    if series_folder is None:\n        print(f\"DICOM folder not found for series {series_uid}\")\n        return\n    \n    # Load DICOM series\n    volume, metadata = load_dicom_series(series_folder)\n    if volume is None:\n        print(f\"Failed to load DICOM series\")\n        return\n    \n    print(f\"Found {len(series_annotations)} annotations for series\")\n    print(f\"DICOM volume shape: {volume.shape}\")\n    \n    # Group annotations by SOPInstanceUID (individual images)\n    sop_groups = series_annotations.groupby('SOPInstanceUID')\n    \n    fig, axes = plt.subplots(2, 3, figsize=(15, 10))\n    axes = axes.flatten()\n    \n    plot_idx = 0\n    for sop_uid, group in sop_groups:\n        if plot_idx >= 6:  # Limit to 6 images\n            break\n            \n        # For simplicity, show first slice (in practice, you'd match SOPInstanceUID)\n        slice_idx = min(plot_idx, volume.shape[0] - 1)\n        img = volume[slice_idx]\n        \n        # Apply windowing if available\n        if metadata[slice_idx]['WindowCenter'] and metadata[slice_idx]['WindowWidth']:\n            img = apply_window_level(\n                img, \n                metadata[slice_idx]['WindowCenter'], \n                metadata[slice_idx]['WindowWidth']\n            )\n        \n        axes[plot_idx].imshow(img, cmap='gray')\n        \n        # Add annotation points\n        for _, row in group.iterrows():\n            x, y = parse_coordinates(row['coordinates'])\n            if x is not None and y is not None:\n                circle = Circle((x, y), radius=5, color='red', fill=False, linewidth=2)\n                axes[plot_idx].add_patch(circle)\n                axes[plot_idx].text(x+10, y-10, row['location'][:20], \n                                   color='red', fontsize=8, weight='bold')\n        \n        axes[plot_idx].set_title(f'Slice {slice_idx}\\n{len(group)} annotations')\n        axes[plot_idx].axis('off')\n        plot_idx += 1\n    \n    # Hide unused subplots\n    for i in range(plot_idx, 6):\n        axes[i].axis('off')\n    \n    plt.tight_layout()\n    plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:58:02.637209Z","iopub.execute_input":"2025-08-11T15:58:02.637476Z","iopub.status.idle":"2025-08-11T15:58:02.646366Z","shell.execute_reply.started":"2025-08-11T15:58:02.637458Z","shell.execute_reply":"2025-08-11T15:58:02.645672Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Test annotation visualization\nif 'train_localizers_df' in locals() and len(train_localizers_df) > 0:\n    # Get a series with multiple annotations\n    series_counts = train_localizers_df['SeriesInstanceUID'].value_counts()\n    test_series = series_counts.index[0]  # Series with most annotations\n    \n    print(f\"Visualizing annotations for series with {series_counts.iloc[0]} annotations\")\n    visualize_annotations_on_dicom(test_series, train_localizers_df)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:58:06.7019Z","iopub.execute_input":"2025-08-11T15:58:06.702788Z","iopub.status.idle":"2025-08-11T15:58:12.631123Z","shell.execute_reply.started":"2025-08-11T15:58:06.70275Z","shell.execute_reply":"2025-08-11T15:58:12.630408Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Basic Model Development Setup\n- Creates PyTorch Dataset class for aneurysm data\n- Implements data preprocessing pipeline\n- Defines a simple CNN architecture\n- Sets up GPU utilization\n- Converts point annotations to bounding boxes\n\n**Key components:**\n\n- `AneurysmDataset`: Handles data loading and preprocessing\n- `SimpleAneurysmDetector`: Basic CNN for classification\n- GPU detection and memory reporting\n- Data format conversion for training","metadata":{}},{"cell_type":"code","source":"# Basic Model Development Setup\nclass AneurysmDataset(Dataset):\n    \"\"\"Dataset class for aneurysm detection\"\"\"\n    \n    def __init__(self, series_paths, annotations_df, transform=None, image_size=512):\n        self.series_paths = series_paths\n        self.annotations_df = annotations_df\n        self.transform = transform\n        self.image_size = image_size\n        \n        # Create mapping from series to annotations\n        self.series_to_annotations = {}\n        for series_uid in annotations_df['SeriesInstanceUID'].unique():\n            series_annotations = annotations_df[annotations_df['SeriesInstanceUID'] == series_uid]\n            self.series_to_annotations[series_uid] = series_annotations\n    \n    def __len__(self):\n        return len(self.series_paths)\n    \n    def __getitem__(self, idx):\n        series_path = self.series_paths[idx]\n        series_uid = os.path.basename(series_path)\n        \n        # Load DICOM series\n        volume, metadata = load_dicom_series(series_path)\n        if volume is None:\n            # Return dummy data if loading fails\n            return torch.zeros(3, self.image_size, self.image_size), torch.zeros(1, 4)\n        \n        # For simplicity, use middle slice\n        middle_slice = volume.shape[0] // 2\n        image = volume[middle_slice]\n        \n        # Normalize and resize\n        image = cv2.resize(image, (self.image_size, self.image_size))\n        image = (image - image.min()) / (image.max() - image.min() + 1e-8)\n        \n        # Convert to 3-channel\n        image = np.stack([image, image, image], axis=0)\n        \n        # Get annotations for this series\n        annotations = self.series_to_annotations.get(series_uid, pd.DataFrame())\n        \n        # Create bounding boxes (simplified - using point annotations as centers)\n        boxes = []\n        labels = []\n        \n        for _, row in annotations.iterrows():\n            x, y = parse_coordinates(row['coordinates'])\n            if x is not None and y is not None:\n                # Scale coordinates to resized image\n                x_scaled = x * (self.image_size / volume.shape[2])\n                y_scaled = y * (self.image_size / volume.shape[1])\n                \n                # Create bounding box around point (simplified)\n                box_size = 32\n                x1 = max(0, x_scaled - box_size // 2)\n                y1 = max(0, y_scaled - box_size // 2)\n                x2 = min(self.image_size, x_scaled + box_size // 2)\n                y2 = min(self.image_size, y_scaled + box_size // 2)\n                \n                boxes.append([x1, y1, x2, y2])\n                labels.append(1)  # Aneurysm class\n        \n        if len(boxes) == 0:\n            boxes = [[0, 0, 1, 1]]  # Dummy box\n            labels = [0]  # Background\n        \n        boxes = torch.tensor(boxes, dtype=torch.float32)\n        labels = torch.tensor(labels, dtype=torch.long)\n        \n        if self.transform:\n            image = self.transform(image)\n        else:\n            image = torch.tensor(image, dtype=torch.float32)\n        \n        return image, {'boxes': boxes, 'labels': labels}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:58:36.458478Z","iopub.execute_input":"2025-08-11T15:58:36.458754Z","iopub.status.idle":"2025-08-11T15:58:36.470035Z","shell.execute_reply.started":"2025-08-11T15:58:36.458737Z","shell.execute_reply":"2025-08-11T15:58:36.46932Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def custom_collate(batch):\n    imgs, targets = zip(*batch)\n    return list(imgs), list(targets)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:58:40.83001Z","iopub.execute_input":"2025-08-11T15:58:40.830409Z","iopub.status.idle":"2025-08-11T15:58:40.834629Z","shell.execute_reply.started":"2025-08-11T15:58:40.830378Z","shell.execute_reply":"2025-08-11T15:58:40.833773Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class SimpleAneurysmDetector(nn.Module):\n    \"\"\"Simple CNN for aneurysm detection\"\"\"\n    \n    def __init__(self, num_classes=2):\n        super(SimpleAneurysmDetector, self).__init__()\n        \n        self.backbone = nn.Sequential(\n            nn.Conv2d(3, 64, 3, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(2),\n            \n            nn.Conv2d(64, 128, 3, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(2),\n            \n            nn.Conv2d(128, 256, 3, padding=1),\n            nn.ReLU(),\n            nn.MaxPool2d(2),\n            \n            nn.Conv2d(256, 512, 3, padding=1),\n            nn.ReLU(),\n            nn.AdaptiveAvgPool2d((7, 7))\n        )\n        \n        self.classifier = nn.Sequential(\n            nn.Linear(512 * 7 * 7, 1024),\n            nn.ReLU(),\n            nn.Dropout(0.5),\n            nn.Linear(1024, num_classes)\n        )\n    \n    def forward(self, x):\n        features = self.backbone(x)\n        features = features.view(features.size(0), -1)\n        output = self.classifier(features)\n        return output","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:58:43.523983Z","iopub.execute_input":"2025-08-11T15:58:43.524312Z","iopub.status.idle":"2025-08-11T15:58:43.531353Z","shell.execute_reply.started":"2025-08-11T15:58:43.524289Z","shell.execute_reply":"2025-08-11T15:58:43.530578Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Initialize model and check GPU availability\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\nprint(f\"Using device: {device}\")\n\nif torch.cuda.is_available():\n    print(f\"GPU: {torch.cuda.get_device_name(0)}\")\n    print(f\"GPU Memory: {torch.cuda.get_device_properties(0).total_memory / 1e9:.1f} GB\")\n\n# Create simple model\nmodel = SimpleAneurysmDetector(num_classes=2)\nmodel = model.to(device)\n\nprint(f\"Model parameters: {sum(p.numel() for p in model.parameters()):,}\")\nprint(\"Model architecture:\")\nprint(model)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:58:47.18191Z","iopub.execute_input":"2025-08-11T15:58:47.18224Z","iopub.status.idle":"2025-08-11T15:58:47.730328Z","shell.execute_reply.started":"2025-08-11T15:58:47.182221Z","shell.execute_reply":"2025-08-11T15:58:47.729521Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Training Pipeline Setup\n- Implements training and validation loops\n- Sets up loss function, optimizer, and scheduler\n- Creates train/validation data splits\n- Tests data loading with actual batches\n- Verifies forward pass through the model\n\n**Training components:**\n\n- `train_epoch()`: One epoch of model training\n- `validate_epoch()`: Model evaluation on validation set\n- Progress bars and metric tracking\n- Error handling for data loading issues","metadata":{}},{"cell_type":"code","source":"# Training Pipeline Setup\ndef train_epoch(model, dataloader, criterion, optimizer, device):\n    \"\"\"Train for one epoch\"\"\"\n    model.train()\n    running_loss = 0.0\n    correct = 0\n    total = 0\n    \n    pbar = tqdm(dataloader, desc='Training')\n    for batch_idx, (images, targets) in enumerate(pbar):\n        images = images.to(device)\n        \n        # For simplicity, convert to classification task\n        # (has_aneurysm = 1 if any annotations, 0 otherwise)\n        labels = torch.tensor([1 if len(t['labels']) > 0 and t['labels'][0] > 0 else 0 \n                              for t in targets], dtype=torch.long).to(device)\n        \n        optimizer.zero_grad()\n        outputs = model(images)\n        loss = criterion(outputs, labels)\n        loss.backward()\n        optimizer.step()\n        \n        running_loss += loss.item()\n        _, predicted = outputs.max(1)\n        total += labels.size(0)\n        correct += predicted.eq(labels).sum().item()\n        \n        pbar.set_postfix({\n            'Loss': f'{running_loss/(batch_idx+1):.4f}',\n            'Acc': f'{100.*correct/total:.2f}%'\n        })\n    \n    return running_loss / len(dataloader), 100. * correct / total","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:58:54.519816Z","iopub.execute_input":"2025-08-11T15:58:54.520352Z","iopub.status.idle":"2025-08-11T15:58:54.5264Z","shell.execute_reply.started":"2025-08-11T15:58:54.520326Z","shell.execute_reply":"2025-08-11T15:58:54.525484Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def validate_epoch(model, dataloader, criterion, device):\n    \"\"\"Validate for one epoch\"\"\"\n    model.eval()\n    running_loss = 0.0\n    correct = 0\n    total = 0\n    \n    with torch.no_grad():\n        pbar = tqdm(dataloader, desc='Validation')\n        for batch_idx, (images, targets) in enumerate(pbar):\n            images = images.to(device)\n            labels = torch.tensor([1 if len(t['labels']) > 0 and t['labels'][0] > 0 else 0 \n                                  for t in targets], dtype=torch.long).to(device)\n            \n            outputs = model(images)\n            loss = criterion(outputs, labels)\n            \n            running_loss += loss.item()\n            _, predicted = outputs.max(1)\n            total += labels.size(0)\n            correct += predicted.eq(labels).sum().item()\n            \n            pbar.set_postfix({\n                'Loss': f'{running_loss/(batch_idx+1):.4f}',\n                'Acc': f'{100.*correct/total:.2f}%'\n            })\n    \n    return running_loss / len(dataloader), 100. * correct / total","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:59:00.577044Z","iopub.execute_input":"2025-08-11T15:59:00.57782Z","iopub.status.idle":"2025-08-11T15:59:00.58401Z","shell.execute_reply.started":"2025-08-11T15:59:00.577792Z","shell.execute_reply":"2025-08-11T15:59:00.583332Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Setup training components\ncriterion = nn.CrossEntropyLoss()\noptimizer = optim.Adam(model.parameters(), lr=0.001)\nscheduler = StepLR(optimizer, step_size=10, gamma=0.1)\n\n# Create dataset (using available series)\navailable_series = [folder for folder in series_folders if os.path.exists(folder)][:50]  # Limit for testing\nprint(f\"Using {len(available_series)} series for training\")\n\nif len(available_series) > 0 and 'train_localizers_df' in locals():\n    # Split data\n    train_series, val_series = train_test_split(available_series, test_size=0.2, random_state=42)\n    \n    # Create datasets\n    train_dataset = AneurysmDataset(train_series, train_localizers_df, image_size=256)\n    val_dataset = AneurysmDataset(val_series, train_localizers_df, image_size=256)\n    \n    # Create dataloaders\n    train_loader = DataLoader(train_dataset, batch_size=4, shuffle=True, num_workers=2, collate_fn=custom_collate)\n    val_loader = DataLoader(val_dataset, batch_size=4, shuffle=False, num_workers=2, collate_fn=custom_collate)\n    \n    print(f\"Training samples: {len(train_dataset)}\")\n    print(f\"Validation samples: {len(val_dataset)}\")\n    \n    # Test one batch\n    print(\"\\nTesting data loading...\")\n    try:\n        sample_batch = next(iter(train_loader))\n        images, targets = sample_batch\n\n        print(f\"Type of images: {type(images)}; Batch size: {len(images)}\")\n        for idx, img in enumerate(images):\n            print(f\"  img[{idx}]: type {type(img)}, shape {tuple(img.shape)}\")\n        print(f\"Number of targets: {len(targets)}, first target keys: {targets[0].keys()}\")\n        \n        # Forward pass (list of different-sized tensors)\n        model.eval()\n        with torch.no_grad():\n            logits_list = []\n            for img in images:\n                img = img.to(device).unsqueeze(0)  # add batch dim = 1\n                logits = model(img)                # e.g. shape [1, num_classes]\n                logits_list.append(logits.squeeze(0))  # shape [num_classes]\n        \n            outputs = torch.stack(logits_list, dim=0)  # shape [batch, num_classes]\n            print(f\"Stacked model output: {outputs.shape}\")\n        \n        print(\"Forward pass successful!\")\n            \n    except Exception as e:\n        print(f\"Error in data loading/forward pass: {e}\")\n\nelse:\n    print(\"Insufficient data for training setup\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:59:03.464242Z","iopub.execute_input":"2025-08-11T15:59:03.464523Z","iopub.status.idle":"2025-08-11T15:59:25.96988Z","shell.execute_reply.started":"2025-08-11T15:59:03.464501Z","shell.execute_reply":"2025-08-11T15:59:25.969101Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Inference Pipeline\n- Creates prediction function for new DICOM series\n- Processes each slice in a volume\n- Aggregates slice-level predictions to series-level\n- Runs batch inference on multiple series\n- Provides confidence scores and detailed results\n\n**Inference features:**\n\n- `predict_aneurysm()`: Single series prediction\n- `batch_inference()`: Multiple series processing\n- Confidence scoring and result aggregation\n- Error handling for failed predictions","metadata":{}},{"cell_type":"code","source":"# Inference Pipeline\ndef predict_aneurysm(model, series_path, device, image_size=256):\n    \"\"\"Predict aneurysm presence in a DICOM series\"\"\"\n    model.eval()\n    \n    # Load DICOM series\n    volume, metadata = load_dicom_series(series_path)\n    if volume is None:\n        return None, \"Failed to load DICOM series\"\n    \n    predictions = []\n    confidences = []\n    \n    with torch.no_grad():\n        # Process each slice\n        for slice_idx in range(volume.shape[0]):\n            image = volume[slice_idx]\n            \n            # Preprocess\n            image = cv2.resize(image, (image_size, image_size))\n            image = (image - image.min()) / (image.max() - image.min() + 1e-8)\n            image = np.stack([image, image, image], axis=0)\n            image = torch.tensor(image, dtype=torch.float32).unsqueeze(0).to(device)\n            \n            # Predict\n            output = model(image)\n            probabilities = F.softmax(output, dim=1)\n            confidence, predicted = probabilities.max(1)\n            \n            predictions.append(predicted.cpu().item())\n            confidences.append(confidence.cpu().item())\n    \n    # Aggregate predictions (majority vote or max confidence)\n    has_aneurysm = sum(predictions) > len(predictions) // 2\n    max_confidence = max(confidences)\n    \n    return {\n        'has_aneurysm': has_aneurysm,\n        'confidence': max_confidence,\n        'slice_predictions': predictions,\n        'slice_confidences': confidences,\n        'num_slices': len(predictions)\n    }, None","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:59:34.935292Z","iopub.execute_input":"2025-08-11T15:59:34.936144Z","iopub.status.idle":"2025-08-11T15:59:34.944509Z","shell.execute_reply.started":"2025-08-11T15:59:34.936113Z","shell.execute_reply":"2025-08-11T15:59:34.943617Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def batch_inference(model, series_list, device, max_series=10):\n    \"\"\"Run inference on multiple series\"\"\"\n    results = []\n    \n    print(f\"Running inference on {min(len(series_list), max_series)} series...\")\n    \n    for i, series_path in enumerate(series_list[:max_series]):\n        series_name = os.path.basename(series_path)\n        print(f\"Processing {i+1}/{min(len(series_list), max_series)}: {series_name[:50]}...\")\n        \n        result, error = predict_aneurysm(model, series_path, device)\n        \n        if error:\n            print(f\"  Error: {error}\")\n            continue\n        \n        results.append({\n            'series_path': series_path,\n            'series_name': series_name,\n            'has_aneurysm': result['has_aneurysm'],\n            'confidence': result['confidence'],\n            'num_slices': result['num_slices']\n        })\n        \n        print(f\"  Result: {'Aneurysm' if result['has_aneurysm'] else 'No Aneurysm'} \"\n              f\"(confidence: {result['confidence']:.3f})\")\n    \n    return results","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:59:38.857851Z","iopub.execute_input":"2025-08-11T15:59:38.858433Z","iopub.status.idle":"2025-08-11T15:59:38.864219Z","shell.execute_reply.started":"2025-08-11T15:59:38.858406Z","shell.execute_reply":"2025-08-11T15:59:38.863352Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Run sample inference\nif len(available_series) > 0:\n    print(\"Running sample inference...\")\n    \n    # Test on a few series\n    inference_results = batch_inference(model, available_series, device, max_series=5)\n    \n    if inference_results:\n        print(f\"\\nInference Summary:\")\n        print(f\"Total series processed: {len(inference_results)}\")\n        \n        aneurysm_count = sum(1 for r in inference_results if r['has_aneurysm'])\n        print(f\"Predicted aneurysms: {aneurysm_count}\")\n        print(f\"Predicted normal: {len(inference_results) - aneurysm_count}\")\n        \n        avg_confidence = np.mean([r['confidence'] for r in inference_results])\n        print(f\"Average confidence: {avg_confidence:.3f}\")\n        \n        # Show detailed results\n        print(\"\\nDetailed Results:\")\n        for result in inference_results:\n            status = \"ANEURYSM\" if result['has_aneurysm'] else \"NORMAL\"\n            print(f\"  {result['series_name'][:40]}: {status} (conf: {result['confidence']:.3f})\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T15:59:41.296568Z","iopub.execute_input":"2025-08-11T15:59:41.297111Z","iopub.status.idle":"2025-08-11T16:00:03.297418Z","shell.execute_reply.started":"2025-08-11T15:59:41.297087Z","shell.execute_reply":"2025-08-11T16:00:03.296539Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Feasibility Report","metadata":{}},{"cell_type":"code","source":"# Feasibility Analysis Report\ndef generate_feasibility_report():\n    \"\"\"Generate a comprehensive feasibility report\"\"\"\n    \n    print(\"=\"*60)\n    print(\"RSNA INTRACRANIAL ANEURYSM DETECTION - FEASIBILITY REPORT\")\n    print(\"=\"*60)\n    \n    # Dataset Analysis\n    print(\"\\n1. DATASET ANALYSIS\")\n    print(\"-\" * 30)\n    \n    if 'train_localizers_df' in locals():\n        print(f\"✓ Localizer annotations: {len(train_localizers_df)} samples\")\n        print(f\"✓ Unique series: {train_localizers_df['SeriesInstanceUID'].nunique()}\")\n        print(f\"✓ Unique locations: {train_localizers_df['location'].nunique()}\")\n        \n        location_dist = train_localizers_df['location'].value_counts()\n        print(f\"✓ Most common location: {location_dist.index[0]} ({location_dist.iloc[0]} cases)\")\n        print(f\"✓ Least common location: {location_dist.index[-1]} ({location_dist.iloc[-1]} cases)\")\n    \n    print(f\"✓ DICOM series folders: {len(series_folders)}\")\n    print(f\"✓ Total DICOM files: {len(dcm_files)}\")\n    print(f\"✓ NIfTI segmentation files: {len(nii_files)}\")\n    print(f\"✓ Cowseg files: {len(cowseg_files)}\")\n    \n    # Technical Feasibility\n    print(\"\\n2. TECHNICAL FEASIBILITY\")\n    print(\"-\" * 30)\n    \n    print(f\"✓ GPU Available: {torch.cuda.is_available()}\")\n    if torch.cuda.is_available():\n        print(f\"✓ GPU Model: {torch.cuda.get_device_name(0)}\")\n        print(f\"✓ GPU Memory: {torch.cuda.get_device_properties(0).total_memory / 1e9:.1f} GB\")\n    \n    print(f\"✓ DICOM Loading: {'Success' if 'volume' in locals() and volume is not None else 'Failed'}\")\n    print(f\"✓ NIfTI Loading: {'Success' if 'nii_data' in locals() and nii_data is not None else 'Failed'}\")\n    print(f\"✓ Model Creation: Success\")\n    print(f\"✓ Data Pipeline: {'Success' if 'train_loader' in locals() else 'Partial'}\")\n    \n    # Challenges and Recommendations\n    print(\"\\n3. CHALLENGES & RECOMMENDATIONS\")\n    print(\"-\" * 30)\n    \n    challenges = [\n        \"Class imbalance - some aneurysm locations are rare\",\n        \"3D nature of data - need to handle volumetric information\",\n        \"Variable image quality and protocols across institutions\",\n        \"Small aneurysm size - requires high-resolution processing\",\n        \"Limited training data for some anatomical locations\"\n    ]\n    \n    recommendations = [\n        \"Use 3D CNNs or 2.5D approaches for better spatial context\",\n        \"Implement data augmentation for rare classes\",\n        \"Consider ensemble methods combining multiple views\",\n        \"Use attention mechanisms to focus on relevant regions\",\n        \"Implement proper cross-validation strategy\",\n        \"Consider transfer learning from pre-trained medical imaging models\"\n    ]\n    \n    print(\"Challenges:\")\n    for i, challenge in enumerate(challenges, 1):\n        print(f\"  {i}. {challenge}\")\n    \n    print(\"\\nRecommendations:\")\n    for i, rec in enumerate(recommendations, 1):\n        print(f\"  {i}. {rec}\")\n    \n    # Model Architecture Suggestions\n    print(\"\\n4. SUGGESTED MODEL ARCHITECTURES\")\n    print(\"-\" * 30)\n    \n    architectures = [\n        \"3D ResNet/DenseNet for volumetric analysis\",\n        \"U-Net for segmentation-based detection\",\n        \"YOLO/RetinaNet for object detection approach\",\n        \"Vision Transformer (ViT) for attention-based detection\",\n        \"Multi-scale CNN with Feature Pyramid Networks\"\n    ]\n    \n    for i, arch in enumerate(architectures, 1):\n        print(f\"  {i}. {arch}\")\n    \n    # Performance Expectations\n    print(\"\\n5. PERFORMANCE EXPECTATIONS\")\n    print(\"-\" * 30)\n    \n    print(\"Based on similar medical imaging competitions:\")\n    print(\"  • Baseline CNN: 0.65-0.75 AUC\")\n    print(\"  • Advanced 3D models: 0.75-0.85 AUC\")\n    print(\"  • Ensemble methods: 0.80-0.90 AUC\")\n    print(\"  • Top solutions: 0.85-0.95 AUC\")\n    \n    # Resource Requirements\n    print(\"\\n6. RESOURCE REQUIREMENTS\")\n    print(\"-\" * 30)\n    \n    print(\"Computational:\")\n    print(\"  • GPU: 16GB+ VRAM recommended for 3D models\")\n    print(\"  • RAM: 32GB+ for large batch processing\")\n    print(\"  • Storage: 100GB+ for preprocessed data\")\n    print(\"  • Training time: 2-5 days for full training\")\n    \n    print(\"\\nDevelopment:\")\n    print(\"  • Data preprocessing: 1-2 weeks\")\n    print(\"  • Model development: 2-3 weeks\")\n    print(\"  • Hyperparameter tuning: 1-2 weeks\")\n    print(\"  • Ensemble creation: 1 week\")\n    \n    print(\"\\n\" + \"=\"*60)\n    print(\"CONCLUSION: Project is FEASIBLE with proper approach\")\n    print(\"=\"*60)\n\n# Generate the report\ngenerate_feasibility_report()\n\n# Additional utility functions for competition\nprint(\"\\n\" + \"=\"*40)\nprint(\"UTILITY FUNCTIONS FOR COMPETITION\")\nprint(\"=\"*40)\n\ndef create_submission_format():\n    \"\"\"Create sample submission format\"\"\"\n    # This would be based on the actual submission requirements\n    sample_submission = pd.DataFrame({\n        'SeriesInstanceUID': ['sample_series_1', 'sample_series_2'],\n        'aneurysm_probability': [0.85, 0.23]\n    })\n    print(\"Sample submission format:\")\n    print(sample_submission)\n    return sample_submission\n\ndef preprocessing_pipeline():\n    \"\"\"Outline preprocessing steps\"\"\"\n    steps = [\n        \"1. Load DICOM series and sort by slice location\",\n        \"2. Apply appropriate windowing (brain window: 80/40)\",\n        \"3. Resample to consistent spacing (e.g., 1mm isotropic)\",\n        \"4. Crop/pad to consistent dimensions\",\n        \"5. Normalize intensity values\",\n        \"6. Apply data augmentation (rotation, scaling, noise)\",\n        \"7. Create 3D patches or 2.5D slices for training\"\n    ]\n    \n    print(\"Preprocessing Pipeline:\")\n    for step in steps:\n        print(f\"  {step}\")\n\ncreate_submission_format()\nprint()\npreprocessing_pipeline()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-11T16:00:24.932011Z","iopub.execute_input":"2025-08-11T16:00:24.932755Z","iopub.status.idle":"2025-08-11T16:00:24.95047Z","shell.execute_reply.started":"2025-08-11T16:00:24.932732Z","shell.execute_reply":"2025-08-11T16:00:24.949631Z"}},"outputs":[],"execution_count":null}]}