{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":99552,"databundleVersionId":13851420,"sourceType":"competition"}],"dockerImageVersionId":31089,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"border-left: 4px solid #4e4e4e; padding: 1.5rem; background-color: #1e1e1e; font-family: 'Segoe UI', sans-serif; color: #e0e0e0; line-height: 1.6; font-size: 16px; border-radius: 8px;\">\n\n\n\n  \n  <h3 style=\"color: #ffa07a;\">🩸 What is an Intracranial Aneurysm?</h3>\n  <p style=\"font-size: 0.9rem; color: #888; margin-top: 0.5rem;\">Credit goes to <strong>AC</strong> – <a href=\"https://www.kaggle.com/code/ahsuna123/voxel-by-voxel-3d-cnn-intracranial-aneurysms\" target=\"_blank\" style=\"color: #ffa07a; text-decoration: none;\">Kaggle Notebook</a></p>\n  <div style=\"display: flex; justify-content: center; margin: 2rem 0;\">\n  <img src=\"https://upload.wikimedia.org/wikipedia/commons/8/80/Cerebral_aneurysm_NIH.jpg\" \n       alt=\"Brain Aneurysm Scan\" \n       style=\"width: 25%; border-radius: 10px;\" />\n</div>\n  <p>\n    An intracranial aneurysm is a bulging or ballooning of a blood vessel in the brain. If undetected, it can rupture — leading to life-threatening bleeding. Early detection is crucial. The challenge? These aneurysms are often subtle, hidden in complex anatomy, and differ widely between patients.\n\n\n\n  </p>\n  <p>\n    That’s where machine learning can step in — augmenting radiologist workflows by detecting patterns across thousands of image slices in a fraction of the time.\n  </p>\n\n  <h3 style=\"color: #ffa07a;\">📷 Imaging Modalities Used</h3>\n  <p>We’re working with 3D scans from the following modalities:</p>\n\n  <table style=\"width: 100%; border-collapse: collapse; margin-top: 1rem; font-size: 15px;\">\n    <thead>\n      <tr style=\"background-color: #2e2e2e;\">\n        <th style=\"text-align: left; padding: 8px;\">Modality</th>\n        <th style=\"text-align: left; padding: 8px;\">Full Form</th>\n        <th style=\"text-align: left; padding: 8px;\">What It Shows</th>\n      </tr>\n    </thead>\n    <tbody>\n      <tr>\n        <td style=\"padding: 8px;\">CTA</td>\n        <td style=\"padding: 8px;\">Computed Tomography Angiography</td>\n        <td style=\"padding: 8px;\">Blood vessels using contrast-enhanced CT</td>\n      </tr>\n      <tr style=\"background-color: #2b2b2b;\">\n        <td style=\"padding: 8px;\">MRA</td>\n        <td style=\"padding: 8px;\">Magnetic Resonance Angiography</td>\n        <td style=\"padding: 8px;\">Blood vessels using MRI</td>\n      </tr>\n      <tr>\n        <td style=\"padding: 8px;\">MRI</td>\n        <td style=\"padding: 8px;\">Magnetic Resonance Imaging</td>\n        <td style=\"padding: 8px;\">Structural brain tissue, no contrast</td>\n      </tr>\n    </tbody>\n  </table>\n\n  <img src=\"https://www.nature.com/articles/s41598-023-33182-1/figures/1\" alt=\"3D Scan Slices\" style=\"width: 100%; border-radius: 10px; margin: 1rem 0;\" />\n\n  <p>\n    Each scan is a 3D volume made up of 2D slices. Think of it as flipping through pages of a brain atlas.\n  </p>\n\n  <h3 style=\"color: #ffa07a;\">📌 What Are We Predicting?</h3>\n  <p>Two key tasks define this competition:</p>\n  <ul>\n    <li><strong>Classification</strong> — Does this scan contain an aneurysm?</li>\n    <li><strong>Localization</strong> — If yes, where is it located? (3D coordinates + 1 of 13 brain regions)</li>\n  </ul>\n\n  <h3 style=\"color: #ffa07a;\">🗃️ Dataset Overview</h3>\n  <ul>\n    <li><code>train.csv</code> — Aneurysm presence and location labels</li>\n    <li><code>train_localizers.csv</code> — 3D coordinates for aneurysms (subset)</li>\n    <li><code>DICOM folders</code> — Raw scan data (one folder per patient)</li>\n    <li><code>segmentations</code> — Optional vessel masks in <code>.nii.gz</code> format</li>\n  </ul>\n\n  <h3 style=\"color: #ffa07a;\">🌍 Why This Matters</h3>\n  <p>\n    In real-world hospitals, radiologists spend hours scanning hundreds of slices to find aneurysms. Our model has the potential to become a powerful clinical assistant flagging risk areas, reducing diagnostic time, and minimizing human error.\n  </p>\n  <p>\n    The stakes are high. So is the opportunity.\n  </p>\n</div>\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\ntrain_df = pd.read_csv('/kaggle/input/rsna-intracranial-aneurysm-detection/train.csv')\ntrain_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:13:41.357277Z","iopub.execute_input":"2025-10-17T02:13:41.357768Z","iopub.status.idle":"2025-10-17T02:13:41.428244Z","shell.execute_reply.started":"2025-10-17T02:13:41.357738Z","shell.execute_reply":"2025-10-17T02:13:41.426987Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">Aneurysm Counts per Location</h2>\n<H4>Where do aneurysms tend to occur?\nThis bar chart answers that — using the 13 location_* columns.</H4>","metadata":{}},{"cell_type":"code","source":"import plotly.express as px\n\nlocation_cols = [col for col in train_df.columns if col not in ['SeriesInstanceUID', 'PatientAge', 'PatientSex', 'Modality', 'Aneurysm Present']]\nlocation_counts = train_df[location_cols].sum().sort_values(ascending=False)\n\nfig = px.bar(\n    location_counts,\n    orientation='v',\n    title='📊 Aneurysm Count by Location',\n    labels={'value': 'Count', 'index': 'Location'},\n    color=location_counts.values,\n    color_continuous_scale=[[0, '#00BFC4'], [1, '#C77CFF']],\n    template='plotly_dark',\n    text_auto=True\n)\nfig.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:13:44.539353Z","iopub.execute_input":"2025-10-17T02:13:44.539739Z","iopub.status.idle":"2025-10-17T02:13:47.145267Z","shell.execute_reply.started":"2025-10-17T02:13:44.539714Z","shell.execute_reply":"2025-10-17T02:13:47.143884Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">Imaging Modality Distribution</h2>","metadata":{}},{"cell_type":"code","source":"modality_counts = train_df['Modality'].value_counts()\n\nfig = px.pie(\n    names=modality_counts.index,\n    values=modality_counts.values,\n    title='🧠 Imaging Modality Distribution',\n    hole=0.4,\n    color_discrete_sequence=['#00BFC4', '#C77CFF'],\n    template='plotly_dark'\n)\nfig.update_traces(texttemplate='%{value} (%{percent})', textposition='inside')\n\nfig.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:13:52.116964Z","iopub.execute_input":"2025-10-17T02:13:52.117481Z","iopub.status.idle":"2025-10-17T02:13:52.209283Z","shell.execute_reply.started":"2025-10-17T02:13:52.117439Z","shell.execute_reply":"2025-10-17T02:13:52.207979Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">Age & Sex Demographics</h2>","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\nimport warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)\n\nplt.style.use('dark_background')\nplt.figure(figsize=(10,6))\nsns.histplot(\n    data=train_df,\n    x='PatientAge',\n    hue='Aneurysm Present',\n    bins=30,\n    kde=True,\n    palette={0: '#00BFC4', 1: '#C77CFF'}\n)\nplt.title(\"Age Distribution by Aneurysm Presence\", fontsize=14)\nplt.xlabel(\"Age\")\nplt.ylabel(\"Frequency\")\nplt.grid(alpha=0.2)\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:13:57.434761Z","iopub.execute_input":"2025-10-17T02:13:57.435159Z","iopub.status.idle":"2025-10-17T02:13:59.417806Z","shell.execute_reply.started":"2025-10-17T02:13:57.435133Z","shell.execute_reply":"2025-10-17T02:13:59.416596Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pd.crosstab(train_df['PatientSex'], train_df['Aneurysm Present'], normalize='index') * 100\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:14:03.193579Z","iopub.execute_input":"2025-10-17T02:14:03.194142Z","iopub.status.idle":"2025-10-17T02:14:03.231711Z","shell.execute_reply.started":"2025-10-17T02:14:03.194113Z","shell.execute_reply":"2025-10-17T02:14:03.230419Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">Class Imbalance: Do Most Patients Have Aneurysms?</h2>","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(\n    train_df,\n    x='Aneurysm Present',\n    title='⚖️ Class Imbalance: Any Aneurysm Present',\n    color='Aneurysm Present',\n    text_auto=True,\n    color_discrete_map={0: '#00BFC4', 1: '#C77CFF'},\n    template='plotly_dark'\n)\nfig.update_xaxes(type='category', tickvals=[0, 1], ticktext=[\"No Aneurysm\", \"Aneurysm\"])\nfig.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:14:06.866602Z","iopub.execute_input":"2025-10-17T02:14:06.866954Z","iopub.status.idle":"2025-10-17T02:14:07.014723Z","shell.execute_reply.started":"2025-10-17T02:14:06.866929Z","shell.execute_reply":"2025-10-17T02:14:07.013676Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"aneurysm_counts = train_df['Aneurysm Present'].value_counts()\nprint(aneurysm_counts)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:14:11.898906Z","iopub.execute_input":"2025-10-17T02:14:11.89928Z","iopub.status.idle":"2025-10-17T02:14:11.908503Z","shell.execute_reply.started":"2025-10-17T02:14:11.899252Z","shell.execute_reply":"2025-10-17T02:14:11.907373Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"total_cases = len(train_df)\nprint(f\"Total cases: {total_cases}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:14:14.240545Z","iopub.execute_input":"2025-10-17T02:14:14.241503Z","iopub.status.idle":"2025-10-17T02:14:14.247287Z","shell.execute_reply.started":"2025-10-17T02:14:14.241473Z","shell.execute_reply":"2025-10-17T02:14:14.246207Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Simple % table (can be styled in notebook output)\ngender_ct = pd.crosstab(train_df['PatientSex'], train_df['Aneurysm Present'], normalize='index') * 100\ngender_ct = gender_ct.rename(columns={0: 'No Aneurysm (%)', 1: 'Aneurysm (%)'})\ngender_ct.style.background_gradient(cmap='crest').format(\"{:.1f}%\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:14:36.19062Z","iopub.execute_input":"2025-10-17T02:14:36.190984Z","iopub.status.idle":"2025-10-17T02:14:36.334659Z","shell.execute_reply.started":"2025-10-17T02:14:36.190959Z","shell.execute_reply":"2025-10-17T02:14:36.333608Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">Train Localizer Analysis </h2>","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport ast\n\nlocalizers_df = pd.read_csv('/kaggle/input/rsna-intracranial-aneurysm-detection/train_localizers.csv')\n\n# Convert coordinate strings to dicts\nlocalizers_df['coords'] = localizers_df['coordinates'].apply(ast.literal_eval)\nlocalizers_df['x'] = localizers_df['coords'].apply(lambda d: d['x'])\nlocalizers_df['y'] = localizers_df['coords'].apply(lambda d: d['y'])\n\nlocalizers_df.drop(columns=['coordinates', 'coords'], inplace=True)\nlocalizers_df.head()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:14:39.542473Z","iopub.execute_input":"2025-10-17T02:14:39.542946Z","iopub.status.idle":"2025-10-17T02:14:39.658828Z","shell.execute_reply.started":"2025-10-17T02:14:39.542919Z","shell.execute_reply":"2025-10-17T02:14:39.657284Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\"> Heatmap of Aneurysm Locations in Image Space</h2>","metadata":{}},{"cell_type":"code","source":"fig = px.density_heatmap(\n    localizers_df,\n    x='x',\n    y='y',\n    nbinsx=50,\n    nbinsy=50,\n    title='🧠 Heatmap of Aneurysm Locations in Image Space',\n    color_continuous_scale='Turbo',\n    template='plotly_dark',\n)\nfig.update_yaxes(autorange=\"reversed\")\nfig.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:14:45.898634Z","iopub.execute_input":"2025-10-17T02:14:45.899063Z","iopub.status.idle":"2025-10-17T02:14:45.994044Z","shell.execute_reply.started":"2025-10-17T02:14:45.899029Z","shell.execute_reply":"2025-10-17T02:14:45.992926Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\"> 2D Scatter of Aneurysm Coordinates by Location</h2>","metadata":{}},{"cell_type":"code","source":"fig = px.scatter(\n    localizers_df,\n    x='x',\n    y='y',\n    color='location',\n    title='🧠 2D Scatter of Aneurysm Coordinates by Location',\n    template='plotly_dark',\n    color_discrete_sequence=px.colors.qualitative.Dark24\n)\nfig.update_yaxes(autorange=\"reversed\")\nfig.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:14:50.10082Z","iopub.execute_input":"2025-10-17T02:14:50.101159Z","iopub.status.idle":"2025-10-17T02:14:50.208025Z","shell.execute_reply.started":"2025-10-17T02:14:50.101135Z","shell.execute_reply":"2025-10-17T02:14:50.206867Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">🔁 Aneurysm Co-occurrence Across Brain Regions</h2>\n\n<p style=\"color:#ccc\">\nEach patient in the dataset can have multiple aneurysms in different brain arteries. But are certain aneurysm sites more likely to occur together?\n\nTo investigate this, we treat the location columns as binary labels and compute a co-occurrence matrix — it tells us how often, for instance, an aneurysm in the <code>Left MCA</code> is also accompanied by one in the <code>ACoA</code>, and so on.\n</p>\n","metadata":{}},{"cell_type":"code","source":"df = train_df\n# Identify location columns correctly\nlocation_cols = df.columns[4:-1]  # skip UID, Age, Sex, Modality, and skip final label\nlocation_df = df[location_cols].astype(int)  # just in case they're still object type\n\n# Co-occurrence matrix\nco_matrix = location_df.T.dot(location_df)\n\n# Plot heatmap\nplt.figure(figsize=(12, 10))\nsns.heatmap(co_matrix, cmap=\"magma\", annot=True, fmt=\".0f\", linewidths=0.5)\nplt.title(\"Aneurysm Co-occurrence Matrix\", fontsize=16, color='white')\nplt.xticks(rotation=45, ha='right', fontsize=9)\nplt.yticks(rotation=0, fontsize=9)\nplt.gca().set_facecolor('black')\nplt.gcf().set_facecolor('#111111')\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:14:58.040777Z","iopub.execute_input":"2025-10-17T02:14:58.041119Z","iopub.status.idle":"2025-10-17T02:14:58.901067Z","shell.execute_reply.started":"2025-10-17T02:14:58.041095Z","shell.execute_reply":"2025-10-17T02:14:58.899579Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\n# Assuming 'location_df' is already created and contains the location columns\n# location_df = df[location_cols].astype(int)\n\n# 1. Location Count Distribution (number of 1s per row)\nlocation_counts = location_df.sum(axis=1)\n\n# --- Plotting Section ---\nplt.figure(figsize=(8, 5))\nsns.histplot(location_counts, bins=range(1, location_counts.max()+2), kde=False, color=\"teal\")\nplt.title(\"Distribution of Aneurysm Locations per Patient\")\nplt.xlabel(\"Number of Locations with Aneurysm\")\nplt.ylabel(\"Number of Patients\")\nplt.grid(True)\nplt.tight_layout()\nplt.show()\n\n# --- Printing the Values Section ---\n# Use value_counts() to get the count for each number of locations and sort by the number of locations\nvalue_distribution = location_counts.value_counts().sort_index()\n\nprint(\"Distribution of Aneurysm Locations per Patient:\")\nprint(value_distribution)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:15:04.658953Z","iopub.execute_input":"2025-10-17T02:15:04.65932Z","iopub.status.idle":"2025-10-17T02:15:04.946796Z","shell.execute_reply.started":"2025-10-17T02:15:04.659295Z","shell.execute_reply":"2025-10-17T02:15:04.945503Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Assuming 'df' is already loaded and preprocessed (i.e., binary columns are int)\n\n# Filter only rows where aneurysm is present\ndf_pos = df[df[\"Aneurysm Present\"] == 1]\n\n# List of all aneurysm location columns\nlocation_cols = df_pos.columns[4:-1]  # From 'Left Infraclinoid...' to 'Other Posterior Circulation'\n\n# Group by PatientSex and sum each location\nsex_location = df_pos.groupby(\"PatientSex\")[location_cols].sum().T\n\n# Reset index for plotting\nsex_location = sex_location.reset_index().melt(id_vars=\"index\", var_name=\"Sex\", value_name=\"Count\")\nsex_location = sex_location.rename(columns={\"index\": \"Location\"})\n\n# Plot\nplt.figure(figsize=(16, 6))\nsns.barplot(data=sex_location, x=\"Location\", y=\"Count\", hue=\"Sex\")\nplt.xticks(rotation=45, ha=\"right\")\nplt.title(\"Aneurysm Location Frequency by Sex\")\nplt.ylabel(\"Aneurysm Count\")\nplt.xlabel(\"Location\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:15:09.481851Z","iopub.execute_input":"2025-10-17T02:15:09.48314Z","iopub.status.idle":"2025-10-17T02:15:09.958684Z","shell.execute_reply.started":"2025-10-17T02:15:09.483107Z","shell.execute_reply":"2025-10-17T02:15:09.957673Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# List of location columns\nlocation_cols = df.columns[4:-1]\n\n# Filter only positive cases\ndf_pos = df[df[\"Aneurysm Present\"] == 1]\n\n# Prepare a DataFrame: for each location, collect patients with that location = 1\nage_location = []\n\nfor loc in location_cols:\n    subset = df_pos[df_pos[loc] == 1][[\"PatientAge\"]].copy()\n    subset[\"Location\"] = loc\n    age_location.append(subset)\n\n# Combine all into one DataFrame\nage_location_df = pd.concat(age_location)\n\n# Plot\nplt.figure(figsize=(16, 6))\nsns.boxplot(data=age_location_df, x=\"Location\", y=\"PatientAge\", palette=\"crest\")\nplt.xticks(rotation=45, ha=\"right\")\nplt.title(\"Patient Age Distribution per Aneurysm Location\")\nplt.ylabel(\"Age\")\nplt.xlabel(\"Location\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:15:16.639493Z","iopub.execute_input":"2025-10-17T02:15:16.639822Z","iopub.status.idle":"2025-10-17T02:15:17.102571Z","shell.execute_reply.started":"2025-10-17T02:15:16.639799Z","shell.execute_reply":"2025-10-17T02:15:17.101148Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(16, 6))\nsns.violinplot(data=age_location_df, x=\"Location\", y=\"PatientAge\", palette=\"crest\")\nplt.xticks(rotation=45, ha=\"right\")\nplt.title(\"Patient Age Distribution per Aneurysm Location (Violin Plot)\")\nplt.ylabel(\"Age\")\nplt.xlabel(\"Location\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:15:22.092704Z","iopub.execute_input":"2025-10-17T02:15:22.093911Z","iopub.status.idle":"2025-10-17T02:15:22.689103Z","shell.execute_reply.started":"2025-10-17T02:15:22.093874Z","shell.execute_reply":"2025-10-17T02:15:22.688306Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Filter rows with aneurysm present\ndf_pos = df[df[\"Aneurysm Present\"] == 1].copy()\n\n# Initialize a dictionary to store counts\nmodality_counts = {}\n\nfor loc in location_cols:\n    subset = df_pos[df_pos[loc] == 1]\n    modality_distribution = subset[\"Modality\"].value_counts()\n    modality_counts[loc] = modality_distribution\n\n# Convert to DataFrame and fill missing values with 0\nmodality_df = pd.DataFrame(modality_counts).T.fillna(0).astype(int)\nmodality_df = modality_df[[\"CTA\", \"MRA\"]] if \"CTA\" in modality_df.columns and \"MRA\" in modality_df.columns else modality_df\n\n# Plot heatmap\nplt.figure(figsize=(12, 6))\nsns.heatmap(modality_df.T, annot=True, cmap=\"crest\", fmt=\"d\")\nplt.title(\"Modality Preference per Aneurysm Location\")\nplt.xlabel(\"Location\")\nplt.ylabel(\"Modality\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:15:27.579114Z","iopub.execute_input":"2025-10-17T02:15:27.579456Z","iopub.status.idle":"2025-10-17T02:15:28.191059Z","shell.execute_reply.started":"2025-10-17T02:15:27.579433Z","shell.execute_reply":"2025-10-17T02:15:28.189737Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Filter rows with aneurysm present\ndf_pos = df[df[\"Aneurysm Present\"] == 1].copy()\n\n# Initialize a dictionary to store counts\nmodality_counts = {}\n\nfor loc in location_cols:\n    subset = df_pos[df_pos[loc] == 1]\n    modality_distribution = subset[\"Modality\"].value_counts()\n    modality_counts[loc] = modality_distribution\n\n# Convert to DataFrame and fill missing values with 0\nmodality_df = pd.DataFrame(modality_counts).T.fillna(0).astype(int)\n#modality_df = modality_df[[\"CTA\", \"MRA\"]] if \"CTA\" in modality_df.columns and \"MRA\" in modality_df.columns else modality_df\n\n# Plot heatmap\nplt.figure(figsize=(12, 6))\nsns.heatmap(modality_df.T, annot=True, cmap=\"crest\", fmt=\"d\")\nplt.title(\"Modality Preference per Aneurysm Location\")\nplt.xlabel(\"Location\")\nplt.ylabel(\"Modality\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:15:34.486097Z","iopub.execute_input":"2025-10-17T02:15:34.486438Z","iopub.status.idle":"2025-10-17T02:15:34.9835Z","shell.execute_reply.started":"2025-10-17T02:15:34.486411Z","shell.execute_reply":"2025-10-17T02:15:34.982264Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"modality_df_melted = modality_df.reset_index().melt(id_vars=\"index\", value_name=\"Count\", var_name=\"Modality\")\nmodality_df_melted.rename(columns={\"index\": \"Location\"}, inplace=True)\n\nplt.figure(figsize=(14, 6))\nsns.barplot(data=modality_df_melted, x=\"Location\", y=\"Count\", hue=\"Modality\", palette=\"crest\")\nplt.title(\"Modality Usage per Aneurysm Location\")\nplt.xticks(rotation=45, ha=\"right\")\nplt.xlabel(\"Location\")\nplt.ylabel(\"Patient Count\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:15:40.399325Z","iopub.execute_input":"2025-10-17T02:15:40.399651Z","iopub.status.idle":"2025-10-17T02:15:40.944723Z","shell.execute_reply.started":"2025-10-17T02:15:40.399628Z","shell.execute_reply":"2025-10-17T02:15:40.943501Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import ast\n\ntrain_localizations = pd.read_csv(\"/kaggle/input/rsna-intracranial-aneurysm-detection/train_localizers.csv\")\n# Parse coordinates\ntrain_localizations[\"coord_dict\"] = train_localizations[\"coordinates\"].apply(ast.literal_eval)\ntrain_localizations[\"x\"] = train_localizations[\"coord_dict\"].apply(lambda d: d[\"x\"])\ntrain_localizations[\"y\"] = train_localizations[\"coord_dict\"].apply(lambda d: d[\"y\"])\n\n# Top 4-5 most frequent locations\ntop_locations = train_localizations[\"location\"].value_counts().head(7).index\n\n# Plot heatmap per location\nfor loc in top_locations:\n    subset = train_localizations[train_localizations[\"location\"] == loc]\n    plt.figure(figsize=(6, 5))\n    sns.kdeplot(data=subset, x=\"x\", y=\"y\", fill=True, cmap=\"crest\")\n    plt.title(f\"Spatial Heatmap for {loc}\")\n    plt.xlabel(\"X Coordinate\")\n    plt.ylabel(\"Y Coordinate\")\n    plt.tight_layout()\n    plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:15:46.13981Z","iopub.execute_input":"2025-10-17T02:15:46.14017Z","iopub.status.idle":"2025-10-17T02:15:49.563163Z","shell.execute_reply.started":"2025-10-17T02:15:46.140147Z","shell.execute_reply":"2025-10-17T02:15:49.561344Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\n\ntrain = pd.read_csv(\"/kaggle/input/rsna-intracranial-aneurysm-detection/train.csv\")\nlabel_freq = train[location_cols].sum().sort_values(ascending=False).rename(\"train.csv\")\nlocation_cols = [col for col in train_df.columns if col not in ['SeriesInstanceUID', 'PatientAge', 'PatientSex', 'Modality', 'Aneurysm Present']]\nlabel_freq = train_df[location_cols].sum().sort_values(ascending=False).rename(\"train.csv\")\nlocalizer_freq = train_localizations[\"location\"].value_counts().rename(\"train_localizations.csv\")\nfreq_comparison = pd.concat([label_freq, localizer_freq], axis=1).fillna(0).astype(int)\nax = freq_comparison.plot(kind=\"bar\", figsize=(14, 6), color=['#00BFC4', '#C77CFF'])\nfor p in ax.patches:\n    ax.annotate(f'{p.get_height()}', \n                (p.get_x() + p.get_width() / 2., p.get_height()), \n                ha='center', \n                va='center', \n                xytext=(0, 5), \n                textcoords='offset points')\n\nplt.title(\"Frequency Comparison of Aneurysm Locations\")\nplt.ylabel(\"Count\")\nplt.xticks(rotation=45, ha=\"right\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:15:56.809093Z","iopub.execute_input":"2025-10-17T02:15:56.80942Z","iopub.status.idle":"2025-10-17T02:15:57.432387Z","shell.execute_reply.started":"2025-10-17T02:15:56.809394Z","shell.execute_reply":"2025-10-17T02:15:57.430878Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# IMAGES","metadata":{}},{"cell_type":"code","source":"import os\nimport pandas as pd\n\nseries_dir = '/kaggle/input/rsna-intracranial-aneurysm-detection/series'\nseries_folders = [f for f in os.listdir(series_dir) if os.path.isdir(os.path.join(series_dir, f))]\n\nseries_counts = []\nfor folder in series_folders:\n    dcm_files = os.listdir(os.path.join(series_dir, folder))\n    series_counts.append({'SeriesInstanceUID': folder, 'NumSlices': len(dcm_files)})\n\ndf_series = pd.DataFrame(series_counts)\ndf_series.describe()\n\n\n# ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:16:07.348934Z","iopub.execute_input":"2025-10-17T02:16:07.349295Z","iopub.status.idle":"2025-10-17T02:21:02.436434Z","shell.execute_reply.started":"2025-10-17T02:16:07.349271Z","shell.execute_reply":"2025-10-17T02:21:02.435432Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pydicom\n\nsample_path = os.path.join(series_dir, series_folders[0])\nsample_file = os.listdir(sample_path)[0]\ndcm = pydicom.dcmread(os.path.join(sample_path, sample_file))\n\nprint(f\"Orientation: {dcm.ImageOrientationPatient}\")\nprint(f\"Voxel spacing: {dcm.PixelSpacing}\")\nprint(f\"Slice Thickness: {dcm.SliceThickness}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:21:48.861287Z","iopub.execute_input":"2025-10-17T02:21:48.86168Z","iopub.status.idle":"2025-10-17T02:21:49.673881Z","shell.execute_reply.started":"2025-10-17T02:21:48.861653Z","shell.execute_reply":"2025-10-17T02:21:49.672862Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport matplotlib.widgets as widgets\nimport numpy as np\n\ndef load_series(series_path):\n    files = sorted(os.listdir(series_path), key=lambda x: pydicom.dcmread(os.path.join(series_path, x)).InstanceNumber)\n    images = [pydicom.dcmread(os.path.join(series_path, f)).pixel_array for f in files]\n    return np.stack(images)\n\nvolume = load_series(os.path.join(series_dir, series_folders[0]))\n\n# Scrollable plot\nfrom ipywidgets import interact\n@interact(slice=(0, volume.shape[0]-1))\ndef show_slice(slice=0):\n    plt.imshow(volume[slice], cmap='gray')\n    plt.axis('off')\n    plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:21:54.10682Z","iopub.execute_input":"2025-10-17T02:21:54.107172Z","iopub.status.idle":"2025-10-17T02:21:56.847476Z","shell.execute_reply.started":"2025-10-17T02:21:54.107149Z","shell.execute_reply":"2025-10-17T02:21:56.846471Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport pydicom\n\n# Count slices and shape per series\ndicom_dir = '/kaggle/input/rsna-intracranial-aneurysm-detection/series'\nseries_stats = []\n\nfor series_id in os.listdir(dicom_dir)[:50]:  # limit for speed\n    series_path = os.path.join(dicom_dir, series_id)\n    files = os.listdir(series_path)\n    num_slices = len(files)\n    sample_dcm = pydicom.dcmread(os.path.join(series_path, files[0]))\n    shape = (sample_dcm.Rows, sample_dcm.Columns)\n    series_stats.append({\"SeriesInstanceUID\": series_id, \"Slices\": num_slices, \"Shape\": shape})\n\npd.DataFrame(series_stats).value_counts(\"Shape\").plot(kind=\"barh\", title=\"Common Image Resolutions\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:23:39.508131Z","iopub.execute_input":"2025-10-17T02:23:39.508719Z","iopub.status.idle":"2025-10-17T02:23:41.123654Z","shell.execute_reply.started":"2025-10-17T02:23:39.50868Z","shell.execute_reply":"2025-10-17T02:23:41.122167Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"voxel_data = []\n\nfor series_id in os.listdir(dicom_dir)[:50]:\n    slices = []\n    for f in sorted(os.listdir(os.path.join(dicom_dir, series_id))):\n        path = os.path.join(dicom_dir, series_id, f)\n        dcm = pydicom.dcmread(path)\n        slices.append(dcm)\n\n    try:\n        spacing = slices[0].PixelSpacing\n        thickness = float(slices[0].SliceThickness)\n        voxel_data.append({\n            \"SeriesInstanceUID\": series_id,\n            \"PixelSpacingX\": spacing[0],\n            \"PixelSpacingY\": spacing[1],\n            \"SliceThickness\": thickness,\n            \"NumSlices\": len(slices)\n        })\n    except:\n        continue\n\nvoxel_df = pd.DataFrame(voxel_data)\nsns.histplot(voxel_df[\"SliceThickness\"], bins=20)\nplt.title(\"Slice Thickness Distribution\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:23:46.632312Z","iopub.execute_input":"2025-10-17T02:23:46.632672Z","iopub.status.idle":"2025-10-17T02:25:57.50997Z","shell.execute_reply.started":"2025-10-17T02:23:46.632647Z","shell.execute_reply":"2025-10-17T02:25:57.509021Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"seg_dir = \"/kaggle/input/rsna-intracranial-aneurysm-detection/segmentations\"\nseg_series = [f.replace(\".nii.gz\", \"\") for f in os.listdir(seg_dir)]\nprint(\"Segmented Series:\", len(seg_series))\n\n# % of series with segmentation\nseg_percent = len(seg_series) / len(os.listdir(dicom_dir)) * 100\nprint(f\"{seg_percent:.2f}% of series have vessel segmentation.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-17T02:26:02.246421Z","iopub.execute_input":"2025-10-17T02:26:02.246763Z","iopub.status.idle":"2025-10-17T02:26:02.380974Z","shell.execute_reply.started":"2025-10-17T02:26:02.246738Z","shell.execute_reply":"2025-10-17T02:26:02.379854Z"}},"outputs":[],"execution_count":null}]}