{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":99552,"databundleVersionId":13851420,"sourceType":"competition"}],"dockerImageVersionId":31090,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Imports\n\nLoads the libraries we need.","metadata":{}},{"cell_type":"code","source":"from glob import glob\nimport os, json, random\nimport numpy as np\nimport pandas as pd\nimport polars as pl\nimport pandas as pd\nimport shutil\nimport pydicom\nimport matplotlib.pyplot as plt\n\nfrom ipywidgets import interact, widgets, fixed\nfrom pathlib import Path\nimport os, random, numpy as np, pandas as pd, pydicom, cv2, gc\n\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\n\nimport kaggle_evaluation.rsna_inference_server","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-08T15:50:11.811903Z","iopub.execute_input":"2025-10-08T15:50:11.812084Z","iopub.status.idle":"2025-10-08T15:50:11.817043Z","shell.execute_reply.started":"2025-10-08T15:50:11.812066Z","shell.execute_reply":"2025-10-08T15:50:11.816334Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Dataset locations & reading the training table\n\nPoints to the RSNA competition folders on Kaggle and loads the main label table (`train.csv`). Each row in train.csv corresponds to one imaging series (i.e., one folder of DICOM slices) and includes the target (Aneurysm Present) plus some metadata.\n\n`train.csv` — the main table you’ll train with. One row per series (via `SeriesInstanceUID`). Includes:\n\n- Demographics: `PatientAge`, `PatientSex`.\n\n- `Modality (CTA/MRA/MRI...)`.\n\n- 14 labels:\n\n    - `Aneurysm Present` (binary, “is there any aneurysm in this series?”).\n\n    - 13 location columns (binary each), e.g. “Left Middle Cerebral Artery”, “Anterior Communicating Artery”, etc.\n\n**This is the primary file for training/evaluating a classifier.**\n\n- For a simple baseline, we start with the single target Aneurysm Present.\n\n- Later, we can switch to multi-label training to predict the 13 locations as well.\n\n\nFor a simple 3D CNN demo: use train.csv (series-level labels).","metadata":{}},{"cell_type":"code","source":"# Kaggle dataset root (mounted read-only in Kaggle notebooks)\nBASE = Path(\"/kaggle/input/rsna-intracranial-aneurysm-detection\")\n\n# Where each DICOM series lives: series/<SeriesInstanceUID>/*.dcm\nSERIES_DIR = BASE / \"series\" \n\n# CSV files provided by the competition\nTRAIN_CSV  = BASE / \"train.csv\"             # main labels per SeriesInstanceUID\n# LOCAL_CSV  = BASE / \"train_localizers.csv\"  # per-slice point annotations\n\n# Load the main training table\ntrain_df = pd.read_csv(TRAIN_CSV)\n\n# Quick peek (optional in class)\nprint(\"Rows in train.csv:\", len(train_df))\ntrain_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-08T15:50:24.280902Z","iopub.execute_input":"2025-10-08T15:50:24.281178Z","iopub.status.idle":"2025-10-08T15:50:24.329577Z","shell.execute_reply.started":"2025-10-08T15:50:24.281155Z","shell.execute_reply":"2025-10-08T15:50:24.328878Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- Each row = one imaging series, identified by a unique SeriesInstanceUID.\n\n- A series is a 3D stack of slices from one scan session and one modality (CTA, MRA, MRI T1post, MRI T2, etc.).","metadata":{}},{"cell_type":"markdown","source":"## 01. Load data: Build a tiny training subset (keep it fast and balanced)\n\nFrom the big CSV, pick a small sample (N=60), try to balance positives/negatives, then split into train/val. This keeps our demo model runs quick.","metadata":{}},{"cell_type":"code","source":"# Small fixed voxel size for memory & speed (D, H, W)\nDEPTH, HEIGHT, WIDTH = 48, 96, 96\nBATCH_SIZE = 4\nEPOCHS = 2\nN_SERIES = 60\nSEED = 1337\nrandom.seed(SEED); np.random.seed(SEED)\n\ndf = pd.read_csv(TRAIN_CSV)[[\"SeriesInstanceUID\",\"Aneurysm Present\"]].dropna()\ndf[\"Aneurysm Present\"] = df[\"Aneurysm Present\"].astype(int)\ndf = df[df[\"SeriesInstanceUID\"].apply(lambda s: (SERIES_DIR/s).exists())].reset_index(drop=True)\n\n# Try to balance classes within the small sample\npos = df[df[\"Aneurysm Present\"]==1].sample(min(N_SERIES//2, df[\"Aneurysm Present\"].sum()), \n                                           random_state=SEED)\nneg = df[df[\"Aneurysm Present\"]==0].sample(N_SERIES-len(pos), random_state=SEED)\ndf_small = pd.concat([pos,neg]).sample(frac=1, random_state=SEED).reset_index(drop=True)\n\n# Simple split \nval_n = max(8, int(0.2*len(df_small)))\ndf_val = df_small.iloc[:val_n].reset_index(drop=True)\ndf_tr  = df_small.iloc[val_n:].reset_index(drop=True)\n\nlen(df_tr), len(df_val)  # helps sanity-check: e.g., 48 train / 12 val","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-10-08T15:50:26.797696Z","iopub.execute_input":"2025-10-08T15:50:26.798289Z","iopub.status.idle":"2025-10-08T15:50:31.176752Z","shell.execute_reply.started":"2025-10-08T15:50:26.798266Z","shell.execute_reply":"2025-10-08T15:50:31.17616Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 02. Set up & tf.data input pipeline\n\n- Makes sure TensorFlow sees the GPU\n- Builds a tf.data pipeline that calls our Python loader, batches, and prefetches.","metadata":{}},{"cell_type":"code","source":"def _sort_key(d):\n    \"\"\"Sort slices. Prefer z-position; fall back to InstanceNumber.\"\"\"\n    z = None\n    if hasattr(d, \"ImagePositionPatient\") and len(d.ImagePositionPatient)>=3:\n        try:\n            z = float(d.ImagePositionPatient[2])\n        except:\n            pass\n    if z is not None:\n        return (0, z)\n    return (1, float(getattr(d, \"InstanceNumber\", 0)))\n\ndef load_series(series_id, target_shape=(DEPTH,HEIGHT,WIDTH), clip=(1,99)):\n    \"\"\"Read a series and turn it into a fixed-size, normalized 3D volume (D,H,W,1).\"\"\"\n    sdir = SERIES_DIR/series_id\n    files = [p for p in sdir.glob(\"*.dcm\")]\n    if not files:\n        return np.zeros((*target_shape,1), np.float32)\n\n    # Read headers, ensure pixels exist\n    headers = []\n    for fp in files:\n        try:\n            d = pydicom.dcmread(str(fp), force=True, stop_before_pixels=False)\n            if hasattr(d, \"PixelData\"):\n                headers.append(d)\n        except:\n            pass\n    if not headers:\n        return np.zeros((*target_shape,1), np.float32)\n\n    # Sort slices (z if available, otherwise instance number)\n    headers.sort(key=_sort_key)\n\n    # Extract pixels; handle multi-frame DICOMs too\n    slices = []\n    for d in headers:\n        try:\n            arr = d.pixel_array.astype(np.float32)\n            slope = float(getattr(d,\"RescaleSlope\",1.0))\n            inter = float(getattr(d,\"RescaleIntercept\",0.0))\n            arr = arr*slope + inter\n            if arr.ndim==2:\n                slices.append(arr)\n            elif arr.ndim==3:\n                # handle (F,H,W) or (H,W,F)\n                if arr.shape[0] < 8 and arr.shape[-1] > arr.shape[0]:\n                    arr = np.moveaxis(arr, -1, 0)\n                for k in range(arr.shape[0]):\n                    slices.append(arr[k])\n        except:\n            continue\n    if not slices:\n        return np.zeros((*target_shape,1), np.float32)\n\n    vol = np.stack(slices,0)  # (D0,H0,W0)\n\n    # Robust intensity clamp (1–99th percentile), then normalize to [0,1]\n    lo, hi = np.percentile(vol, clip)\n    vol = np.clip(vol, lo, hi)\n    D0,H0,W0 = vol.shape\n\n    # Resize each slice to target H×W\n    resized = np.zeros((D0, target_shape[1], target_shape[2]), np.float32)\n    for i in range(D0):\n        resized[i] = cv2.resize(vol[i], (target_shape[2], target_shape[1]), interpolation=cv2.INTER_LINEAR)\n\n    # Depth-uniform sampling to target D\n    idx = np.linspace(0, D0-1, target_shape[0]).astype(int)\n    vol = resized[idx]\n\n    # Normalize to [0,1]\n    mn, mx = vol.min(), vol.max()\n    if mx>mn:\n        vol = (vol-mn)/(mx-mn)\n\n    return vol[...,None].astype(np.float32)  # (D,H,W,1)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-08T16:01:45.152913Z","iopub.execute_input":"2025-10-08T16:01:45.153212Z","iopub.status.idle":"2025-10-08T16:01:45.164379Z","shell.execute_reply.started":"2025-10-08T16:01:45.153189Z","shell.execute_reply":"2025-10-08T16:01:45.163576Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Make GPU memory growth \"on demand\" (avoid grabbing all VRAM at once)\ngpus = tf.config.list_physical_devices('GPU')\nfor g in gpus:\n    try:\n        tf.config.experimental.set_memory_growth(g, True)\n    except:\n        pass\n# print(\"GPUs visible to TF:\", gpus)\n\n# Optional mixed precision for speed (safe because output Dense is float32)\nfrom tensorflow.keras import mixed_precision\nmixed_precision.set_global_policy(\"mixed_float16\")\n\ndef make_ds_cache(df_part, shuffle):\n    \"\"\"With memory caching: first epoch is slow, later epochs much faster\"\"\"\n    sids = df_part[\"SeriesInstanceUID\"].astype(str).values\n    ys   = df_part[\"Aneurysm Present\"].astype(np.float32).values\n\n    def py_load(sid_bytes):\n        sid = sid_bytes.decode()\n        x = load_series(sid)  # (D,H,W,1) float32 in [0,1]\n        return x.astype(np.float32)\n\n    def wrapper(sid, y):\n        vol = tf.numpy_function(py_load, [sid], tf.float32)\n        vol.set_shape((DEPTH,HEIGHT,WIDTH,1))\n        return vol, y\n\n    ds = tf.data.Dataset.from_tensor_slices((sids, ys))\n    if shuffle:\n        ds = ds.shuffle(min(len(df_part),256), seed=SEED, reshuffle_each_iteration=True)\n\n    # key: map -> cache -> batch -> prefetch\n    ds = ds.map(wrapper, num_parallel_calls=tf.data.AUTOTUNE, deterministic=False)\n    ds = ds.cache()   # store mapped volumes in memory; reuse instead of reloading\n    ds = ds.batch(BATCH_SIZE).prefetch(tf.data.AUTOTUNE)\n    return ds\n\n# Final train/val datasets\ntrain_ds_cache = make_ds_cache(df_tr, shuffle=True)\nval_ds_cache   = make_ds_cache(df_val, shuffle=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-08T16:01:48.741835Z","iopub.execute_input":"2025-10-08T16:01:48.742111Z","iopub.status.idle":"2025-10-08T16:01:48.783283Z","shell.execute_reply.started":"2025-10-08T16:01:48.742086Z","shell.execute_reply":"2025-10-08T16:01:48.782731Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 03 Simple 3D CNN","metadata":{}},{"cell_type":"code","source":"def bigger_3dcnn():\n    inp = keras.Input((DEPTH, HEIGHT, WIDTH, 1))\n    x = layers.Conv3D(16, 3, padding=\"same\", activation=\"relu\")(inp)\n    x = layers.MaxPool3D()(x)\n    x = layers.Conv3D(32, 3, padding=\"same\", activation=\"relu\")(x)\n    x = layers.MaxPool3D()(x)\n\n    x = layers.Flatten()(x)\n    x = layers.Dense(128, activation=\"relu\")(x)\n    x = layers.Dropout(0.4)(x)\n\n    out = layers.Dense(1, activation=\"sigmoid\", dtype=\"float32\")(x)\n\n    model = keras.Model(inp, out)\n    model.compile(\n        optimizer=keras.optimizers.Adam(1e-3),\n        loss=\"binary_crossentropy\",\n        metrics=[\"accuracy\"]\n    )\n    return model\n\nmodel = bigger_3dcnn()\n\nmodel.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-08T16:01:54.080782Z","iopub.execute_input":"2025-10-08T16:01:54.081391Z","iopub.status.idle":"2025-10-08T16:01:54.987885Z","shell.execute_reply.started":"2025-10-08T16:01:54.081365Z","shell.execute_reply":"2025-10-08T16:01:54.987163Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 04. Training","metadata":{}},{"cell_type":"code","source":"# Optional: touch the datasets once to fill the cache\nfor _ in train_ds_cache.take(1): pass\nfor _ in val_ds_cache.take(1):   pass\n\nhistory = model.fit(\n    train_ds_cache,\n    validation_data=val_ds_cache,\n    epochs=4,\n    verbose=1,\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-08T16:01:57.572286Z","iopub.execute_input":"2025-10-08T16:01:57.572615Z","iopub.status.idle":"2025-10-08T16:06:01.746326Z","shell.execute_reply.started":"2025-10-08T16:01:57.57259Z","shell.execute_reply":"2025-10-08T16:06:01.745607Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 05. Validation","metadata":{}},{"cell_type":"code","source":"val_metrics = model.evaluate(val_ds_cache , return_dict=True, verbose=0)\nprint(val_metrics)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-08T16:19:52.275527Z","iopub.execute_input":"2025-10-08T16:19:52.282412Z","iopub.status.idle":"2025-10-08T16:19:52.322142Z","shell.execute_reply.started":"2025-10-08T16:19:52.282379Z","shell.execute_reply":"2025-10-08T16:19:52.321561Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 06. Save the trained model\n\nAfter training, we save the model file so we can use it later for prediction without retraining every time.","metadata":{}},{"cell_type":"code","source":"# Save the trained model\nMODEL_PATH = \"/kaggle/working/tiny3dcnn_model.keras\"\nmodel.save(MODEL_PATH)\nprint(\"Model saved to:\", MODEL_PATH)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-08T16:19:54.976692Z","iopub.execute_input":"2025-10-08T16:19:54.97697Z","iopub.status.idle":"2025-10-08T16:19:57.459148Z","shell.execute_reply.started":"2025-10-08T16:19:54.976949Z","shell.execute_reply":"2025-10-08T16:19:57.458541Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 07. Load the Trained Model and Define Labels\n\nLoad the saved model back into memory for inference.\n  \nAlso define the column names that Kaggle expects in final submission — one for each artery location plus the “Aneurysm Present” column.","metadata":{}},{"cell_type":"code","source":"# Load the saved model only once (outside predict)\nMODEL_PATH = \"/kaggle/working/tiny3dcnn_model.keras\"\ninference_model = keras.models.load_model(MODEL_PATH, compile=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-08T16:19:59.74785Z","iopub.execute_input":"2025-10-08T16:19:59.748463Z","iopub.status.idle":"2025-10-08T16:20:00.949047Z","shell.execute_reply.started":"2025-10-08T16:19:59.748439Z","shell.execute_reply":"2025-10-08T16:20:00.948241Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Define label columns and constants (same as demo submission)\n\nID_COL = \"SeriesInstanceUID\"\n\nLABEL_COLS = [\n    \"Left Infraclinoid Internal Carotid Artery\",\n    \"Right Infraclinoid Internal Carotid Artery\",\n    \"Left Supraclinoid Internal Carotid Artery\",\n    \"Right Supraclinoid Internal Carotid Artery\",\n    \"Left Middle Cerebral Artery\",\n    \"Right Middle Cerebral Artery\",\n    \"Anterior Communicating Artery\",\n    \"Left Anterior Cerebral Artery\",\n    \"Right Anterior Cerebral Artery\",\n    \"Left Posterior Communicating Artery\",\n    \"Right Posterior Communicating Artery\",\n    \"Basilar Tip\",\n    \"Other Posterior Circulation\",\n    \"Aneurysm Present\",\n]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-08T16:20:02.581684Z","iopub.execute_input":"2025-10-08T16:20:02.581952Z","iopub.status.idle":"2025-10-08T16:20:02.586512Z","shell.execute_reply.started":"2025-10-08T16:20:02.581933Z","shell.execute_reply":"2025-10-08T16:20:02.585624Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 08. Define the Prediction Function\n\nThis function tells Kaggle how to make predictions using our model.  \n\nEvery time Kaggle sends a new brain scan folder, the function loads that folder, runs our model to predict “Aneurysm Present,” fills the other columns with 0.5, and returns the results in the correct table format.","metadata":{}},{"cell_type":"code","source":"# Load the saved model only once (outside predict)\n\ndef predict(series_path: str) -> pl.DataFrame | pd.DataFrame:\n    \"\"\"Make a prediction for one series folder.\"\"\"\n\n    # Get the series ID (the folder name)\n    series_id = os.path.basename(series_path)\n\n    # ---- load and preprocess the DICOM series ----\n    # Use the same loader we defined before\n    vol = load_series(series_id)          # (D,H,W,1)\n    vol = np.expand_dims(vol, 0)          # (1,D,H,W,1) -> batch dimension\n\n    # ---- model prediction ----\n    pred_prob = inference_model.predict(vol, verbose=0)[0][0]  # single float between 0-1\n\n    # ---- build final prediction table ----\n    # 13 other columns = 0.5, last one from model\n    row = [series_id] + [0.5] * 13 + [float(pred_prob)]\n    predictions = pl.DataFrame(\n        data=[row],\n        schema=[ID_COL, *LABEL_COLS],\n        orient=\"row\",\n    )\n\n    # ---- sanity checks (keep these) ----\n    if isinstance(predictions, pl.DataFrame):\n        assert predictions.columns == [ID_COL, *LABEL_COLS]\n    elif isinstance(predictions, pd.DataFrame):\n        assert (predictions.columns == [ID_COL, *LABEL_COLS]).all()\n    else:\n        raise TypeError(\"The predict function must return a DataFrame\")\n\n    # ---- cleanup to avoid disk overflow ----\n    shutil.rmtree(\"/kaggle/shared\", ignore_errors=True)\n\n    # Kaggle expects the ID column to be dropped before returning\n    return predictions.drop(ID_COL)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-08T16:20:05.200766Z","iopub.execute_input":"2025-10-08T16:20:05.201259Z","iopub.status.idle":"2025-10-08T16:20:05.207175Z","shell.execute_reply.started":"2025-10-08T16:20:05.201232Z","shell.execute_reply":"2025-10-08T16:20:05.206529Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 09. Connect to Kaggle Evaluation Server\n\n- This section sets up the Kaggle inference server.\n\n- It allows our notebook to talk to Kaggle’s hidden test data.\n\n- During submission, Kaggle sends one scan at a time to this server, and our function sends back predictions.\n\n- Locally, it also saves a small example file (submission.parquet) so we can preview the format before submitting.","metadata":{}},{"cell_type":"code","source":"inference_server = kaggle_evaluation.rsna_inference_server.RSNAInferenceServer(predict)\n\nif os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n    inference_server.serve()\nelse:\n    inference_server.run_local_gateway()\n    display(pl.read_parquet('/kaggle/working/submission.parquet'))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-08T16:20:10.541076Z","iopub.execute_input":"2025-10-08T16:20:10.541962Z","iopub.status.idle":"2025-10-08T16:20:15.944313Z","shell.execute_reply.started":"2025-10-08T16:20:10.541923Z","shell.execute_reply":"2025-10-08T16:20:15.943517Z"}},"outputs":[],"execution_count":null}]}