{
  "id": 611846,
  "title": "1st Place Solution",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/611846",
  "author_name": "tomoon33",
  "post_date": "2025-10-15T03:36:01.913000",
  "votes": 158,
  "comment_count": 36,
  "views": 0,
  "content": "<p>I thank RSNA, the organizers, the contributing radiologists and institutions, and the Kaggle team for hosting this impactful challenge. This write-up summarizes my 1st place solution. The core of my approach is a robust, coarse-to-fine pipeline that uses vessel segmentation to guide a region-of-interest (ROI) based classifier, producing location-aware predictions.</p>\n<h2>Solution Overview</h2>\n<ul>\n<li><p><strong>High-Level Pipeline</strong></p>\n<ol>\n<li><strong>Preprocessing:</strong> Convert and standardize DICOM series into NIfTI volumes.</li>\n<li><strong>Vessel Segmentation &amp; ROI Extraction:</strong> Use a coarse-to-fine nnU-Net approach. A fast, low-resolution model first finds a candidate region. This allows the more detailed, high-resolution models to focus only on this ROI, which improves both accuracy and speed.</li>\n<li><strong>ROI Classification:</strong> A 3D classification model, using the detailed vessel masks as input, predicts the probabilities for the 13 anatomical locations and the overall aneurysm presence.</li></ol></li>\n<li><p><strong>Key Design Principles</strong></p>\n<ul>\n<li><strong>Coarse-to-Fine Efficiency:</strong> A fast, low-resolution model first scans the entire volume to find a candidate region. This allows the more detailed, high-resolution models to focus only on this ROI. This improves both accuracy and speed.</li>\n<li><strong>Using Segmentation as a Structural Guide:</strong> By providing the classification model with an explicit vessel mask, I give it a detailed map of the vessel structure. This helps the model to focus on the vessels, where aneurysms are located.</li></ul></li>\n</ul>\n<h2>Data Preparation</h2>\n<ul>\n<li>I excluded approximately 60 series due to data quality issues, such as orientation anomalies, corrupted DICOM files, and implausible slice spacing.</li>\n<li>I used multilabel-stratified 5-fold cross-validation. </li>\n</ul>\n<h2>Pipeline</h2>\n<h3>1. Preprocessing</h3>\n<ul>\n<li><strong>Filter Slices:</strong> Within each series, I retained only the images matching the majority <code>Rows × Cols</code> and <code>PixelSpacing</code> configuration and filtered out slices with outlier interslice spacing to ensure consistency.</li>\n<li><strong>Convert DICOM to NIfTI:</strong> I used <code>dcm2niix</code> for conversion. If it failed, I first ran <code>gdcmconv --raw</code> as a fallback before retrying the conversion [1].</li>\n<li><strong>Standardize Orientation:</strong> All volumes were reoriented to a consistent anatomical orientation using nnU-Net’s <code>SimpleITKIOWithReorient</code>.</li>\n<li><strong>Normalize Intensity:</strong> I applied nnU-Net’s standard per-volume z-score normalization to standardize image intensities.</li>\n</ul>\n<h3>2. nnU-Net Segmentation + ROI Extraction (Coarse-to-Fine)</h3>\n<p>This stage uses a sequence of three nnU-Net v2 [2] models (<code>nnUNetResEncUNetMPlans</code>, <code>3d_fullres</code> configuration) to first locate a coarse ROI and then produce detailed vessel segmentations within it.</p>\n<ul>\n<li><p><strong>Model 1: Coarse Vessel Localization</strong></p>\n<ul>\n<li><strong>Spacing:</strong> (1.0, 1.0, 1.0) mm</li>\n<li><strong>Classes:</strong> 3 vessel groups (Posterior+Basilar / MCA / Other)</li>\n<li><strong>Loss:</strong> Dice + Cross-Entropy</li>\n<li><strong>Purpose:</strong> To perform a fast, low-resolution scan to efficiently find a single ROI candidate for the high-resolution models. This model also supports an optional orientation correction step described later.</li></ul></li>\n<li><p><strong>Model 2: Fine Segmentation (Balanced)</strong></p>\n<ul>\n<li><strong>Spacing:</strong> (0.80, 0.45, 0.44) mm</li>\n<li><strong>Loss:</strong> Dice + Cross-Entropy + SkeletonRecall (weight=1) [3]</li>\n<li><strong>Purpose:</strong> To generate a precise vessel segmentation. The SkeletonRecall loss improves the connectivity of thin vessels, which standard losses might miss.</li></ul></li>\n<li><p><strong>Model 3: Fine Segmentation (Recall-Focused)</strong></p>\n<ul>\n<li><strong>Spacing:</strong> (0.80, 0.45, 0.44) mm</li>\n<li><strong>Loss:</strong> Tversky + Cross-Entropy + SkeletonRecall (weight=3)</li>\n<li><strong>Purpose:</strong> To complement Model 2 by prioritizing recall, making it more sensitive to detecting hard-to-find vessel segments.</li></ul></li>\n<li><p><strong>Augmentation Strategy</strong></p>\n<ul>\n<li>I disabled left–right mirroring for the fine models (2 and 3) to preserve anatomical asymmetry.</li>\n<li>I used stronger intensity and geometric augmentations [4].</li>\n<li>I added low-resolution simulation transforms to make the models robust to scans with thick slices.</li></ul></li>\n<li><p><strong>Inference Process (Two-Stage)</strong></p>\n<ul>\n<li><strong>Stage 1 (Coarse Scan):</strong> I run Model 1 with a sliding window (overlap=0.2) and binarize the output to get a foreground mask. I then apply DBSCAN clustering to the mask to remove scattered false positives. The centroid of the largest cluster is used to crop a fixed-size ROI (140×140×140 mm).</li>\n<li><strong>Stage 2 (Fine Inference):</strong> I run Models 2 and 3 with a higher overlap (0.3) only within the coarse ROI. The vessel segmentation from Model 2 is used to compute a tight bounding box, which is then re-cropped with margins to create the final ROI for the classifier.</li></ul></li>\n</ul>\n<p>The detailed segmentations from this stage are passed to the classification model as masks.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F9edf9e0846e14151f33a1ece358473a9%2Fflow_segmentation.png?generation=1760833847347432&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F82fa14c88dfc1dc57448f40bca3ee07d%2Fseg_losses.png?generation=1760833856913829&amp;alt=media\" alt=\"\"></p>\n<h3>3. ROI Classification (13 locations + Aneurysm Present)</h3>\n<p>Once the ROI and vessel masks are prepared, a 3D classification model predicts the final probabilities.</p>\n<ul>\n<li><p><strong>Model Architecture</strong></p>\n<ul>\n<li><strong>Input Size:</strong> The model takes ROI volumes of size 128 × 256 × 256 voxels.</li>\n<li><strong>Backbone:</strong> The core of the model is an nnU-Net pre-trained for the vessel segmentation task. This approach was more accurate and faster to train than standard 2.5D or 3D timm backbones.<ul>\n<li><strong>Decoder Simplification:</strong> To improve efficiency, I simplified the decoder by removing its final, computationally-heavy block. This had no negative impact on performance.</li></ul></li>\n<li><strong>Auxiliary Detection Task:</strong> To help the model learn specific features of aneurysms, I added an auxiliary task that uses the decoder features to reconstruct a small binary sphere (5-pixel radius) at the location of each annotated aneurysm.</li>\n<li><strong>Location-Specific Prediction Head:</strong> To predict the 13 location-specific probabilities, I use the following process:<ol>\n<li><strong>Per-Location Feature Pooling:</strong> A \"Vessel Region-Masked Pooling\" layer uses the vessel masks to extract feature vectors corresponding to each of the 13 anatomical locations from the decoder's feature maps. (<em>Note: In practice, I apply this pooling using masks from both fine segmentation models and concatenate the results for more complete features.</em>)</li>\n<li><strong>Feature Fusion:</strong> These 13 feature vectors are combined with a global feature vector from the encoder (via Global Average Pooling).</li>\n<li><strong>Inter-Location Modeling:</strong> The combined features are fed into a \"Location-Aware Transformer\" to model relationships between different vessel locations.</li>\n<li><strong>Classification:</strong> Finally, an MLP head predicts the probability for each of the 13 locations.</li></ol></li>\n<li><strong>\"Aneurysm Present\" Prediction Head:</strong> For the overall presence prediction, I pool features over the entire vessel structure (a union of all vessel masks) and combine them with the encoder's global features. This aggregated feature vector is passed to a separate MLP head.</li>\n<li><strong>Output Design:</strong> I treated each of the 14 labels as an independent binary classification problem. This design helps with the severe class imbalance, as positive cases for any single location are very rare.</li></ul></li>\n<li><p><strong>Training Details</strong></p>\n<ul>\n<li><strong>Loss Functions:</strong><ul>\n<li><strong>13 Locations:</strong> <code>BCEWithLogitsLoss</code>.</li>\n<li><strong>Aneurysm Present:</strong> <code>BCEWithLogitsLoss</code>.</li>\n<li><strong>Auxiliary Sphere Segmentation:</strong> A combination of Balanced BCE [5] and Focal-Tversky++ loss [6, 7]. This combination works well for highly sparse targets and helps prevent the over-confidence that can occur with Dice-like losses.</li>\n<li><strong>Loss Weights:</strong> The final loss was a weighted sum of the three components. I set the weights to 0.1 for the 13 location losses, 0.05 for the Aneurysm Present loss, and 1.0 for the auxiliary sphere segmentation loss. The main goal was to prioritize learning the precise location of aneurysms, so the sphere segmentation task had the highest weight. I found that higher weights on the classification losses led to overfitting, so this balance was important.</li></ul></li>\n<li><strong>Data Augmentation</strong><ul>\n<li><strong>Intensity Transforms:</strong> Gaussian noise, Gaussian smoothing, intensity shift/scale, contrast adjustment, Gaussian sharpening, and intensity inversion.</li>\n<li><strong>Geometric Transforms:</strong> Random flips (z, y, x axes), small rotations (±10°), scaling/shearing (±10%), mild grid distortions, and a simulated low-resolution transform.</li></ul></li>\n<li><strong>Optimizer and Schedule:</strong> I used the AdamW optimizer with a learning rate of 1e-4 and an effective batch size of 8 (achieved with gradient accumulation). A standard cosine annealing schedule with a warmup period was used.</li>\n<li><strong>EMA Weights:</strong> I used the Exponential Moving Average (EMA) of the model weights for inference.</li></ul></li>\n<li><p><strong>Inference</strong></p>\n<ul>\n<li><strong>Ensembling:</strong> The final predictions are an average of the models from 4 of the 5 cross-validation folds.</li>\n<li><strong>Test-Time Augmentation (TTA):</strong> I averaged the predictions from the original volume and a left-right flipped version of it.</li></ul></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F98e18ca93230a3f26bf2181021ce481a%2Fmodel_overview.png?generation=1760490733871010&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F187c3d0327547a692d8e922dc5f1d5cd%2Fvessel_pooling.png?generation=1760490747217372&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F0dd935cbc51a7d50bc69779c8f12d918%2Ftransformer.png?generation=1760490761936793&amp;alt=media\"></p>\n<h3>4. Other Details</h3>\n<ul>\n<li><strong>Fail-safe Mechanism:</strong> If an anomaly occurred during the segmentation or ROI extraction steps, the pipeline would not attempt a prediction. Instead, it fell back to a set of pre-determined probabilities. For each class, this probability was the mean of the out-of-fold predictions from my cross-validation set.</li>\n<li><strong>Optional Orientation Correction:</strong> I implemented a method to fix misoriented scans. It analyzed the spatial arrangement of the three vessel groups from the coarse segmentation. By comparing the relative positions of these groups to their expected anatomical locations, it estimated the correct axis permutation. This worked perfectly on the training set, fixing all identified orientation issues. However, it had no measurable effect on the leaderboard score, likely because the test set did not contain such orientation errors.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2Ff954825e84986adf4e7b537f9160b1b0%2F3classes_for_ori_corr.png?generation=1760833867211206&amp;alt=media\"></p>\n<h2>Processing Time</h2>\n<ul>\n<li><strong>Training</strong><ul>\n<li>The training times below were measured on a single NVIDIA RTX 4090.</li>\n<li>The 96×192×192 input size was mainly used for faster experimentation and tuning.</li></ul></li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Input size</th>\n<th>Epochs</th>\n<th>Training Time</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>nnU-Net (Model 2)</td>\n<td>64×192×192</td>\n<td>1000</td>\n<td>14 h</td>\n</tr>\n<tr>\n<td>ROI classifier</td>\n<td>96×192×192</td>\n<td>30</td>\n<td>6 h</td>\n</tr>\n<tr>\n<td>ROI classifier</td>\n<td>128×256×256</td>\n<td>30</td>\n<td>12 h</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><strong>Inference</strong><ul>\n<li>Inference times were measured on the Kaggle notebook using two T4 GPUs.</li>\n<li>The times reported are an average over 100 samples.</li></ul></li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Step</th>\n<th>Time per Series</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Preprocessing</td>\n<td>4.03 ± 3.78 s</td>\n</tr>\n<tr>\n<td>Vessel segmentation</td>\n<td>10.65 ± 4.10 s</td>\n</tr>\n<tr>\n<td>ROI classification</td>\n<td>3.33 ± 0.09 s</td>\n</tr>\n</tbody>\n</table>\n<h2>Ablation Study</h2>\n<p>To validate the effectiveness of the key components in my ROI classification model, I conducted an ablation study. Note that this was a simplified evaluation; hyperparameters such as the number of epochs and learning rate were not re-tuned for each experiment. For efficiency, these experiments were run on folds 0, 1, and 2, using a reduced input size of 96×192×192.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Input Size</th>\n<th>AUC (Aneurysm Present)</th>\n<th>AUC (13 Locations)</th>\n<th>Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Final model (full resolution)</td>\n<td>128×256×256</td>\n<td>0.915</td>\n<td>0.916</td>\n<td>0.916</td>\n</tr>\n<tr>\n<td>Final model</td>\n<td>96×192×192</td>\n<td>0.907</td>\n<td>0.898</td>\n<td>0.902</td>\n</tr>\n<tr>\n<td>Without Location-Aware Transformer</td>\n<td>96×192×192</td>\n<td>0.899</td>\n<td>0.894</td>\n<td>0.896</td>\n</tr>\n<tr>\n<td>Using Dice instead of FocalTversky++</td>\n<td>96×192×192</td>\n<td>0.902</td>\n<td>0.896</td>\n<td>0.899</td>\n</tr>\n<tr>\n<td>Setting all loss weights to 1.0</td>\n<td>96×192×192</td>\n<td>0.890</td>\n<td>0.877</td>\n<td>0.884</td>\n</tr>\n<tr>\n<td>Without backbone pretraining</td>\n<td>96×192×192</td>\n<td>0.777</td>\n<td>0.811</td>\n<td>0.794</td>\n</tr>\n<tr>\n<td>Without segmentation model 3</td>\n<td>96×192×192</td>\n<td>0.899</td>\n<td>0.883</td>\n<td>0.891</td>\n</tr>\n<tr>\n<td>Without auxiliary segmentation loss</td>\n<td>96×192×192</td>\n<td>0.880</td>\n<td>0.871</td>\n<td>0.876</td>\n</tr>\n</tbody>\n</table>\n<p>The results show several key points:</p>\n<ul>\n<li>Pretraining the backbone on the vessel segmentation task was the most important factor. It greatly improved the score and helped the model train much faster. This was very helpful for running many experiments.</li>\n<li>The Location-Aware Transformer and the FocalTversky++ loss seemed to help at first, but their final contribution to the score was small. This is likely because other improvements and tuning had a larger overall effect.</li>\n</ul>\n<h2>Design Journey</h2>\n<p>My final design was the result of several iterations:</p>\n<ol>\n<li>I initially tried a single, end-to-end 3D classifier that had auxiliary heads for vessel segmentation and aneurysm localization. While it could detect the presence of an aneurysm, it failed to predict the 13 specific locations accurately.</li>\n<li>I then observed that a standard nnU-Net for vessel segmentation trained easily and generalized well across all modalities. This led me to change to a two-stage, vessel-first pipeline, where the segmentation acts as a strong guide for the subsequent classification task.</li>\n<li>I also experimented with a simpler model that took a 2-channel input: the image volume concatenated with a single binary vessel mask. To get a prediction for a specific label, I would feed the model the corresponding mask (e.g., the mask for one location, or the union mask for \"Aneurysm Present\"). Although the model itself was simple, this approach required running 14 separate forward passes to get all predictions for a single patient series, which was too computationally expensive.</li>\n<li>This led to my final approach using a single backbone pass with the region-masked pooling, which provided a good balance of accuracy and computational efficiency.</li>\n</ol>\n<h2>References</h2>\n<p>[1] <a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/598083\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/598083</a><br>\n[2] Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J., &amp; Maier-Hein, K. H. (2021). nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods, 18(2), 203-211.<br>\n[3] Kirchhoff, Yannick, et al. \"Skeleton recall loss for connectivity conserving and resource efficient segmentation of thin tubular structures.\" European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024.<br>\n[4] <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/mic-dkfz-2nd-place-solution-3d-nnu-net-blob-regres\" target=\"_blank\">https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/mic-dkfz-2nd-place-solution-3d-nnu-net-blob-regres</a><br>\n[5] <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/writeups/yu4u-tattaka-4th-place-solution-source-codes-submi\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/writeups/yu4u-tattaka-4th-place-solution-source-codes-submi</a><br>\n[6] Yeung, Michael, et al. \"Calibrating the dice loss to handle neural network overconfidence for biomedical image segmentation.\" Journal of Digital Imaging 36.2 (2023): 739-752.<br>\n[7] <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/writeups/tomoon33-6th-place-solution\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/writeups/tomoon33-6th-place-solution</a></p>\n<h2>Code Release</h2>\n<ul>\n<li>Training code: <a href=\"https://github.com/uchiyama33/rsna2025_1st_place\" target=\"_blank\">https://github.com/uchiyama33/rsna2025_1st_place</a></li>\n<li>Submission notebook: <a href=\"https://www.kaggle.com/code/tomoon33/rsna2025-submission-1st-place\" target=\"_blank\">https://www.kaggle.com/code/tomoon33/rsna2025-submission-1st-place</a></li>\n</ul>",
  "messages": [
    {
      "id": 3302102,
      "postDate": "2025-10-15T03:36:01.913Z",
      "content": "<p>I thank RSNA, the organizers, the contributing radiologists and institutions, and the Kaggle team for hosting this impactful challenge. This write-up summarizes my 1st place solution. The core of my approach is a robust, coarse-to-fine pipeline that uses vessel segmentation to guide a region-of-interest (ROI) based classifier, producing location-aware predictions.</p>\n<h2>Solution Overview</h2>\n<ul>\n<li><p><strong>High-Level Pipeline</strong></p>\n<ol>\n<li><strong>Preprocessing:</strong> Convert and standardize DICOM series into NIfTI volumes.</li>\n<li><strong>Vessel Segmentation &amp; ROI Extraction:</strong> Use a coarse-to-fine nnU-Net approach. A fast, low-resolution model first finds a candidate region. This allows the more detailed, high-resolution models to focus only on this ROI, which improves both accuracy and speed.</li>\n<li><strong>ROI Classification:</strong> A 3D classification model, using the detailed vessel masks as input, predicts the probabilities for the 13 anatomical locations and the overall aneurysm presence.</li></ol></li>\n<li><p><strong>Key Design Principles</strong></p>\n<ul>\n<li><strong>Coarse-to-Fine Efficiency:</strong> A fast, low-resolution model first scans the entire volume to find a candidate region. This allows the more detailed, high-resolution models to focus only on this ROI. This improves both accuracy and speed.</li>\n<li><strong>Using Segmentation as a Structural Guide:</strong> By providing the classification model with an explicit vessel mask, I give it a detailed map of the vessel structure. This helps the model to focus on the vessels, where aneurysms are located.</li></ul></li>\n</ul>\n<h2>Data Preparation</h2>\n<ul>\n<li>I excluded approximately 60 series due to data quality issues, such as orientation anomalies, corrupted DICOM files, and implausible slice spacing.</li>\n<li>I used multilabel-stratified 5-fold cross-validation. </li>\n</ul>\n<h2>Pipeline</h2>\n<h3>1. Preprocessing</h3>\n<ul>\n<li><strong>Filter Slices:</strong> Within each series, I retained only the images matching the majority <code>Rows × Cols</code> and <code>PixelSpacing</code> configuration and filtered out slices with outlier interslice spacing to ensure consistency.</li>\n<li><strong>Convert DICOM to NIfTI:</strong> I used <code>dcm2niix</code> for conversion. If it failed, I first ran <code>gdcmconv --raw</code> as a fallback before retrying the conversion [1].</li>\n<li><strong>Standardize Orientation:</strong> All volumes were reoriented to a consistent anatomical orientation using nnU-Net’s <code>SimpleITKIOWithReorient</code>.</li>\n<li><strong>Normalize Intensity:</strong> I applied nnU-Net’s standard per-volume z-score normalization to standardize image intensities.</li>\n</ul>\n<h3>2. nnU-Net Segmentation + ROI Extraction (Coarse-to-Fine)</h3>\n<p>This stage uses a sequence of three nnU-Net v2 [2] models (<code>nnUNetResEncUNetMPlans</code>, <code>3d_fullres</code> configuration) to first locate a coarse ROI and then produce detailed vessel segmentations within it.</p>\n<ul>\n<li><p><strong>Model 1: Coarse Vessel Localization</strong></p>\n<ul>\n<li><strong>Spacing:</strong> (1.0, 1.0, 1.0) mm</li>\n<li><strong>Classes:</strong> 3 vessel groups (Posterior+Basilar / MCA / Other)</li>\n<li><strong>Loss:</strong> Dice + Cross-Entropy</li>\n<li><strong>Purpose:</strong> To perform a fast, low-resolution scan to efficiently find a single ROI candidate for the high-resolution models. This model also supports an optional orientation correction step described later.</li></ul></li>\n<li><p><strong>Model 2: Fine Segmentation (Balanced)</strong></p>\n<ul>\n<li><strong>Spacing:</strong> (0.80, 0.45, 0.44) mm</li>\n<li><strong>Loss:</strong> Dice + Cross-Entropy + SkeletonRecall (weight=1) [3]</li>\n<li><strong>Purpose:</strong> To generate a precise vessel segmentation. The SkeletonRecall loss improves the connectivity of thin vessels, which standard losses might miss.</li></ul></li>\n<li><p><strong>Model 3: Fine Segmentation (Recall-Focused)</strong></p>\n<ul>\n<li><strong>Spacing:</strong> (0.80, 0.45, 0.44) mm</li>\n<li><strong>Loss:</strong> Tversky + Cross-Entropy + SkeletonRecall (weight=3)</li>\n<li><strong>Purpose:</strong> To complement Model 2 by prioritizing recall, making it more sensitive to detecting hard-to-find vessel segments.</li></ul></li>\n<li><p><strong>Augmentation Strategy</strong></p>\n<ul>\n<li>I disabled left–right mirroring for the fine models (2 and 3) to preserve anatomical asymmetry.</li>\n<li>I used stronger intensity and geometric augmentations [4].</li>\n<li>I added low-resolution simulation transforms to make the models robust to scans with thick slices.</li></ul></li>\n<li><p><strong>Inference Process (Two-Stage)</strong></p>\n<ul>\n<li><strong>Stage 1 (Coarse Scan):</strong> I run Model 1 with a sliding window (overlap=0.2) and binarize the output to get a foreground mask. I then apply DBSCAN clustering to the mask to remove scattered false positives. The centroid of the largest cluster is used to crop a fixed-size ROI (140×140×140 mm).</li>\n<li><strong>Stage 2 (Fine Inference):</strong> I run Models 2 and 3 with a higher overlap (0.3) only within the coarse ROI. The vessel segmentation from Model 2 is used to compute a tight bounding box, which is then re-cropped with margins to create the final ROI for the classifier.</li></ul></li>\n</ul>\n<p>The detailed segmentations from this stage are passed to the classification model as masks.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F9edf9e0846e14151f33a1ece358473a9%2Fflow_segmentation.png?generation=1760833847347432&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F82fa14c88dfc1dc57448f40bca3ee07d%2Fseg_losses.png?generation=1760833856913829&amp;alt=media\" alt=\"\"></p>\n<h3>3. ROI Classification (13 locations + Aneurysm Present)</h3>\n<p>Once the ROI and vessel masks are prepared, a 3D classification model predicts the final probabilities.</p>\n<ul>\n<li><p><strong>Model Architecture</strong></p>\n<ul>\n<li><strong>Input Size:</strong> The model takes ROI volumes of size 128 × 256 × 256 voxels.</li>\n<li><strong>Backbone:</strong> The core of the model is an nnU-Net pre-trained for the vessel segmentation task. This approach was more accurate and faster to train than standard 2.5D or 3D timm backbones.<ul>\n<li><strong>Decoder Simplification:</strong> To improve efficiency, I simplified the decoder by removing its final, computationally-heavy block. This had no negative impact on performance.</li></ul></li>\n<li><strong>Auxiliary Detection Task:</strong> To help the model learn specific features of aneurysms, I added an auxiliary task that uses the decoder features to reconstruct a small binary sphere (5-pixel radius) at the location of each annotated aneurysm.</li>\n<li><strong>Location-Specific Prediction Head:</strong> To predict the 13 location-specific probabilities, I use the following process:<ol>\n<li><strong>Per-Location Feature Pooling:</strong> A \"Vessel Region-Masked Pooling\" layer uses the vessel masks to extract feature vectors corresponding to each of the 13 anatomical locations from the decoder's feature maps. (<em>Note: In practice, I apply this pooling using masks from both fine segmentation models and concatenate the results for more complete features.</em>)</li>\n<li><strong>Feature Fusion:</strong> These 13 feature vectors are combined with a global feature vector from the encoder (via Global Average Pooling).</li>\n<li><strong>Inter-Location Modeling:</strong> The combined features are fed into a \"Location-Aware Transformer\" to model relationships between different vessel locations.</li>\n<li><strong>Classification:</strong> Finally, an MLP head predicts the probability for each of the 13 locations.</li></ol></li>\n<li><strong>\"Aneurysm Present\" Prediction Head:</strong> For the overall presence prediction, I pool features over the entire vessel structure (a union of all vessel masks) and combine them with the encoder's global features. This aggregated feature vector is passed to a separate MLP head.</li>\n<li><strong>Output Design:</strong> I treated each of the 14 labels as an independent binary classification problem. This design helps with the severe class imbalance, as positive cases for any single location are very rare.</li></ul></li>\n<li><p><strong>Training Details</strong></p>\n<ul>\n<li><strong>Loss Functions:</strong><ul>\n<li><strong>13 Locations:</strong> <code>BCEWithLogitsLoss</code>.</li>\n<li><strong>Aneurysm Present:</strong> <code>BCEWithLogitsLoss</code>.</li>\n<li><strong>Auxiliary Sphere Segmentation:</strong> A combination of Balanced BCE [5] and Focal-Tversky++ loss [6, 7]. This combination works well for highly sparse targets and helps prevent the over-confidence that can occur with Dice-like losses.</li>\n<li><strong>Loss Weights:</strong> The final loss was a weighted sum of the three components. I set the weights to 0.1 for the 13 location losses, 0.05 for the Aneurysm Present loss, and 1.0 for the auxiliary sphere segmentation loss. The main goal was to prioritize learning the precise location of aneurysms, so the sphere segmentation task had the highest weight. I found that higher weights on the classification losses led to overfitting, so this balance was important.</li></ul></li>\n<li><strong>Data Augmentation</strong><ul>\n<li><strong>Intensity Transforms:</strong> Gaussian noise, Gaussian smoothing, intensity shift/scale, contrast adjustment, Gaussian sharpening, and intensity inversion.</li>\n<li><strong>Geometric Transforms:</strong> Random flips (z, y, x axes), small rotations (±10°), scaling/shearing (±10%), mild grid distortions, and a simulated low-resolution transform.</li></ul></li>\n<li><strong>Optimizer and Schedule:</strong> I used the AdamW optimizer with a learning rate of 1e-4 and an effective batch size of 8 (achieved with gradient accumulation). A standard cosine annealing schedule with a warmup period was used.</li>\n<li><strong>EMA Weights:</strong> I used the Exponential Moving Average (EMA) of the model weights for inference.</li></ul></li>\n<li><p><strong>Inference</strong></p>\n<ul>\n<li><strong>Ensembling:</strong> The final predictions are an average of the models from 4 of the 5 cross-validation folds.</li>\n<li><strong>Test-Time Augmentation (TTA):</strong> I averaged the predictions from the original volume and a left-right flipped version of it.</li></ul></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F98e18ca93230a3f26bf2181021ce481a%2Fmodel_overview.png?generation=1760490733871010&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F187c3d0327547a692d8e922dc5f1d5cd%2Fvessel_pooling.png?generation=1760490747217372&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F0dd935cbc51a7d50bc69779c8f12d918%2Ftransformer.png?generation=1760490761936793&amp;alt=media\"></p>\n<h3>4. Other Details</h3>\n<ul>\n<li><strong>Fail-safe Mechanism:</strong> If an anomaly occurred during the segmentation or ROI extraction steps, the pipeline would not attempt a prediction. Instead, it fell back to a set of pre-determined probabilities. For each class, this probability was the mean of the out-of-fold predictions from my cross-validation set.</li>\n<li><strong>Optional Orientation Correction:</strong> I implemented a method to fix misoriented scans. It analyzed the spatial arrangement of the three vessel groups from the coarse segmentation. By comparing the relative positions of these groups to their expected anatomical locations, it estimated the correct axis permutation. This worked perfectly on the training set, fixing all identified orientation issues. However, it had no measurable effect on the leaderboard score, likely because the test set did not contain such orientation errors.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2Ff954825e84986adf4e7b537f9160b1b0%2F3classes_for_ori_corr.png?generation=1760833867211206&amp;alt=media\"></p>\n<h2>Processing Time</h2>\n<ul>\n<li><strong>Training</strong><ul>\n<li>The training times below were measured on a single NVIDIA RTX 4090.</li>\n<li>The 96×192×192 input size was mainly used for faster experimentation and tuning.</li></ul></li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Input size</th>\n<th>Epochs</th>\n<th>Training Time</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>nnU-Net (Model 2)</td>\n<td>64×192×192</td>\n<td>1000</td>\n<td>14 h</td>\n</tr>\n<tr>\n<td>ROI classifier</td>\n<td>96×192×192</td>\n<td>30</td>\n<td>6 h</td>\n</tr>\n<tr>\n<td>ROI classifier</td>\n<td>128×256×256</td>\n<td>30</td>\n<td>12 h</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><strong>Inference</strong><ul>\n<li>Inference times were measured on the Kaggle notebook using two T4 GPUs.</li>\n<li>The times reported are an average over 100 samples.</li></ul></li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Step</th>\n<th>Time per Series</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Preprocessing</td>\n<td>4.03 ± 3.78 s</td>\n</tr>\n<tr>\n<td>Vessel segmentation</td>\n<td>10.65 ± 4.10 s</td>\n</tr>\n<tr>\n<td>ROI classification</td>\n<td>3.33 ± 0.09 s</td>\n</tr>\n</tbody>\n</table>\n<h2>Ablation Study</h2>\n<p>To validate the effectiveness of the key components in my ROI classification model, I conducted an ablation study. Note that this was a simplified evaluation; hyperparameters such as the number of epochs and learning rate were not re-tuned for each experiment. For efficiency, these experiments were run on folds 0, 1, and 2, using a reduced input size of 96×192×192.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Input Size</th>\n<th>AUC (Aneurysm Present)</th>\n<th>AUC (13 Locations)</th>\n<th>Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Final model (full resolution)</td>\n<td>128×256×256</td>\n<td>0.915</td>\n<td>0.916</td>\n<td>0.916</td>\n</tr>\n<tr>\n<td>Final model</td>\n<td>96×192×192</td>\n<td>0.907</td>\n<td>0.898</td>\n<td>0.902</td>\n</tr>\n<tr>\n<td>Without Location-Aware Transformer</td>\n<td>96×192×192</td>\n<td>0.899</td>\n<td>0.894</td>\n<td>0.896</td>\n</tr>\n<tr>\n<td>Using Dice instead of FocalTversky++</td>\n<td>96×192×192</td>\n<td>0.902</td>\n<td>0.896</td>\n<td>0.899</td>\n</tr>\n<tr>\n<td>Setting all loss weights to 1.0</td>\n<td>96×192×192</td>\n<td>0.890</td>\n<td>0.877</td>\n<td>0.884</td>\n</tr>\n<tr>\n<td>Without backbone pretraining</td>\n<td>96×192×192</td>\n<td>0.777</td>\n<td>0.811</td>\n<td>0.794</td>\n</tr>\n<tr>\n<td>Without segmentation model 3</td>\n<td>96×192×192</td>\n<td>0.899</td>\n<td>0.883</td>\n<td>0.891</td>\n</tr>\n<tr>\n<td>Without auxiliary segmentation loss</td>\n<td>96×192×192</td>\n<td>0.880</td>\n<td>0.871</td>\n<td>0.876</td>\n</tr>\n</tbody>\n</table>\n<p>The results show several key points:</p>\n<ul>\n<li>Pretraining the backbone on the vessel segmentation task was the most important factor. It greatly improved the score and helped the model train much faster. This was very helpful for running many experiments.</li>\n<li>The Location-Aware Transformer and the FocalTversky++ loss seemed to help at first, but their final contribution to the score was small. This is likely because other improvements and tuning had a larger overall effect.</li>\n</ul>\n<h2>Design Journey</h2>\n<p>My final design was the result of several iterations:</p>\n<ol>\n<li>I initially tried a single, end-to-end 3D classifier that had auxiliary heads for vessel segmentation and aneurysm localization. While it could detect the presence of an aneurysm, it failed to predict the 13 specific locations accurately.</li>\n<li>I then observed that a standard nnU-Net for vessel segmentation trained easily and generalized well across all modalities. This led me to change to a two-stage, vessel-first pipeline, where the segmentation acts as a strong guide for the subsequent classification task.</li>\n<li>I also experimented with a simpler model that took a 2-channel input: the image volume concatenated with a single binary vessel mask. To get a prediction for a specific label, I would feed the model the corresponding mask (e.g., the mask for one location, or the union mask for \"Aneurysm Present\"). Although the model itself was simple, this approach required running 14 separate forward passes to get all predictions for a single patient series, which was too computationally expensive.</li>\n<li>This led to my final approach using a single backbone pass with the region-masked pooling, which provided a good balance of accuracy and computational efficiency.</li>\n</ol>\n<h2>References</h2>\n<p>[1] <a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/598083\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/598083</a><br>\n[2] Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J., &amp; Maier-Hein, K. H. (2021). nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods, 18(2), 203-211.<br>\n[3] Kirchhoff, Yannick, et al. \"Skeleton recall loss for connectivity conserving and resource efficient segmentation of thin tubular structures.\" European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024.<br>\n[4] <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/mic-dkfz-2nd-place-solution-3d-nnu-net-blob-regres\" target=\"_blank\">https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/mic-dkfz-2nd-place-solution-3d-nnu-net-blob-regres</a><br>\n[5] <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/writeups/yu4u-tattaka-4th-place-solution-source-codes-submi\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/writeups/yu4u-tattaka-4th-place-solution-source-codes-submi</a><br>\n[6] Yeung, Michael, et al. \"Calibrating the dice loss to handle neural network overconfidence for biomedical image segmentation.\" Journal of Digital Imaging 36.2 (2023): 739-752.<br>\n[7] <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/writeups/tomoon33-6th-place-solution\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/writeups/tomoon33-6th-place-solution</a></p>\n<h2>Code Release</h2>\n<ul>\n<li>Training code: <a href=\"https://github.com/uchiyama33/rsna2025_1st_place\" target=\"_blank\">https://github.com/uchiyama33/rsna2025_1st_place</a></li>\n<li>Submission notebook: <a href=\"https://www.kaggle.com/code/tomoon33/rsna2025-submission-1st-place\" target=\"_blank\">https://www.kaggle.com/code/tomoon33/rsna2025-submission-1st-place</a></li>\n</ul>",
      "rawMarkdown": "I thank RSNA, the organizers, the contributing radiologists and institutions, and the Kaggle team for hosting this impactful challenge. This write-up summarizes my 1st place solution. The core of my approach is a robust, coarse-to-fine pipeline that uses vessel segmentation to guide a region-of-interest (ROI) based classifier, producing location-aware predictions.\n\n## Solution Overview\n\n- **High-Level Pipeline**\n  1.  **Preprocessing:** Convert and standardize DICOM series into NIfTI volumes.\n  2.  **Vessel Segmentation & ROI Extraction:** Use a coarse-to-fine nnU-Net approach. A fast, low-resolution model first finds a candidate region. This allows the more detailed, high-resolution models to focus only on this ROI, which improves both accuracy and speed.\n  3.  **ROI Classification:** A 3D classification model, using the detailed vessel masks as input, predicts the probabilities for the 13 anatomical locations and the overall aneurysm presence.\n\n- **Key Design Principles**\n    - **Coarse-to-Fine Efficiency:** A fast, low-resolution model first scans the entire volume to find a candidate region. This allows the more detailed, high-resolution models to focus only on this ROI. This improves both accuracy and speed.\n    - **Using Segmentation as a Structural Guide:** By providing the classification model with an explicit vessel mask, I give it a detailed map of the vessel structure. This helps the model to focus on the vessels, where aneurysms are located.\n\n## Data Preparation\n- I excluded approximately 60 series due to data quality issues, such as orientation anomalies, corrupted DICOM files, and implausible slice spacing.\n- I used multilabel-stratified 5-fold cross-validation. \n\n## Pipeline\n\n### 1. Preprocessing\n- **Filter Slices:** Within each series, I retained only the images matching the majority `Rows × Cols` and `PixelSpacing` configuration and filtered out slices with outlier interslice spacing to ensure consistency.\n- **Convert DICOM to NIfTI:** I used `dcm2niix` for conversion. If it failed, I first ran `gdcmconv --raw` as a fallback before retrying the conversion [1].\n- **Standardize Orientation:** All volumes were reoriented to a consistent anatomical orientation using nnU-Net’s `SimpleITKIOWithReorient`.\n- **Normalize Intensity:** I applied nnU-Net’s standard per-volume z-score normalization to standardize image intensities.\n\n### 2. nnU-Net Segmentation + ROI Extraction (Coarse-to-Fine)\nThis stage uses a sequence of three nnU-Net v2 [2] models (`nnUNetResEncUNetMPlans`, `3d_fullres` configuration) to first locate a coarse ROI and then produce detailed vessel segmentations within it.\n\n- **Model 1: Coarse Vessel Localization**\n  - **Spacing:** (1.0, 1.0, 1.0) mm\n  - **Classes:** 3 vessel groups (Posterior+Basilar / MCA / Other)\n  - **Loss:** Dice + Cross-Entropy\n  - **Purpose:** To perform a fast, low-resolution scan to efficiently find a single ROI candidate for the high-resolution models. This model also supports an optional orientation correction step described later.\n\n- **Model 2: Fine Segmentation (Balanced)**\n  - **Spacing:** (0.80, 0.45, 0.44) mm\n  - **Loss:** Dice + Cross-Entropy + SkeletonRecall (weight=1) [3]\n  - **Purpose:** To generate a precise vessel segmentation. The SkeletonRecall loss improves the connectivity of thin vessels, which standard losses might miss.\n\n- **Model 3: Fine Segmentation (Recall-Focused)**\n  - **Spacing:** (0.80, 0.45, 0.44) mm\n  - **Loss:** Tversky + Cross-Entropy + SkeletonRecall (weight=3)\n  - **Purpose:** To complement Model 2 by prioritizing recall, making it more sensitive to detecting hard-to-find vessel segments.\n\n- **Augmentation Strategy**\n  - I disabled left–right mirroring for the fine models (2 and 3) to preserve anatomical asymmetry.\n  - I used stronger intensity and geometric augmentations [4].\n  - I added low-resolution simulation transforms to make the models robust to scans with thick slices.\n\n- **Inference Process (Two-Stage)**\n  - **Stage 1 (Coarse Scan):** I run Model 1 with a sliding window (overlap=0.2) and binarize the output to get a foreground mask. I then apply DBSCAN clustering to the mask to remove scattered false positives. The centroid of the largest cluster is used to crop a fixed-size ROI (140×140×140 mm).\n  - **Stage 2 (Fine Inference):** I run Models 2 and 3 with a higher overlap (0.3) only within the coarse ROI. The vessel segmentation from Model 2 is used to compute a tight bounding box, which is then re-cropped with margins to create the final ROI for the classifier.\n\nThe detailed segmentations from this stage are passed to the classification model as masks.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F9edf9e0846e14151f33a1ece358473a9%2Fflow_segmentation.png?generation=1760833847347432&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F82fa14c88dfc1dc57448f40bca3ee07d%2Fseg_losses.png?generation=1760833856913829&alt=media)\n\n### 3. ROI Classification (13 locations + Aneurysm Present)\nOnce the ROI and vessel masks are prepared, a 3D classification model predicts the final probabilities.\n\n- **Model Architecture**\n  - **Input Size:** The model takes ROI volumes of size 128 × 256 × 256 voxels.\n  - **Backbone:** The core of the model is an nnU-Net pre-trained for the vessel segmentation task. This approach was more accurate and faster to train than standard 2.5D or 3D timm backbones.\n      - **Decoder Simplification:** To improve efficiency, I simplified the decoder by removing its final, computationally-heavy block. This had no negative impact on performance.\n  - **Auxiliary Detection Task:** To help the model learn specific features of aneurysms, I added an auxiliary task that uses the decoder features to reconstruct a small binary sphere (5-pixel radius) at the location of each annotated aneurysm.\n  - **Location-Specific Prediction Head:** To predict the 13 location-specific probabilities, I use the following process:\n      1.  **Per-Location Feature Pooling:** A \"Vessel Region-Masked Pooling\" layer uses the vessel masks to extract feature vectors corresponding to each of the 13 anatomical locations from the decoder's feature maps. (*Note: In practice, I apply this pooling using masks from both fine segmentation models and concatenate the results for more complete features.*)\n      2.  **Feature Fusion:** These 13 feature vectors are combined with a global feature vector from the encoder (via Global Average Pooling).\n      3.  **Inter-Location Modeling:** The combined features are fed into a \"Location-Aware Transformer\" to model relationships between different vessel locations.\n      4.  **Classification:** Finally, an MLP head predicts the probability for each of the 13 locations.\n  - **\"Aneurysm Present\" Prediction Head:** For the overall presence prediction, I pool features over the entire vessel structure (a union of all vessel masks) and combine them with the encoder's global features. This aggregated feature vector is passed to a separate MLP head.\n  - **Output Design:** I treated each of the 14 labels as an independent binary classification problem. This design helps with the severe class imbalance, as positive cases for any single location are very rare.\n\n- **Training Details**\n  - **Loss Functions:**\n      - **13 Locations:** `BCEWithLogitsLoss`.\n      - **Aneurysm Present:** `BCEWithLogitsLoss`.\n      - **Auxiliary Sphere Segmentation:** A combination of Balanced BCE [5] and Focal-Tversky++ loss [6, 7]. This combination works well for highly sparse targets and helps prevent the over-confidence that can occur with Dice-like losses.\n      - **Loss Weights:** The final loss was a weighted sum of the three components. I set the weights to 0.1 for the 13 location losses, 0.05 for the Aneurysm Present loss, and 1.0 for the auxiliary sphere segmentation loss. The main goal was to prioritize learning the precise location of aneurysms, so the sphere segmentation task had the highest weight. I found that higher weights on the classification losses led to overfitting, so this balance was important.\n  - **Data Augmentation**\n      - **Intensity Transforms:** Gaussian noise, Gaussian smoothing, intensity shift/scale, contrast adjustment, Gaussian sharpening, and intensity inversion.\n      - **Geometric Transforms:** Random flips (z, y, x axes), small rotations (±10°), scaling/shearing (±10%), mild grid distortions, and a simulated low-resolution transform.\n  - **Optimizer and Schedule:** I used the AdamW optimizer with a learning rate of 1e-4 and an effective batch size of 8 (achieved with gradient accumulation). A standard cosine annealing schedule with a warmup period was used.\n  - **EMA Weights:** I used the Exponential Moving Average (EMA) of the model weights for inference.\n\n- **Inference**\n  - **Ensembling:** The final predictions are an average of the models from 4 of the 5 cross-validation folds.\n  - **Test-Time Augmentation (TTA):** I averaged the predictions from the original volume and a left-right flipped version of it.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F98e18ca93230a3f26bf2181021ce481a%2Fmodel_overview.png?generation=1760490733871010&alt=media)\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F187c3d0327547a692d8e922dc5f1d5cd%2Fvessel_pooling.png?generation=1760490747217372&alt=media\" width=\"80%\">\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F0dd935cbc51a7d50bc69779c8f12d918%2Ftransformer.png?generation=1760490761936793&alt=media\" width=\"60%\">\n\n### 4. Other Details\n- **Fail-safe Mechanism:** If an anomaly occurred during the segmentation or ROI extraction steps, the pipeline would not attempt a prediction. Instead, it fell back to a set of pre-determined probabilities. For each class, this probability was the mean of the out-of-fold predictions from my cross-validation set.\n- **Optional Orientation Correction:** I implemented a method to fix misoriented scans. It analyzed the spatial arrangement of the three vessel groups from the coarse segmentation. By comparing the relative positions of these groups to their expected anatomical locations, it estimated the correct axis permutation. This worked perfectly on the training set, fixing all identified orientation issues. However, it had no measurable effect on the leaderboard score, likely because the test set did not contain such orientation errors.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2Ff954825e84986adf4e7b537f9160b1b0%2F3classes_for_ori_corr.png?generation=1760833867211206&alt=media\" width=\"40%\">\n\n## Processing Time\n\n- **Training**\n  - The training times below were measured on a single NVIDIA RTX 4090.\n  - The 96×192×192 input size was mainly used for faster experimentation and tuning.\n\n| Model | Input size | Epochs | Training Time |\n|---|---|---|---|\n| nnU-Net (Model 2) | 64×192×192 | 1000 | 14 h |\n| ROI classifier | 96×192×192 | 30 | 6 h |\n| ROI classifier | 128×256×256 | 30 | 12 h |\n\n\n- **Inference**\n  - Inference times were measured on the Kaggle notebook using two T4 GPUs.\n  - The times reported are an average over 100 samples.\n\n| Step | Time per Series |\n|---|---|\n| Preprocessing | 4.03 ± 3.78 s |\n| Vessel segmentation | 10.65 ± 4.10 s |\n| ROI classification | 3.33 ± 0.09 s |\n \n \n## Ablation Study\nTo validate the effectiveness of the key components in my ROI classification model, I conducted an ablation study. Note that this was a simplified evaluation; hyperparameters such as the number of epochs and learning rate were not re-tuned for each experiment. For efficiency, these experiments were run on folds 0, 1, and 2, using a reduced input size of 96×192×192.\n\n| Model | Input Size | AUC (Aneurysm Present) | AUC (13 Locations) | Score |\n|---|---|---|---|---|\n| Final model (full resolution) | 128×256×256 | 0.915 | 0.916 | 0.916 |\n| Final model | 96×192×192 | 0.907 | 0.898 | 0.902 |\n| Without Location-Aware Transformer | 96×192×192 | 0.899 | 0.894 | 0.896 |\n| Using Dice instead of FocalTversky++ | 96×192×192 | 0.902 | 0.896 | 0.899 |\n| Setting all loss weights to 1.0 | 96×192×192 | 0.890 | 0.877 | 0.884 |\n| Without backbone pretraining | 96×192×192 | 0.777 | 0.811 | 0.794 |\n| Without segmentation model 3 | 96×192×192 | 0.899 | 0.883 | 0.891 |\n| Without auxiliary segmentation loss | 96×192×192 | 0.880 | 0.871 | 0.876 |\n\nThe results show several key points:\n- Pretraining the backbone on the vessel segmentation task was the most important factor. It greatly improved the score and helped the model train much faster. This was very helpful for running many experiments.\n- The Location-Aware Transformer and the FocalTversky++ loss seemed to help at first, but their final contribution to the score was small. This is likely because other improvements and tuning had a larger overall effect.\n\n\n## Design Journey\nMy final design was the result of several iterations:\n\n1.  I initially tried a single, end-to-end 3D classifier that had auxiliary heads for vessel segmentation and aneurysm localization. While it could detect the presence of an aneurysm, it failed to predict the 13 specific locations accurately.\n2.  I then observed that a standard nnU-Net for vessel segmentation trained easily and generalized well across all modalities. This led me to change to a two-stage, vessel-first pipeline, where the segmentation acts as a strong guide for the subsequent classification task.\n3.  I also experimented with a simpler model that took a 2-channel input: the image volume concatenated with a single binary vessel mask. To get a prediction for a specific label, I would feed the model the corresponding mask (e.g., the mask for one location, or the union mask for \"Aneurysm Present\"). Although the model itself was simple, this approach required running 14 separate forward passes to get all predictions for a single patient series, which was too computationally expensive.\n4.  This led to my final approach using a single backbone pass with the region-masked pooling, which provided a good balance of accuracy and computational efficiency.\n\n\n## References\n[1] https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/598083\n[2] Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J., & Maier-Hein, K. H. (2021). nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods, 18(2), 203-211.\n[3] Kirchhoff, Yannick, et al. \"Skeleton recall loss for connectivity conserving and resource efficient segmentation of thin tubular structures.\" European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024.\n[4] https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/mic-dkfz-2nd-place-solution-3d-nnu-net-blob-regres\n[5] https://www.kaggle.com/competitions/czii-cryo-et-object-identification/writeups/yu4u-tattaka-4th-place-solution-source-codes-submi\n[6] Yeung, Michael, et al. \"Calibrating the dice loss to handle neural network overconfidence for biomedical image segmentation.\" Journal of Digital Imaging 36.2 (2023): 739-752.\n[7] https://www.kaggle.com/competitions/czii-cryo-et-object-identification/writeups/tomoon33-6th-place-solution\n\n## Code Release\n- Training code: https://github.com/uchiyama33/rsna2025_1st_place\n- Submission notebook: https://www.kaggle.com/code/tomoon33/rsna2025-submission-1st-place",
      "votes": 158
    },
    {
      "id": 3302258,
      "postDate": "2025-10-15T12:18:38.653Z",
      "content": "<p>Congrats!! Glad Skeleton Recall Loss could provide some benefit here as well 💯</p>",
      "rawMarkdown": "Congrats!! Glad Skeleton Recall Loss could provide some benefit here as well 💯\n",
      "votes": 3
    },
    {
      "id": 3344478,
      "postDate": "2025-11-22T15:50:37.223Z",
      "content": "<p>Hey, could you share the W&amp;B logs and the train/validation loss curves?</p>",
      "rawMarkdown": "Hey, could you share the W&B logs and the train/validation loss curves?",
      "votes": 1
    },
    {
      "id": 3303796,
      "postDate": "2025-10-19T01:10:20.273Z",
      "content": "<p>I've just updated my solution write-up with a more detailed version. The code will be released shortly.</p>",
      "rawMarkdown": "I've just updated my solution write-up with a more detailed version. The code will be released shortly.",
      "votes": 4,
      "replies": [
        {
          "id": 3304763,
          "postDate": "2025-10-21T09:48:11.697Z",
          "content": "<p>Okay, we are waiting for it 👍</p>",
          "rawMarkdown": "Okay, we are waiting for it 👍"
        },
        {
          "id": 3304835,
          "postDate": "2025-10-21T11:58:36.647Z",
          "content": "<p>I've updated my solution write-up again. The code has been released, and I've added a new section describing the ablation study. I hope you find it helpful!</p>",
          "rawMarkdown": "I've updated my solution write-up again. The code has been released, and I've added a new section describing the ablation study. I hope you find it helpful!"
        }
      ]
    },
    {
      "id": 3302359,
      "postDate": "2025-10-15T16:29:47.380Z",
      "content": "<p>Congratulations on taking 1st place solo!!👏</p>\n<p>How was this architecture derived? It's nothing short of spectacular.<br>\nI'd love to hear about the process that led to this model.</p>",
      "rawMarkdown": "Congratulations on taking 1st place solo!!👏\n\nHow was this architecture derived? It's nothing short of spectacular.\nI'd love to hear about the process that led to this model.",
      "votes": 1,
      "replies": [
        {
          "id": 3303789,
          "postDate": "2025-10-19T00:51:24.873Z",
          "content": "<p>I just published an updated write-up with a “Design Journey” section that explains how this architecture came together. Please check that section.</p>",
          "rawMarkdown": "I just published an updated write-up with a “Design Journey” section that explains how this architecture came together. Please check that section.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3302107,
      "postDate": "2025-10-15T04:06:23.260Z",
      "content": "<p>Congratulations on the win! The auxiliary tasks using nnunet are really good. I wonder, what was the Dice score of your nnunet for vessel segmentation? I managed to get one close to 0.78. I could not do much with the ROIs afterwards, because I did not have enough compute to generate entire train data segmentations (disk space issue). I only tried a PoC with 128x128x128 volumes using another segmentation model, but nothing worked. </p>",
      "rawMarkdown": "Congratulations on the win! The auxiliary tasks using nnunet are really good. I wonder, what was the Dice score of your nnunet for vessel segmentation? I managed to get one close to 0.78. I could not do much with the ROIs afterwards, because I did not have enough compute to generate entire train data segmentations (disk space issue). I only tried a PoC with 128x128x128 volumes using another segmentation model, but nothing worked. ",
      "votes": 1,
      "replies": [
        {
          "id": 3302506,
          "postDate": "2025-10-16T01:28:09.953Z",
          "content": "<p>A Dice of 0.78 is very good. For my nnU-Net (Model 2), the validation Dice was around 0.70. I think there was still room to improve.</p>\n<p>However, in my solution the Dice itself was not so important. I used the vessel segmentation for masking. So I focused to have high recall, especially on small regions and peripheral branches. </p>",
          "rawMarkdown": "A Dice of 0.78 is very good. For my nnU-Net (Model 2), the validation Dice was around 0.70. I think there was still room to improve.\n\nHowever, in my solution the Dice itself was not so important. I used the vessel segmentation for masking. So I focused to have high recall, especially on small regions and peripheral branches. ",
          "votes": 1,
          "replies": [
            {
              "id": 3302509,
              "postDate": "2025-10-16T01:33:57.357Z",
              "content": "<p>Makes sense, thanks. I have some more questions, but I will wait for the detailed solution first :D <br>\nCongratulations again!</p>",
              "rawMarkdown": "Makes sense, thanks. I have some more questions, but I will wait for the detailed solution first :D \nCongratulations again!"
            }
          ]
        }
      ]
    },
    {
      "id": 3302525,
      "postDate": "2025-10-16T02:21:24.563Z",
      "content": "<p>very perfect solution. would you like to share more detail about your solution. or share you inference and training code. thanks very much</p>",
      "rawMarkdown": "very perfect solution. would you like to share more detail about your solution. or share you inference and training code. thanks very much",
      "votes": 2
    },
    {
      "id": 3346908,
      "postDate": "2025-11-24T18:00:17.983Z",
      "content": "<ul>\n<li>In ./classification/models, I saw a class called AneurysmModel. Is it used for segmentation?</li>\n<li>If I want to replace Stage 2 classification with segmentation, is there any existing configuration to train Stage 2 as a segmentation model?\"</li>\n</ul>",
      "rawMarkdown": "- In ./classification/models, I saw a class called AneurysmModel. Is it used for segmentation?\n- If I want to replace Stage 2 classification with segmentation, is there any existing configuration to train Stage 2 as a segmentation model?\""
    },
    {
      "id": 3304670,
      "postDate": "2025-10-21T04:11:13.540Z",
      "content": "<p>Thank You so much for sharing <a href=\"https://www.kaggle.com/tomoon33\" target=\"_blank\">@tomoon33</a> </p>",
      "rawMarkdown": "Thank You so much for sharing @tomoon33 "
    },
    {
      "id": 3303903,
      "postDate": "2025-10-19T08:12:33.447Z",
      "content": "<p>Great solution….<br>\nCould you please explain how vessel segmentation is performed since the dataset only contains the vessel location of where aneurysm&nbsp;is&nbsp;present? What dataset did you use to train the coarse nn-Unet</p>",
      "rawMarkdown": "Great solution....\nCould you please explain how vessel segmentation is performed since the dataset only contains the vessel location of where aneurysm is present? What dataset did you use to train the coarse nn-Unet"
    },
    {
      "id": 3303642,
      "postDate": "2025-10-18T15:19:22.977Z",
      "content": "<p>Congratulations on the win! </p>",
      "rawMarkdown": "Congratulations on the win! "
    },
    {
      "id": 3303580,
      "postDate": "2025-10-18T13:15:10.180Z",
      "content": "<p>Congratulations on your win, amazing work!</p>\n<p>Your writeup diagrams are wonderful. May I ask which tool you used to create them?</p>",
      "rawMarkdown": "Congratulations on your win, amazing work!\n\nYour writeup diagrams are wonderful. May I ask which tool you used to create them?",
      "replies": [
        {
          "id": 3303794,
          "postDate": "2025-10-19T01:07:00.090Z",
          "content": "<p>Thank you! I made the diagrams in PowerPoint.</p>",
          "rawMarkdown": "Thank you! I made the diagrams in PowerPoint."
        }
      ]
    },
    {
      "id": 3303512,
      "postDate": "2025-10-18T08:14:35.303Z",
      "content": "<p>Congratulations on 1st place. Your pipeline looks very robust. I’m particularly curious about how you managed inference speed with the coarse-to-fine approach did it add much latency, or was it still efficient at scale? Would love to try something similar in my experiments.</p>",
      "rawMarkdown": "Congratulations on 1st place. Your pipeline looks very robust. I’m particularly curious about how you managed inference speed with the coarse-to-fine approach did it add much latency, or was it still efficient at scale? Would love to try something similar in my experiments.",
      "replies": [
        {
          "id": 3303791,
          "postDate": "2025-10-19T00:58:06.633Z",
          "content": "<p>Thank you! I can’t share exact timings, but the coarse‑to‑fine setup did not increase end‑to‑end latency. On large CTA volumes, it was typically faster and much more memory‑stable.</p>",
          "rawMarkdown": "Thank you! I can’t share exact timings, but the coarse‑to‑fine setup did not increase end‑to‑end latency. On large CTA volumes, it was typically faster and much more memory‑stable."
        }
      ]
    },
    {
      "id": 3303471,
      "postDate": "2025-10-18T05:39:21.310Z",
      "content": "<p>Congratulations! Could you explain how you developed this model from nnUNet to the current complete version? What inspired you?</p>",
      "rawMarkdown": "Congratulations! Could you explain how you developed this model from nnUNet to the current complete version? What inspired you?",
      "replies": [
        {
          "id": 3303792,
          "postDate": "2025-10-19T00:59:56.993Z",
          "content": "<p>Thank you! I just published an updated write-up with a “Design Journey” section that explains how the system evolved from an nnU-Net baseline to the complete pipeline, and what inspired the key design choices. Please check that section.</p>",
          "rawMarkdown": "Thank you! I just published an updated write-up with a “Design Journey” section that explains how the system evolved from an nnU-Net baseline to the complete pipeline, and what inspired the key design choices. Please check that section.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3303443,
      "postDate": "2025-10-18T03:31:33.297Z",
      "content": "<p>Really love this solution! Congrats</p>",
      "rawMarkdown": "Really love this solution! Congrats"
    },
    {
      "id": 3302875,
      "postDate": "2025-10-16T16:57:02.450Z",
      "content": "<p>Super congratulations! <br>\nVery well deserved, beautiful work. </p>\n<p>M</p>",
      "rawMarkdown": "Super congratulations! \nVery well deserved, beautiful work. \n\nM"
    },
    {
      "id": 3302797,
      "postDate": "2025-10-16T14:26:42.750Z",
      "content": "<p>congrats and thanks for the nice summary. would it be possible to make the cleaned dataset publicly available?</p>",
      "rawMarkdown": "congrats and thanks for the nice summary. would it be possible to make the cleaned dataset publicly available?",
      "replies": [
        {
          "id": 3303793,
          "postDate": "2025-10-19T01:04:24.867Z",
          "content": "<p>Thank you! I’ll release the exact cleaning scripts and a list of excluded series so you can reproduce the cleaned set from the official data. </p>",
          "rawMarkdown": "Thank you! I’ll release the exact cleaning scripts and a list of excluded series so you can reproduce the cleaned set from the official data. "
        }
      ]
    },
    {
      "id": 3302777,
      "postDate": "2025-10-16T13:34:41.003Z",
      "content": "<p>It's really good </p>",
      "rawMarkdown": "It's really good "
    },
    {
      "id": 3302152,
      "postDate": "2025-10-15T07:16:47.807Z",
      "content": "<p>Congratulations! Nice usage of nnU-Net 😍</p>",
      "rawMarkdown": "Congratulations! Nice usage of nnU-Net 😍"
    },
    {
      "id": 3302134,
      "postDate": "2025-10-15T06:24:05.283Z",
      "content": "<p>Congratulations on the win! Brilliant design! I was wondering, how much did the location-aware transformers contribute to the results? Also, how were the weights in the loss function allocated?</p>",
      "rawMarkdown": "Congratulations on the win! Brilliant design! I was wondering, how much did the location-aware transformers contribute to the results? Also, how were the weights in the loss function allocated?",
      "replies": [
        {
          "id": 3302511,
          "postDate": "2025-10-16T01:38:43.723Z",
          "content": "<p>Thanks! In my experiments, the location‑aware transformer added about +0.02 to the macro AUC across the 13 location labels. For the loss, I put relatively larger weight on the auxiliary segmentation loss, so the model prioritized learning aneurysm localization. </p>\n<p>I plan to add more details to the solution write‑up later.</p>",
          "rawMarkdown": "Thanks! In my experiments, the location‑aware transformer added about +0.02 to the macro AUC across the 13 location labels. For the loss, I put relatively larger weight on the auxiliary segmentation loss, so the model prioritized learning aneurysm localization. \n\nI plan to add more details to the solution write‑up later."
        }
      ]
    },
    {
      "id": 3302132,
      "postDate": "2025-10-15T06:21:13.230Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tomoon33\" target=\"_blank\">@tomoon33</a> <br>\nWould you mind sharing why spacing = (1.0, 1.0, 1.0) and  (0.80, 0.45, 0.44) were chosen? It is by experiment or what?<br>\nThanks.</p>",
      "rawMarkdown": "Congratulations @tomoon33 \nWould you mind sharing why spacing = (1.0, 1.0, 1.0) and  (0.80, 0.45, 0.44) were chosen? It is by experiment or what?\nThanks.",
      "replies": [
        {
          "id": 3302512,
          "postDate": "2025-10-16T01:44:17.667Z",
          "content": "<p>Thanks for the question.</p>\n<p>(0.80, 0.45, 0.44): This came from nnU‑Net’s auto‑configuration. I just used the suggested value.<br>\n(1.0, 1.0, 1.0): An isotropic, coarser resampling for a “coarse scan” path to save time and memory. It has not been carefully tuned.</p>",
          "rawMarkdown": "Thanks for the question.\n\n(0.80, 0.45, 0.44): This came from nnU‑Net’s auto‑configuration. I just used the suggested value.\n(1.0, 1.0, 1.0): An isotropic, coarser resampling for a “coarse scan” path to save time and memory. It has not been carefully tuned."
        }
      ]
    },
    {
      "id": 3302113,
      "postDate": "2025-10-15T04:20:50.673Z",
      "content": "<p>Congratulations on your win—well deserved! Your solution wrap-up was very clear.<br>\nI also have a couple of questions: Did you encounter any aneurysms located outside the segmentation ROI, and if so, how did you address them? I'm also curious about your Dice scores for the Right and Left PCA, which are quite small structures and appear to have significant missing labels in the provided segmentation data.</p>",
      "rawMarkdown": "Congratulations on your win—well deserved! Your solution wrap-up was very clear.\nI also have a couple of questions: Did you encounter any aneurysms located outside the segmentation ROI, and if so, how did you address them? I'm also curious about your Dice scores for the Right and Left PCA, which are quite small structures and appear to have significant missing labels in the provided segmentation data.",
      "replies": [
        {
          "id": 3302514,
          "postDate": "2025-10-16T01:54:33.410Z",
          "content": "<p>Thanks for the kind words.</p>\n<p>Aneurysms outside ROI: In my cleaned training set, I found 10 aneurysm labels that fell outside the vessel‑based ROI. They were mostly superior outliers. I considered expanding the superior margin, but since the count was small and a larger ROI would raise memory/compute, I decided not to change it and accepted those rare cases.</p>\n<p>Dice for R/L PCA: I didn’t compute Dice specifically for the PCA labels, so I don’t have exact numbers to share. Because PCA is tiny, I prioritized recall by adding SkeletonRecall loss. This made the segmentation somewhat over‑inclusive, but it eliminated misses.</p>",
          "rawMarkdown": "Thanks for the kind words.\n\nAneurysms outside ROI: In my cleaned training set, I found 10 aneurysm labels that fell outside the vessel‑based ROI. They were mostly superior outliers. I considered expanding the superior margin, but since the count was small and a larger ROI would raise memory/compute, I decided not to change it and accepted those rare cases.\n\nDice for R/L PCA: I didn’t compute Dice specifically for the PCA labels, so I don’t have exact numbers to share. Because PCA is tiny, I prioritized recall by adding SkeletonRecall loss. This made the segmentation somewhat over‑inclusive, but it eliminated misses.",
          "votes": 1,
          "replies": [
            {
              "id": 3302563,
              "postDate": "2025-10-16T03:52:57.343Z",
              "content": "<p>Thanks for the detailed reply!</p>",
              "rawMarkdown": "Thanks for the detailed reply!"
            }
          ]
        }
      ]
    },
    {
      "id": 3302104,
      "postDate": "2025-10-15T03:49:35.173Z",
      "content": "<p><a href=\"https://www.kaggle.com/tomoon33\" target=\"_blank\">@tomoon33</a>  It seems nnUnet can have better representation than YOLO. I've tried my BYU solution + <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> previous RSNA solution but result is not expected like yours. I really like your development of ROI classification stage, Great work. </p>",
      "rawMarkdown": "@tomoon33  It seems nnUnet can have better representation than YOLO. I've tried my BYU solution + @ren4yu previous RSNA solution but result is not expected like yours. I really like your development of ROI classification stage, Great work. ",
      "replies": [
        {
          "id": 3302147,
          "postDate": "2025-10-15T06:56:53.390Z",
          "content": "<p>My intuition that excluding bad data probably plays a big role. However, It needs professional knowledge.</p>",
          "rawMarkdown": "My intuition that excluding bad data probably plays a big role. However, It needs professional knowledge."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3302258,
      "author_name": "Maximilian Rokuss",
      "author_url": "",
      "post_date": "2025-10-15T12:18:38.653000",
      "content": "<p>Congrats!! Glad Skeleton Recall Loss could provide some benefit here as well 💯</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3344478,
      "author_name": "xuananle",
      "author_url": "",
      "post_date": "2025-11-22T15:50:37.223000",
      "content": "<p>Hey, could you share the W&amp;B logs and the train/validation loss curves?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3303796,
      "author_name": "tomoon33",
      "author_url": "",
      "post_date": "2025-10-19T01:10:20.273000",
      "content": "<p>I've just updated my solution write-up with a more detailed version. The code will be released shortly.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 3304763,
          "author_name": "Irakoze Ntawigenga Kelly",
          "author_url": "",
          "post_date": "2025-10-21T09:48:11.697000",
          "content": "<p>Okay, we are waiting for it 👍</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3304835,
          "author_name": "tomoon33",
          "author_url": "",
          "post_date": "2025-10-21T11:58:36.647000",
          "content": "<p>I've updated my solution write-up again. The code has been released, and I've added a new section describing the ablation study. I hope you find it helpful!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3302359,
      "author_name": "Ataracsia",
      "author_url": "",
      "post_date": "2025-10-15T16:29:47.380000",
      "content": "<p>Congratulations on taking 1st place solo!!👏</p>\n<p>How was this architecture derived? It's nothing short of spectacular.<br>\nI'd love to hear about the process that led to this model.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3303789,
          "author_name": "tomoon33",
          "author_url": "",
          "post_date": "2025-10-19T00:51:24.873000",
          "content": "<p>I just published an updated write-up with a “Design Journey” section that explains how this architecture came together. Please check that section.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3302107,
      "author_name": "Satwik",
      "author_url": "",
      "post_date": "2025-10-15T04:06:23.260000",
      "content": "<p>Congratulations on the win! The auxiliary tasks using nnunet are really good. I wonder, what was the Dice score of your nnunet for vessel segmentation? I managed to get one close to 0.78. I could not do much with the ROIs afterwards, because I did not have enough compute to generate entire train data segmentations (disk space issue). I only tried a PoC with 128x128x128 volumes using another segmentation model, but nothing worked. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3302506,
          "author_name": "tomoon33",
          "author_url": "",
          "post_date": "2025-10-16T01:28:09.953000",
          "content": "<p>A Dice of 0.78 is very good. For my nnU-Net (Model 2), the validation Dice was around 0.70. I think there was still room to improve.</p>\n<p>However, in my solution the Dice itself was not so important. I used the vessel segmentation for masking. So I focused to have high recall, especially on small regions and peripheral branches. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3302509,
              "author_name": "Satwik",
              "author_url": "",
              "post_date": "2025-10-16T01:33:57.357000",
              "content": "<p>Makes sense, thanks. I have some more questions, but I will wait for the detailed solution first :D <br>\nCongratulations again!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3302525,
      "author_name": "Bill Fan",
      "author_url": "",
      "post_date": "2025-10-16T02:21:24.563000",
      "content": "<p>very perfect solution. would you like to share more detail about your solution. or share you inference and training code. thanks very much</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3346908,
      "author_name": "Đỗ Quốc Minh Nghĩa",
      "author_url": "",
      "post_date": "2025-11-24T18:00:17.983000",
      "content": "<ul>\n<li>In ./classification/models, I saw a class called AneurysmModel. Is it used for segmentation?</li>\n<li>If I want to replace Stage 2 classification with segmentation, is there any existing configuration to train Stage 2 as a segmentation model?\"</li>\n</ul>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3304670,
      "author_name": "yeshwant",
      "author_url": "",
      "post_date": "2025-10-21T04:11:13.540000",
      "content": "<p>Thank You so much for sharing <a href=\"https://www.kaggle.com/tomoon33\" target=\"_blank\">@tomoon33</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3303903,
      "author_name": "Jeslin Thomaskutty",
      "author_url": "",
      "post_date": "2025-10-19T08:12:33.447000",
      "content": "<p>Great solution….<br>\nCould you please explain how vessel segmentation is performed since the dataset only contains the vessel location of where aneurysm&nbsp;is&nbsp;present? What dataset did you use to train the coarse nn-Unet</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3303642,
      "author_name": "Khushi Yadav",
      "author_url": "",
      "post_date": "2025-10-18T15:19:22.977000",
      "content": "<p>Congratulations on the win! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3303580,
      "author_name": "vsemionov",
      "author_url": "",
      "post_date": "2025-10-18T13:15:10.180000",
      "content": "<p>Congratulations on your win, amazing work!</p>\n<p>Your writeup diagrams are wonderful. May I ask which tool you used to create them?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3303794,
          "author_name": "tomoon33",
          "author_url": "",
          "post_date": "2025-10-19T01:07:00.090000",
          "content": "<p>Thank you! I made the diagrams in PowerPoint.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3303512,
      "author_name": "Dimpi Mittal",
      "author_url": "",
      "post_date": "2025-10-18T08:14:35.303000",
      "content": "<p>Congratulations on 1st place. Your pipeline looks very robust. I’m particularly curious about how you managed inference speed with the coarse-to-fine approach did it add much latency, or was it still efficient at scale? Would love to try something similar in my experiments.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3303791,
          "author_name": "tomoon33",
          "author_url": "",
          "post_date": "2025-10-19T00:58:06.633000",
          "content": "<p>Thank you! I can’t share exact timings, but the coarse‑to‑fine setup did not increase end‑to‑end latency. On large CTA volumes, it was typically faster and much more memory‑stable.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3303471,
      "author_name": "Seeing Times",
      "author_url": "",
      "post_date": "2025-10-18T05:39:21.310000",
      "content": "<p>Congratulations! Could you explain how you developed this model from nnUNet to the current complete version? What inspired you?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3303792,
          "author_name": "tomoon33",
          "author_url": "",
          "post_date": "2025-10-19T00:59:56.993000",
          "content": "<p>Thank you! I just published an updated write-up with a “Design Journey” section that explains how the system evolved from an nnU-Net baseline to the complete pipeline, and what inspired the key design choices. Please check that section.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3303443,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2025-10-18T03:31:33.297000",
      "content": "<p>Really love this solution! Congrats</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3302875,
      "author_name": "MD The Dramatist",
      "author_url": "",
      "post_date": "2025-10-16T16:57:02.450000",
      "content": "<p>Super congratulations! <br>\nVery well deserved, beautiful work. </p>\n<p>M</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3302797,
      "author_name": "Ma Edward",
      "author_url": "",
      "post_date": "2025-10-16T14:26:42.750000",
      "content": "<p>congrats and thanks for the nice summary. would it be possible to make the cleaned dataset publicly available?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3303793,
          "author_name": "tomoon33",
          "author_url": "",
          "post_date": "2025-10-19T01:04:24.867000",
          "content": "<p>Thank you! I’ll release the exact cleaning scripts and a list of excluded series so you can reproduce the cleaned set from the official data. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3302777,
      "author_name": "Manjunath_uppar",
      "author_url": "",
      "post_date": "2025-10-16T13:34:41.003000",
      "content": "<p>It's really good </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3302152,
      "author_name": "FabianIsensee",
      "author_url": "",
      "post_date": "2025-10-15T07:16:47.807000",
      "content": "<p>Congratulations! Nice usage of nnU-Net 😍</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3302134,
      "author_name": "Shuolin Liu",
      "author_url": "",
      "post_date": "2025-10-15T06:24:05.283000",
      "content": "<p>Congratulations on the win! Brilliant design! I was wondering, how much did the location-aware transformers contribute to the results? Also, how were the weights in the loss function allocated?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3302511,
          "author_name": "tomoon33",
          "author_url": "",
          "post_date": "2025-10-16T01:38:43.723000",
          "content": "<p>Thanks! In my experiments, the location‑aware transformer added about +0.02 to the macro AUC across the 13 location labels. For the loss, I put relatively larger weight on the auxiliary segmentation loss, so the model prioritized learning aneurysm localization. </p>\n<p>I plan to add more details to the solution write‑up later.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3302132,
      "author_name": "Dennis",
      "author_url": "",
      "post_date": "2025-10-15T06:21:13.230000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tomoon33\" target=\"_blank\">@tomoon33</a> <br>\nWould you mind sharing why spacing = (1.0, 1.0, 1.0) and  (0.80, 0.45, 0.44) were chosen? It is by experiment or what?<br>\nThanks.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3302512,
          "author_name": "tomoon33",
          "author_url": "",
          "post_date": "2025-10-16T01:44:17.667000",
          "content": "<p>Thanks for the question.</p>\n<p>(0.80, 0.45, 0.44): This came from nnU‑Net’s auto‑configuration. I just used the suggested value.<br>\n(1.0, 1.0, 1.0): An isotropic, coarser resampling for a “coarse scan” path to save time and memory. It has not been carefully tuned.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3302113,
      "author_name": "Luca",
      "author_url": "",
      "post_date": "2025-10-15T04:20:50.673000",
      "content": "<p>Congratulations on your win—well deserved! Your solution wrap-up was very clear.<br>\nI also have a couple of questions: Did you encounter any aneurysms located outside the segmentation ROI, and if so, how did you address them? I'm also curious about your Dice scores for the Right and Left PCA, which are quite small structures and appear to have significant missing labels in the provided segmentation data.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3302514,
          "author_name": "tomoon33",
          "author_url": "",
          "post_date": "2025-10-16T01:54:33.410000",
          "content": "<p>Thanks for the kind words.</p>\n<p>Aneurysms outside ROI: In my cleaned training set, I found 10 aneurysm labels that fell outside the vessel‑based ROI. They were mostly superior outliers. I considered expanding the superior margin, but since the count was small and a larger ROI would raise memory/compute, I decided not to change it and accepted those rare cases.</p>\n<p>Dice for R/L PCA: I didn’t compute Dice specifically for the PCA labels, so I don’t have exact numbers to share. Because PCA is tiny, I prioritized recall by adding SkeletonRecall loss. This made the segmentation somewhat over‑inclusive, but it eliminated misses.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3302563,
              "author_name": "Luca",
              "author_url": "",
              "post_date": "2025-10-16T03:52:57.343000",
              "content": "<p>Thanks for the detailed reply!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3302104,
      "author_name": "Tom",
      "author_url": "",
      "post_date": "2025-10-15T03:49:35.173000",
      "content": "<p><a href=\"https://www.kaggle.com/tomoon33\" target=\"_blank\">@tomoon33</a>  It seems nnUnet can have better representation than YOLO. I've tried my BYU solution + <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> previous RSNA solution but result is not expected like yours. I really like your development of ROI classification stage, Great work. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3302147,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2025-10-15T06:56:53.390000",
          "content": "<p>My intuition that excluding bad data probably plays a big role. However, It needs professional knowledge.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3302102": "I thank RSNA, the organizers, the contributing radiologists and institutions, and the Kaggle team for hosting this impactful challenge. This write-up summarizes my 1st place solution. The core of my approach is a robust, coarse-to-fine pipeline that uses vessel segmentation to guide a region-of-interest (ROI) based classifier, producing location-aware predictions.\n\n## Solution Overview\n\n- **High-Level Pipeline**\n  1.  **Preprocessing:** Convert and standardize DICOM series into NIfTI volumes.\n  2.  **Vessel Segmentation & ROI Extraction:** Use a coarse-to-fine nnU-Net approach. A fast, low-resolution model first finds a candidate region. This allows the more detailed, high-resolution models to focus only on this ROI, which improves both accuracy and speed.\n  3.  **ROI Classification:** A 3D classification model, using the detailed vessel masks as input, predicts the probabilities for the 13 anatomical locations and the overall aneurysm presence.\n\n- **Key Design Principles**\n    - **Coarse-to-Fine Efficiency:** A fast, low-resolution model first scans the entire volume to find a candidate region. This allows the more detailed, high-resolution models to focus only on this ROI. This improves both accuracy and speed.\n    - **Using Segmentation as a Structural Guide:** By providing the classification model with an explicit vessel mask, I give it a detailed map of the vessel structure. This helps the model to focus on the vessels, where aneurysms are located.\n\n## Data Preparation\n- I excluded approximately 60 series due to data quality issues, such as orientation anomalies, corrupted DICOM files, and implausible slice spacing.\n- I used multilabel-stratified 5-fold cross-validation. \n\n## Pipeline\n\n### 1. Preprocessing\n- **Filter Slices:** Within each series, I retained only the images matching the majority `Rows × Cols` and `PixelSpacing` configuration and filtered out slices with outlier interslice spacing to ensure consistency.\n- **Convert DICOM to NIfTI:** I used `dcm2niix` for conversion. If it failed, I first ran `gdcmconv --raw` as a fallback before retrying the conversion [1].\n- **Standardize Orientation:** All volumes were reoriented to a consistent anatomical orientation using nnU-Net’s `SimpleITKIOWithReorient`.\n- **Normalize Intensity:** I applied nnU-Net’s standard per-volume z-score normalization to standardize image intensities.\n\n### 2. nnU-Net Segmentation + ROI Extraction (Coarse-to-Fine)\nThis stage uses a sequence of three nnU-Net v2 [2] models (`nnUNetResEncUNetMPlans`, `3d_fullres` configuration) to first locate a coarse ROI and then produce detailed vessel segmentations within it.\n\n- **Model 1: Coarse Vessel Localization**\n  - **Spacing:** (1.0, 1.0, 1.0) mm\n  - **Classes:** 3 vessel groups (Posterior+Basilar / MCA / Other)\n  - **Loss:** Dice + Cross-Entropy\n  - **Purpose:** To perform a fast, low-resolution scan to efficiently find a single ROI candidate for the high-resolution models. This model also supports an optional orientation correction step described later.\n\n- **Model 2: Fine Segmentation (Balanced)**\n  - **Spacing:** (0.80, 0.45, 0.44) mm\n  - **Loss:** Dice + Cross-Entropy + SkeletonRecall (weight=1) [3]\n  - **Purpose:** To generate a precise vessel segmentation. The SkeletonRecall loss improves the connectivity of thin vessels, which standard losses might miss.\n\n- **Model 3: Fine Segmentation (Recall-Focused)**\n  - **Spacing:** (0.80, 0.45, 0.44) mm\n  - **Loss:** Tversky + Cross-Entropy + SkeletonRecall (weight=3)\n  - **Purpose:** To complement Model 2 by prioritizing recall, making it more sensitive to detecting hard-to-find vessel segments.\n\n- **Augmentation Strategy**\n  - I disabled left–right mirroring for the fine models (2 and 3) to preserve anatomical asymmetry.\n  - I used stronger intensity and geometric augmentations [4].\n  - I added low-resolution simulation transforms to make the models robust to scans with thick slices.\n\n- **Inference Process (Two-Stage)**\n  - **Stage 1 (Coarse Scan):** I run Model 1 with a sliding window (overlap=0.2) and binarize the output to get a foreground mask. I then apply DBSCAN clustering to the mask to remove scattered false positives. The centroid of the largest cluster is used to crop a fixed-size ROI (140×140×140 mm).\n  - **Stage 2 (Fine Inference):** I run Models 2 and 3 with a higher overlap (0.3) only within the coarse ROI. The vessel segmentation from Model 2 is used to compute a tight bounding box, which is then re-cropped with margins to create the final ROI for the classifier.\n\nThe detailed segmentations from this stage are passed to the classification model as masks.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F9edf9e0846e14151f33a1ece358473a9%2Fflow_segmentation.png?generation=1760833847347432&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F82fa14c88dfc1dc57448f40bca3ee07d%2Fseg_losses.png?generation=1760833856913829&alt=media)\n\n### 3. ROI Classification (13 locations + Aneurysm Present)\nOnce the ROI and vessel masks are prepared, a 3D classification model predicts the final probabilities.\n\n- **Model Architecture**\n  - **Input Size:** The model takes ROI volumes of size 128 × 256 × 256 voxels.\n  - **Backbone:** The core of the model is an nnU-Net pre-trained for the vessel segmentation task. This approach was more accurate and faster to train than standard 2.5D or 3D timm backbones.\n      - **Decoder Simplification:** To improve efficiency, I simplified the decoder by removing its final, computationally-heavy block. This had no negative impact on performance.\n  - **Auxiliary Detection Task:** To help the model learn specific features of aneurysms, I added an auxiliary task that uses the decoder features to reconstruct a small binary sphere (5-pixel radius) at the location of each annotated aneurysm.\n  - **Location-Specific Prediction Head:** To predict the 13 location-specific probabilities, I use the following process:\n      1.  **Per-Location Feature Pooling:** A \"Vessel Region-Masked Pooling\" layer uses the vessel masks to extract feature vectors corresponding to each of the 13 anatomical locations from the decoder's feature maps. (*Note: In practice, I apply this pooling using masks from both fine segmentation models and concatenate the results for more complete features.*)\n      2.  **Feature Fusion:** These 13 feature vectors are combined with a global feature vector from the encoder (via Global Average Pooling).\n      3.  **Inter-Location Modeling:** The combined features are fed into a \"Location-Aware Transformer\" to model relationships between different vessel locations.\n      4.  **Classification:** Finally, an MLP head predicts the probability for each of the 13 locations.\n  - **\"Aneurysm Present\" Prediction Head:** For the overall presence prediction, I pool features over the entire vessel structure (a union of all vessel masks) and combine them with the encoder's global features. This aggregated feature vector is passed to a separate MLP head.\n  - **Output Design:** I treated each of the 14 labels as an independent binary classification problem. This design helps with the severe class imbalance, as positive cases for any single location are very rare.\n\n- **Training Details**\n  - **Loss Functions:**\n      - **13 Locations:** `BCEWithLogitsLoss`.\n      - **Aneurysm Present:** `BCEWithLogitsLoss`.\n      - **Auxiliary Sphere Segmentation:** A combination of Balanced BCE [5] and Focal-Tversky++ loss [6, 7]. This combination works well for highly sparse targets and helps prevent the over-confidence that can occur with Dice-like losses.\n      - **Loss Weights:** The final loss was a weighted sum of the three components. I set the weights to 0.1 for the 13 location losses, 0.05 for the Aneurysm Present loss, and 1.0 for the auxiliary sphere segmentation loss. The main goal was to prioritize learning the precise location of aneurysms, so the sphere segmentation task had the highest weight. I found that higher weights on the classification losses led to overfitting, so this balance was important.\n  - **Data Augmentation**\n      - **Intensity Transforms:** Gaussian noise, Gaussian smoothing, intensity shift/scale, contrast adjustment, Gaussian sharpening, and intensity inversion.\n      - **Geometric Transforms:** Random flips (z, y, x axes), small rotations (±10°), scaling/shearing (±10%), mild grid distortions, and a simulated low-resolution transform.\n  - **Optimizer and Schedule:** I used the AdamW optimizer with a learning rate of 1e-4 and an effective batch size of 8 (achieved with gradient accumulation). A standard cosine annealing schedule with a warmup period was used.\n  - **EMA Weights:** I used the Exponential Moving Average (EMA) of the model weights for inference.\n\n- **Inference**\n  - **Ensembling:** The final predictions are an average of the models from 4 of the 5 cross-validation folds.\n  - **Test-Time Augmentation (TTA):** I averaged the predictions from the original volume and a left-right flipped version of it.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F98e18ca93230a3f26bf2181021ce481a%2Fmodel_overview.png?generation=1760490733871010&alt=media)\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F187c3d0327547a692d8e922dc5f1d5cd%2Fvessel_pooling.png?generation=1760490747217372&alt=media\" width=\"80%\">\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2F0dd935cbc51a7d50bc69779c8f12d918%2Ftransformer.png?generation=1760490761936793&alt=media\" width=\"60%\">\n\n### 4. Other Details\n- **Fail-safe Mechanism:** If an anomaly occurred during the segmentation or ROI extraction steps, the pipeline would not attempt a prediction. Instead, it fell back to a set of pre-determined probabilities. For each class, this probability was the mean of the out-of-fold predictions from my cross-validation set.\n- **Optional Orientation Correction:** I implemented a method to fix misoriented scans. It analyzed the spatial arrangement of the three vessel groups from the coarse segmentation. By comparing the relative positions of these groups to their expected anatomical locations, it estimated the correct axis permutation. This worked perfectly on the training set, fixing all identified orientation issues. However, it had no measurable effect on the leaderboard score, likely because the test set did not contain such orientation errors.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F15217057%2Ff954825e84986adf4e7b537f9160b1b0%2F3classes_for_ori_corr.png?generation=1760833867211206&alt=media\" width=\"40%\">\n\n## Processing Time\n\n- **Training**\n  - The training times below were measured on a single NVIDIA RTX 4090.\n  - The 96×192×192 input size was mainly used for faster experimentation and tuning.\n\n| Model | Input size | Epochs | Training Time |\n|---|---|---|---|\n| nnU-Net (Model 2) | 64×192×192 | 1000 | 14 h |\n| ROI classifier | 96×192×192 | 30 | 6 h |\n| ROI classifier | 128×256×256 | 30 | 12 h |\n\n\n- **Inference**\n  - Inference times were measured on the Kaggle notebook using two T4 GPUs.\n  - The times reported are an average over 100 samples.\n\n| Step | Time per Series |\n|---|---|\n| Preprocessing | 4.03 ± 3.78 s |\n| Vessel segmentation | 10.65 ± 4.10 s |\n| ROI classification | 3.33 ± 0.09 s |\n \n \n## Ablation Study\nTo validate the effectiveness of the key components in my ROI classification model, I conducted an ablation study. Note that this was a simplified evaluation; hyperparameters such as the number of epochs and learning rate were not re-tuned for each experiment. For efficiency, these experiments were run on folds 0, 1, and 2, using a reduced input size of 96×192×192.\n\n| Model | Input Size | AUC (Aneurysm Present) | AUC (13 Locations) | Score |\n|---|---|---|---|---|\n| Final model (full resolution) | 128×256×256 | 0.915 | 0.916 | 0.916 |\n| Final model | 96×192×192 | 0.907 | 0.898 | 0.902 |\n| Without Location-Aware Transformer | 96×192×192 | 0.899 | 0.894 | 0.896 |\n| Using Dice instead of FocalTversky++ | 96×192×192 | 0.902 | 0.896 | 0.899 |\n| Setting all loss weights to 1.0 | 96×192×192 | 0.890 | 0.877 | 0.884 |\n| Without backbone pretraining | 96×192×192 | 0.777 | 0.811 | 0.794 |\n| Without segmentation model 3 | 96×192×192 | 0.899 | 0.883 | 0.891 |\n| Without auxiliary segmentation loss | 96×192×192 | 0.880 | 0.871 | 0.876 |\n\nThe results show several key points:\n- Pretraining the backbone on the vessel segmentation task was the most important factor. It greatly improved the score and helped the model train much faster. This was very helpful for running many experiments.\n- The Location-Aware Transformer and the FocalTversky++ loss seemed to help at first, but their final contribution to the score was small. This is likely because other improvements and tuning had a larger overall effect.\n\n\n## Design Journey\nMy final design was the result of several iterations:\n\n1.  I initially tried a single, end-to-end 3D classifier that had auxiliary heads for vessel segmentation and aneurysm localization. While it could detect the presence of an aneurysm, it failed to predict the 13 specific locations accurately.\n2.  I then observed that a standard nnU-Net for vessel segmentation trained easily and generalized well across all modalities. This led me to change to a two-stage, vessel-first pipeline, where the segmentation acts as a strong guide for the subsequent classification task.\n3.  I also experimented with a simpler model that took a 2-channel input: the image volume concatenated with a single binary vessel mask. To get a prediction for a specific label, I would feed the model the corresponding mask (e.g., the mask for one location, or the union mask for \"Aneurysm Present\"). Although the model itself was simple, this approach required running 14 separate forward passes to get all predictions for a single patient series, which was too computationally expensive.\n4.  This led to my final approach using a single backbone pass with the region-masked pooling, which provided a good balance of accuracy and computational efficiency.\n\n\n## References\n[1] https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/598083\n[2] Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J., & Maier-Hein, K. H. (2021). nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods, 18(2), 203-211.\n[3] Kirchhoff, Yannick, et al. \"Skeleton recall loss for connectivity conserving and resource efficient segmentation of thin tubular structures.\" European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024.\n[4] https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/mic-dkfz-2nd-place-solution-3d-nnu-net-blob-regres\n[5] https://www.kaggle.com/competitions/czii-cryo-et-object-identification/writeups/yu4u-tattaka-4th-place-solution-source-codes-submi\n[6] Yeung, Michael, et al. \"Calibrating the dice loss to handle neural network overconfidence for biomedical image segmentation.\" Journal of Digital Imaging 36.2 (2023): 739-752.\n[7] https://www.kaggle.com/competitions/czii-cryo-et-object-identification/writeups/tomoon33-6th-place-solution\n\n## Code Release\n- Training code: https://github.com/uchiyama33/rsna2025_1st_place\n- Submission notebook: https://www.kaggle.com/code/tomoon33/rsna2025-submission-1st-place",
    "3302258": "Congrats!! Glad Skeleton Recall Loss could provide some benefit here as well 💯\n",
    "3344478": "Hey, could you share the W&B logs and the train/validation loss curves?",
    "3303796": "I've just updated my solution write-up with a more detailed version. The code will be released shortly.",
    "3302359": "Congratulations on taking 1st place solo!!👏\n\nHow was this architecture derived? It's nothing short of spectacular.\nI'd love to hear about the process that led to this model.",
    "3302107": "Congratulations on the win! The auxiliary tasks using nnunet are really good. I wonder, what was the Dice score of your nnunet for vessel segmentation? I managed to get one close to 0.78. I could not do much with the ROIs afterwards, because I did not have enough compute to generate entire train data segmentations (disk space issue). I only tried a PoC with 128x128x128 volumes using another segmentation model, but nothing worked. ",
    "3302525": "very perfect solution. would you like to share more detail about your solution. or share you inference and training code. thanks very much",
    "3346908": "- In ./classification/models, I saw a class called AneurysmModel. Is it used for segmentation?\n- If I want to replace Stage 2 classification with segmentation, is there any existing configuration to train Stage 2 as a segmentation model?\"",
    "3304670": "Thank You so much for sharing @tomoon33 ",
    "3303903": "Great solution....\nCould you please explain how vessel segmentation is performed since the dataset only contains the vessel location of where aneurysm is present? What dataset did you use to train the coarse nn-Unet",
    "3303642": "Congratulations on the win! ",
    "3303580": "Congratulations on your win, amazing work!\n\nYour writeup diagrams are wonderful. May I ask which tool you used to create them?",
    "3303512": "Congratulations on 1st place. Your pipeline looks very robust. I’m particularly curious about how you managed inference speed with the coarse-to-fine approach did it add much latency, or was it still efficient at scale? Would love to try something similar in my experiments.",
    "3303471": "Congratulations! Could you explain how you developed this model from nnUNet to the current complete version? What inspired you?",
    "3303443": "Really love this solution! Congrats",
    "3302875": "Super congratulations! \nVery well deserved, beautiful work. \n\nM",
    "3302797": "congrats and thanks for the nice summary. would it be possible to make the cleaned dataset publicly available?",
    "3302777": "It's really good ",
    "3302152": "Congratulations! Nice usage of nnU-Net 😍",
    "3302134": "Congratulations on the win! Brilliant design! I was wondering, how much did the location-aware transformers contribute to the results? Also, how were the weights in the loss function allocated?",
    "3302132": "Congratulations @tomoon33 \nWould you mind sharing why spacing = (1.0, 1.0, 1.0) and  (0.80, 0.45, 0.44) were chosen? It is by experiment or what?\nThanks.",
    "3302113": "Congratulations on your win—well deserved! Your solution wrap-up was very clear.\nI also have a couple of questions: Did you encounter any aneurysms located outside the segmentation ROI, and if so, how did you address them? I'm also curious about your Dice scores for the Right and Left PCA, which are quite small structures and appear to have significant missing labels in the provided segmentation data.",
    "3302104": "@tomoon33  It seems nnUnet can have better representation than YOLO. I've tried my BYU solution + @ren4yu previous RSNA solution but result is not expected like yours. I really like your development of ROI classification stage, Great work. "
  }
}