{
  "id": 611867,
  "title": "2nd Place Solution",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/611867",
  "author_name": "Pengcheng Shi",
  "post_date": "2025-10-15T08:41:30.042000",
  "votes": 41,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Thanks to <a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a>, RSNA, and Kaggle for organizing this Intracranial Aneurysm Detection competition! A special shoutout to the insightful discussions (especially <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a>'s detailed post on multi-frame DICOM analysis) that helped us gain a deeper understanding of the competition data.</p>\n<p>A detailed version of this work is available in the paper at: <a href=\"https://arxiv.org/abs/2606.26706\" target=\"_blank\">https://arxiv.org/abs/2606.26706</a></p>\n<h2>Overview</h2>\n<p>Our approach focused on simplicity and generality to handle the diverse data in this classification-focused task. Key elements:</p>\n<ul>\n<li><p><strong>Stage 1</strong>: Fast 2D tri-axial ROI extraction using an nnU-Net 2D segmentation model to crop binary vascular regions efficiently, validated on all training cases locally.</p></li>\n<li><p><strong>Stage 2</strong>: 3D multi-task learning (segmentation of vessels and aneurysms + classification) based on nnU-Net, with enhancements like cross-attention pooling, modality 4-class heads, and targeted oversampling for rare classes. All data resized to uniform 224x224x224; heavy TTA (8x) including left-right flips with label swapping.</p></li>\n</ul>\n<p>Built entirely on the open-source nnU-Net framework—big thanks to <a href=\"https://www.kaggle.com/fabianisensee\" target=\"_blank\">@fabianisensee</a> and team for its robust 3D segmentation baseline.</p>\n<p>Inference optimized for speed (encoder + classification head only), running on 2x T4 GPUs in ~9 hours for conservative ensembles.</p>\n<h2>Who We Are</h2>\n<p>We're a mix of algorithm engineers and PhD researchers: Pengcheng Shi, Yan Lu, and Jiawei Chen from Medical Image Insights in Shanghai. Kaiyuan Yang and Houjing Huang from UZH in Zurich. Kaiyuan Yang, Houjing Huang, and Pengcheng Shi are also organizers of the MICCAI TopCoW (<a href=\"https://topcow24.grand-challenge.org/\" target=\"_blank\">https://topcow24.grand-challenge.org/</a>) / TopBrain (<a href=\"https://topbrain2025.grand-challenge.org/\" target=\"_blank\">https://topbrain2025.grand-challenge.org/</a>) challenges that benchmarked the segmentation of Circle of Willis (CoW) and whole brain vessel anatomy. Most of us are new to Kaggle (our background is in MICCAI events), and we noticed differences: Kaggle emphasizes data wrangling, efficiency under resource limits, and flat metrics (pure classification here, no localization/segmentation scores). In the TopCoW summary pre-print, we have previously explored the potential of TopCoW segmentation model at locating aneurysm with the CoW anatomy, which inspired many ideas used in our current solution.</p>\n<h2>Data Handling Challenges</h2>\n<p>Kaggle's data diversity (spacing, modalities, multi-frame DICOMs) required heavy preprocessing. Multi-frame issues were tricky—we had limited DICOM experience, and test sets had deleted fields, breaking dicom2nifti. We switched to pydicom, mapping spacing by slice shapes (e.g., &lt;45 slices → 5mm), and trained a T2-specific orientation classifier for corrections. This ate up time but ensured full test coverage. No try-except fallbacks to 0.5 predictions in final inference—maximized robustness.</p>\n<p>We used flipping TTA and multi-fold model ensembling to increase robustness of the model. Promising results showed that increasing in the public leaderboard led to consistent improvement in the private leaderboard.</p>\n<h2>Stage 1: 2D Tri-Axial ROI Extraction</h2>\n<p>To crop binary vascular regions efficiently:</p>\n<ul>\n<li>Sample 3 slices per axis (1/4, 1/2, 3/4 positions) from iterative vascular ROI data → 9 slices total.</li>\n<li>Train nnU-Net 2D config on sliding windows for cropped segmentation.</li>\n<li>Inference: Merging sliding window patches into the batch dimension enhances both speed and robustness.</li>\n<li>Local tests: Handled 100% of training cases correctly.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4254412%2Fc7dd148e42375ae82883854f426db4ff%2Fstage1.png?generation=1760691510047622&amp;alt=media\" alt=\"stage1\">\nThis stage was key for quick, universal ROI on varied data.</li>\n</ul>\n<h2>Stage 2: 3D Multi-Task Learning</h2>\n<p>Unified all spacing/modalities by resizing to 224x224x224.</p>\n<p><strong>Architecture</strong>: nnU-Net 3D with multi-task heads (vessel/aneurysm seg + classification).</p>\n<p><strong>Enhancements</strong>:</p>\n<ul>\n<li>Cross-attention pooling surpassed average pooling in convergence speed.</li>\n<li>Modality classification head enhanced training stability.</li>\n<li>Inference: Only encoder and class head forwarded to reduce latency.</li>\n<li>Data aug: Left-right flips with swapped labels and masks for classification and segmentation, effectively doubling dataset size.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4254412%2Ffbfc6b089e512c99a842a91afe975f31%2Fvessel_flip.png?generation=1760637300043403&amp;alt=media\" alt=\"vessel_flip\">\n<em>Left: Original vascular anatomy segmentation mask | Right: Augmented mask after horizontal flipping with swapped labels</em></li>\n<li>Aneurysm masks were iteratively generated via model inference and manual refinement based on aneurysm center points.</li>\n<li>13-class vascular segmentation refined aneurysm masks using center-point distance heatmaps.</li>\n<li>Rare classes were oversampled and assigned higher cross-entropy weights for both classification and segmentation.</li>\n<li>Vessel and aneurysm segmentation stabilized classification convergence.</li>\n<li>TTA: 8x (flips; swap left/right labels on outputs).</li>\n<li>Final submissions: Due to platform constraints, we cut model count and ran a conservative two-fold ensemble on 2×T4 GPUs to prevent timeouts (runs were ~9 hours).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4254412%2F0940c203faf2ab7400461b7f45fa746e%2Fstage2.png?generation=1764323763634081&amp;alt=media\" alt=\"stage2\"></li>\n</ul>\n<h2>Training Details</h2>\n<p>Based on nnU-Net defaults, trained from scratch.</p>\n<p><strong>Loss</strong>: CE with heatmap weighting for aneurysm centers.</p>\n<p><strong>Augmentations</strong>: Standard nnU-Net (rotations, scaling, noise) + custom left-right flips and label swaps.</p>\n<p><strong>External data</strong>: </p>\n<ul>\n<li>TopCoW Training Data and its External Testsets: <a href=\"https://zenodo.org/records/15692630\" target=\"_blank\">https://zenodo.org/records/15692630</a> (Note: The LargeIA dataset was excluded from training.)</li>\n<li>TopBrain annotations: <a href=\"https://zenodo.org/records/16878417\" target=\"_blank\">https://zenodo.org/records/16878417</a></li>\n</ul>\n<p><strong>Training data annotation and correction</strong>: </p>\n<ul>\n<li>Aneurysm seg mask: Center position provided in the challenge data csv combined with light manual annotation to iteratively annotate and train our own aneurysm segmentation model.</li>\n<li>The provided rsna aneurysm annotations contain a few cases that mixed up left vs right side aneurysms, supra- vs infra-clinoid ICA aneurysms,and other position label mix-ups, which were manually corrected if identified.</li>\n<li>Vessel seg mask customized by merging provided cow-seg masks with TopCoW and TopBrain annotations; a few hard cases, especially of T2 and T1-post (some T1-post even have artery flow void), were manually corrected.</li>\n</ul>\n<h2>Results and Insights</h2>\n<h3>Ablation Study</h3>\n<table>\n<thead>\n<tr>\n<th>Experiment Configuration</th>\n<th>Public Score</th>\n<th>Private Score</th>\n<th>Notes</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline</td>\n<td>0.84407</td>\n<td>0.81268</td>\n<td>• First version segmentation annotation data<br>• Epoch 100, Fold 0<br>• resize 224×224×224<br>• cls_modality</td>\n</tr>\n<tr>\n<td>Baseline + TTA 4x</td>\n<td>0.87056</td>\n<td>0.82508</td>\n<td>• First version segmentation annotation data<br>• Epoch 100, Fold 0<br>• resize 224×224×224<br>• cls_modality<br>• Test-Time Augmentation 4x</td>\n</tr>\n<tr>\n<td>Improved Data + TTA 4x</td>\n<td>0.88832</td>\n<td>0.85718</td>\n<td>• Improved segmentation annotation data<br>• Epoch 100, Fold 0<br>• resize 224×224×224<br>• cls_modality<br>• Test-Time Augmentation 4x</td>\n</tr>\n<tr>\n<td>Improved Data + TTA 8x</td>\n<td>0.89805</td>\n<td>0.86228</td>\n<td>• Improved segmentation annotation data<br>• Left-right data augmentation with label swap<br>• Epoch 100, Fold 0<br>• resize 224×224×224<br>• cls_modality<br>• Test-Time Augmentation 8x</td>\n</tr>\n<tr>\n<td>Improved Data + TTA 8x + Ensemble</td>\n<td>0.90035</td>\n<td>0.86727</td>\n<td>• Improved segmentation annotation data<br>• Left-right data augmentation with label swap<br>• Epoch 100, Fold 0 + Epoch 250, Fold 1 (Ensemble)<br>• resize 224×224×224<br>• cls_modality<br>• Test-Time Augmentation 8x</td>\n</tr>\n</tbody>\n</table>\n<h3>Key Technical Insights</h3>\n<ul>\n<li>ROI Extraction: Our Stage 1, 2D tri-axial ROI extraction significantly improved inference speed.</li>\n<li>Vessel Segmentation: Incorporating vessel segmentation clearly enhanced aneurysm segmentation and classification performance.</li>\n<li>Noise Handling: Our noise-handling techniques performed more effectively on the public dataset.</li>\n<li>Dataset Shift: Removing the 0.5 prediction fallback revealed a consistent 3–4% performance gap between the public and private leaderboards. This suggests a differing distribution of abnormal cases between the two test sets.</li>\n</ul>\n<h3>Performance Evolution</h3>\n<ul>\n<li>A consistent performance improvement was observed across all experiment iterations.</li>\n<li>Top Result: \"Improved Data + TTA 8x + Ensemble\" scored <strong>0.90035</strong> (Public LB, 1st) and <strong>0.86732</strong> (Private LB, 2nd).</li>\n</ul>\n<h2>Code Availability</h2>\n<ul>\n<li>Training and inference code: <a href=\"https://github.com/PengchengShi1220/RSNA2025_Intracranial-Aneurysm-Detection\" target=\"_blank\">https://github.com/PengchengShi1220/RSNA2025_Intracranial-Aneurysm-Detection</a></li>\n<li>Inference demo: <a href=\"https://www.kaggle.com/code/pengchengshi/bravecowcow-2nd-place-inference-demo\" target=\"_blank\">https://www.kaggle.com/code/pengchengshi/bravecowcow-2nd-place-inference-demo</a></li>\n<li>Final submission inference: <a href=\"https://www.kaggle.com/code/pengchengshi/bravecowcow-2nd-place-inference-final-submission\" target=\"_blank\">https://www.kaggle.com/code/pengchengshi/bravecowcow-2nd-place-inference-final-submission</a></li>\n<li>Stage 1 model checkpoint: <a href=\"https://www.kaggle.com/models/pengchengshi/dataset180_2d_vessel_box_seg_stable\" target=\"_blank\">https://www.kaggle.com/models/pengchengshi/dataset180_2d_vessel_box_seg_stable</a></li>\n<li>Stage 2 model checkpoint: <a href=\"https://www.kaggle.com/models/pengchengshi/dataset660_26classes_resize224_4661\" target=\"_blank\">https://www.kaggle.com/models/pengchengshi/dataset660_26classes_resize224_4661</a></li>\n</ul>\n<h2>Data Availability</h2>\n<ul>\n<li>Segmentation labels: <a href=\"https://huggingface.co/datasets/spc819/rsna2025-aneurysm-26class-seg\" target=\"_blank\">https://huggingface.co/datasets/spc819/rsna2025-aneurysm-26class-seg</a></li>\n</ul>\n<h2>3D Slicer Plugin</h2>\n<ul>\n<li>Source code: <a href=\"https://github.com/murong-xu/SlicerBraveCowCow\" target=\"_blank\">https://github.com/murong-xu/SlicerBraveCowCow</a></li>\n<li>Inference backend: <a href=\"https://github.com/huanghoujing/bravecowcow_inference_docker\" target=\"_blank\">https://github.com/huanghoujing/bravecowcow_inference_docker</a></li>\n</ul>\n<h2>Acknowledgements</h2>\n<p>Thanks to Medical Image Insights and UZH for compute support, Bjoern Menze and the Helmut Horten Foundation for funding support. We are grateful to RSNA/Kaggle hosts, nnU-Net devs, and forum contributors.</p>\n<h2>Citation</h2>\n<p>Please cite our work if it is helpful for your research:</p>\n<pre><code>@misc{shi2026intracranialaneurysmclassificationsegmentation,\n      title={Intracranial Aneurysm Classification and Segmentation via Tri-Axial ROI and Multi-Task Learning},\n      author={Pengcheng Shi and Kaiyuan Yang and Houjing Huang and Jiawei Chen and Yan Lu and Jiaqi Liu and Murong Xu and Minghui Zhang and Bjoern Menze and Xinglin Zhang},\n      year={2026},\n      eprint={2606.26706},\n      archivePrefix={arXiv},\n      primaryClass={cs.CV},\n      url={https://arxiv.org/abs/2606.26706},\n}\n</code></pre>\n<h2>References</h2>\n<ul>\n<li>Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J., &amp; Maier-Hein, K. H. (2021). nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nature Methods, 18(2), 203-211.</li>\n<li>Yang, Musio, Ma, et al. \"Benchmarking the CoW with the TopCoW Challenge: Topology-Aware Anatomical Segmentation of the Circle of Willis for CTA and MRA.\" arXiv (2025): arXiv-2312</li>\n</ul>",
  "messages": [
    {
      "id": 3302177,
      "postDate": "2025-10-15T08:41:30.043Z",
      "content": "<p>Thanks to <a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a>, RSNA, and Kaggle for organizing this Intracranial Aneurysm Detection competition! A special shoutout to the insightful discussions (especially <a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a>'s detailed post on multi-frame DICOM analysis) that helped us gain a deeper understanding of the competition data.</p>\n<p>A detailed version of this work is available in the paper at: <a href=\"https://arxiv.org/abs/2606.26706\" target=\"_blank\">https://arxiv.org/abs/2606.26706</a></p>\n<h2>Overview</h2>\n<p>Our approach focused on simplicity and generality to handle the diverse data in this classification-focused task. Key elements:</p>\n<ul>\n<li><p><strong>Stage 1</strong>: Fast 2D tri-axial ROI extraction using an nnU-Net 2D segmentation model to crop binary vascular regions efficiently, validated on all training cases locally.</p></li>\n<li><p><strong>Stage 2</strong>: 3D multi-task learning (segmentation of vessels and aneurysms + classification) based on nnU-Net, with enhancements like cross-attention pooling, modality 4-class heads, and targeted oversampling for rare classes. All data resized to uniform 224x224x224; heavy TTA (8x) including left-right flips with label swapping.</p></li>\n</ul>\n<p>Built entirely on the open-source nnU-Net framework—big thanks to <a href=\"https://www.kaggle.com/fabianisensee\" target=\"_blank\">@fabianisensee</a> and team for its robust 3D segmentation baseline.</p>\n<p>Inference optimized for speed (encoder + classification head only), running on 2x T4 GPUs in ~9 hours for conservative ensembles.</p>\n<h2>Who We Are</h2>\n<p>We're a mix of algorithm engineers and PhD researchers: Pengcheng Shi, Yan Lu, and Jiawei Chen from Medical Image Insights in Shanghai. Kaiyuan Yang and Houjing Huang from UZH in Zurich. Kaiyuan Yang, Houjing Huang, and Pengcheng Shi are also organizers of the MICCAI TopCoW (<a href=\"https://topcow24.grand-challenge.org/\" target=\"_blank\">https://topcow24.grand-challenge.org/</a>) / TopBrain (<a href=\"https://topbrain2025.grand-challenge.org/\" target=\"_blank\">https://topbrain2025.grand-challenge.org/</a>) challenges that benchmarked the segmentation of Circle of Willis (CoW) and whole brain vessel anatomy. Most of us are new to Kaggle (our background is in MICCAI events), and we noticed differences: Kaggle emphasizes data wrangling, efficiency under resource limits, and flat metrics (pure classification here, no localization/segmentation scores). In the TopCoW summary pre-print, we have previously explored the potential of TopCoW segmentation model at locating aneurysm with the CoW anatomy, which inspired many ideas used in our current solution.</p>\n<h2>Data Handling Challenges</h2>\n<p>Kaggle's data diversity (spacing, modalities, multi-frame DICOMs) required heavy preprocessing. Multi-frame issues were tricky—we had limited DICOM experience, and test sets had deleted fields, breaking dicom2nifti. We switched to pydicom, mapping spacing by slice shapes (e.g., &lt;45 slices → 5mm), and trained a T2-specific orientation classifier for corrections. This ate up time but ensured full test coverage. No try-except fallbacks to 0.5 predictions in final inference—maximized robustness.</p>\n<p>We used flipping TTA and multi-fold model ensembling to increase robustness of the model. Promising results showed that increasing in the public leaderboard led to consistent improvement in the private leaderboard.</p>\n<h2>Stage 1: 2D Tri-Axial ROI Extraction</h2>\n<p>To crop binary vascular regions efficiently:</p>\n<ul>\n<li>Sample 3 slices per axis (1/4, 1/2, 3/4 positions) from iterative vascular ROI data → 9 slices total.</li>\n<li>Train nnU-Net 2D config on sliding windows for cropped segmentation.</li>\n<li>Inference: Merging sliding window patches into the batch dimension enhances both speed and robustness.</li>\n<li>Local tests: Handled 100% of training cases correctly.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4254412%2Fc7dd148e42375ae82883854f426db4ff%2Fstage1.png?generation=1760691510047622&amp;alt=media\" alt=\"stage1\">\nThis stage was key for quick, universal ROI on varied data.</li>\n</ul>\n<h2>Stage 2: 3D Multi-Task Learning</h2>\n<p>Unified all spacing/modalities by resizing to 224x224x224.</p>\n<p><strong>Architecture</strong>: nnU-Net 3D with multi-task heads (vessel/aneurysm seg + classification).</p>\n<p><strong>Enhancements</strong>:</p>\n<ul>\n<li>Cross-attention pooling surpassed average pooling in convergence speed.</li>\n<li>Modality classification head enhanced training stability.</li>\n<li>Inference: Only encoder and class head forwarded to reduce latency.</li>\n<li>Data aug: Left-right flips with swapped labels and masks for classification and segmentation, effectively doubling dataset size.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4254412%2Ffbfc6b089e512c99a842a91afe975f31%2Fvessel_flip.png?generation=1760637300043403&amp;alt=media\" alt=\"vessel_flip\">\n<em>Left: Original vascular anatomy segmentation mask | Right: Augmented mask after horizontal flipping with swapped labels</em></li>\n<li>Aneurysm masks were iteratively generated via model inference and manual refinement based on aneurysm center points.</li>\n<li>13-class vascular segmentation refined aneurysm masks using center-point distance heatmaps.</li>\n<li>Rare classes were oversampled and assigned higher cross-entropy weights for both classification and segmentation.</li>\n<li>Vessel and aneurysm segmentation stabilized classification convergence.</li>\n<li>TTA: 8x (flips; swap left/right labels on outputs).</li>\n<li>Final submissions: Due to platform constraints, we cut model count and ran a conservative two-fold ensemble on 2×T4 GPUs to prevent timeouts (runs were ~9 hours).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4254412%2F0940c203faf2ab7400461b7f45fa746e%2Fstage2.png?generation=1764323763634081&amp;alt=media\" alt=\"stage2\"></li>\n</ul>\n<h2>Training Details</h2>\n<p>Based on nnU-Net defaults, trained from scratch.</p>\n<p><strong>Loss</strong>: CE with heatmap weighting for aneurysm centers.</p>\n<p><strong>Augmentations</strong>: Standard nnU-Net (rotations, scaling, noise) + custom left-right flips and label swaps.</p>\n<p><strong>External data</strong>: </p>\n<ul>\n<li>TopCoW Training Data and its External Testsets: <a href=\"https://zenodo.org/records/15692630\" target=\"_blank\">https://zenodo.org/records/15692630</a> (Note: The LargeIA dataset was excluded from training.)</li>\n<li>TopBrain annotations: <a href=\"https://zenodo.org/records/16878417\" target=\"_blank\">https://zenodo.org/records/16878417</a></li>\n</ul>\n<p><strong>Training data annotation and correction</strong>: </p>\n<ul>\n<li>Aneurysm seg mask: Center position provided in the challenge data csv combined with light manual annotation to iteratively annotate and train our own aneurysm segmentation model.</li>\n<li>The provided rsna aneurysm annotations contain a few cases that mixed up left vs right side aneurysms, supra- vs infra-clinoid ICA aneurysms,and other position label mix-ups, which were manually corrected if identified.</li>\n<li>Vessel seg mask customized by merging provided cow-seg masks with TopCoW and TopBrain annotations; a few hard cases, especially of T2 and T1-post (some T1-post even have artery flow void), were manually corrected.</li>\n</ul>\n<h2>Results and Insights</h2>\n<h3>Ablation Study</h3>\n<table>\n<thead>\n<tr>\n<th>Experiment Configuration</th>\n<th>Public Score</th>\n<th>Private Score</th>\n<th>Notes</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Baseline</td>\n<td>0.84407</td>\n<td>0.81268</td>\n<td>• First version segmentation annotation data<br>• Epoch 100, Fold 0<br>• resize 224×224×224<br>• cls_modality</td>\n</tr>\n<tr>\n<td>Baseline + TTA 4x</td>\n<td>0.87056</td>\n<td>0.82508</td>\n<td>• First version segmentation annotation data<br>• Epoch 100, Fold 0<br>• resize 224×224×224<br>• cls_modality<br>• Test-Time Augmentation 4x</td>\n</tr>\n<tr>\n<td>Improved Data + TTA 4x</td>\n<td>0.88832</td>\n<td>0.85718</td>\n<td>• Improved segmentation annotation data<br>• Epoch 100, Fold 0<br>• resize 224×224×224<br>• cls_modality<br>• Test-Time Augmentation 4x</td>\n</tr>\n<tr>\n<td>Improved Data + TTA 8x</td>\n<td>0.89805</td>\n<td>0.86228</td>\n<td>• Improved segmentation annotation data<br>• Left-right data augmentation with label swap<br>• Epoch 100, Fold 0<br>• resize 224×224×224<br>• cls_modality<br>• Test-Time Augmentation 8x</td>\n</tr>\n<tr>\n<td>Improved Data + TTA 8x + Ensemble</td>\n<td>0.90035</td>\n<td>0.86727</td>\n<td>• Improved segmentation annotation data<br>• Left-right data augmentation with label swap<br>• Epoch 100, Fold 0 + Epoch 250, Fold 1 (Ensemble)<br>• resize 224×224×224<br>• cls_modality<br>• Test-Time Augmentation 8x</td>\n</tr>\n</tbody>\n</table>\n<h3>Key Technical Insights</h3>\n<ul>\n<li>ROI Extraction: Our Stage 1, 2D tri-axial ROI extraction significantly improved inference speed.</li>\n<li>Vessel Segmentation: Incorporating vessel segmentation clearly enhanced aneurysm segmentation and classification performance.</li>\n<li>Noise Handling: Our noise-handling techniques performed more effectively on the public dataset.</li>\n<li>Dataset Shift: Removing the 0.5 prediction fallback revealed a consistent 3–4% performance gap between the public and private leaderboards. This suggests a differing distribution of abnormal cases between the two test sets.</li>\n</ul>\n<h3>Performance Evolution</h3>\n<ul>\n<li>A consistent performance improvement was observed across all experiment iterations.</li>\n<li>Top Result: \"Improved Data + TTA 8x + Ensemble\" scored <strong>0.90035</strong> (Public LB, 1st) and <strong>0.86732</strong> (Private LB, 2nd).</li>\n</ul>\n<h2>Code Availability</h2>\n<ul>\n<li>Training and inference code: <a href=\"https://github.com/PengchengShi1220/RSNA2025_Intracranial-Aneurysm-Detection\" target=\"_blank\">https://github.com/PengchengShi1220/RSNA2025_Intracranial-Aneurysm-Detection</a></li>\n<li>Inference demo: <a href=\"https://www.kaggle.com/code/pengchengshi/bravecowcow-2nd-place-inference-demo\" target=\"_blank\">https://www.kaggle.com/code/pengchengshi/bravecowcow-2nd-place-inference-demo</a></li>\n<li>Final submission inference: <a href=\"https://www.kaggle.com/code/pengchengshi/bravecowcow-2nd-place-inference-final-submission\" target=\"_blank\">https://www.kaggle.com/code/pengchengshi/bravecowcow-2nd-place-inference-final-submission</a></li>\n<li>Stage 1 model checkpoint: <a href=\"https://www.kaggle.com/models/pengchengshi/dataset180_2d_vessel_box_seg_stable\" target=\"_blank\">https://www.kaggle.com/models/pengchengshi/dataset180_2d_vessel_box_seg_stable</a></li>\n<li>Stage 2 model checkpoint: <a href=\"https://www.kaggle.com/models/pengchengshi/dataset660_26classes_resize224_4661\" target=\"_blank\">https://www.kaggle.com/models/pengchengshi/dataset660_26classes_resize224_4661</a></li>\n</ul>\n<h2>Data Availability</h2>\n<ul>\n<li>Segmentation labels: <a href=\"https://huggingface.co/datasets/spc819/rsna2025-aneurysm-26class-seg\" target=\"_blank\">https://huggingface.co/datasets/spc819/rsna2025-aneurysm-26class-seg</a></li>\n</ul>\n<h2>3D Slicer Plugin</h2>\n<ul>\n<li>Source code: <a href=\"https://github.com/murong-xu/SlicerBraveCowCow\" target=\"_blank\">https://github.com/murong-xu/SlicerBraveCowCow</a></li>\n<li>Inference backend: <a href=\"https://github.com/huanghoujing/bravecowcow_inference_docker\" target=\"_blank\">https://github.com/huanghoujing/bravecowcow_inference_docker</a></li>\n</ul>\n<h2>Acknowledgements</h2>\n<p>Thanks to Medical Image Insights and UZH for compute support, Bjoern Menze and the Helmut Horten Foundation for funding support. We are grateful to RSNA/Kaggle hosts, nnU-Net devs, and forum contributors.</p>\n<h2>Citation</h2>\n<p>Please cite our work if it is helpful for your research:</p>\n<pre><code>@misc{shi2026intracranialaneurysmclassificationsegmentation,\n      title={Intracranial Aneurysm Classification and Segmentation via Tri-Axial ROI and Multi-Task Learning},\n      author={Pengcheng Shi and Kaiyuan Yang and Houjing Huang and Jiawei Chen and Yan Lu and Jiaqi Liu and Murong Xu and Minghui Zhang and Bjoern Menze and Xinglin Zhang},\n      year={2026},\n      eprint={2606.26706},\n      archivePrefix={arXiv},\n      primaryClass={cs.CV},\n      url={https://arxiv.org/abs/2606.26706},\n}\n</code></pre>\n<h2>References</h2>\n<ul>\n<li>Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J., &amp; Maier-Hein, K. H. (2021). nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nature Methods, 18(2), 203-211.</li>\n<li>Yang, Musio, Ma, et al. \"Benchmarking the CoW with the TopCoW Challenge: Topology-Aware Anatomical Segmentation of the Circle of Willis for CTA and MRA.\" arXiv (2025): arXiv-2312</li>\n</ul>",
      "rawMarkdown": "Thanks to @evancalabrese, RSNA, and Kaggle for organizing this Intracranial Aneurysm Detection competition! A special shoutout to the insightful discussions (especially @honglihang's detailed post on multi-frame DICOM analysis) that helped us gain a deeper understanding of the competition data.\n\nA detailed version of this work is available in the paper at: [https://arxiv.org/abs/2606.26706](https://arxiv.org/abs/2606.26706)\n\n## Overview\n\nOur approach focused on simplicity and generality to handle the diverse data in this classification-focused task. Key elements:\n\n- **Stage 1**: Fast 2D tri-axial ROI extraction using an nnU-Net 2D segmentation model to crop binary vascular regions efficiently, validated on all training cases locally.\n\n- **Stage 2**: 3D multi-task learning (segmentation of vessels and aneurysms + classification) based on nnU-Net, with enhancements like cross-attention pooling, modality 4-class heads, and targeted oversampling for rare classes. All data resized to uniform 224x224x224; heavy TTA (8x) including left-right flips with label swapping.\n\nBuilt entirely on the open-source nnU-Net framework—big thanks to @fabianisensee and team for its robust 3D segmentation baseline.\n\nInference optimized for speed (encoder + classification head only), running on 2x T4 GPUs in ~9 hours for conservative ensembles.\n\n## Who We Are\n\nWe're a mix of algorithm engineers and PhD researchers: Pengcheng Shi, Yan Lu, and Jiawei Chen from Medical Image Insights in Shanghai. Kaiyuan Yang and Houjing Huang from UZH in Zurich. Kaiyuan Yang, Houjing Huang, and Pengcheng Shi are also organizers of the MICCAI TopCoW (https://topcow24.grand-challenge.org/) / TopBrain (https://topbrain2025.grand-challenge.org/) challenges that benchmarked the segmentation of Circle of Willis (CoW) and whole brain vessel anatomy. Most of us are new to Kaggle (our background is in MICCAI events), and we noticed differences: Kaggle emphasizes data wrangling, efficiency under resource limits, and flat metrics (pure classification here, no localization/segmentation scores). In the TopCoW summary pre-print, we have previously explored the potential of TopCoW segmentation model at locating aneurysm with the CoW anatomy, which inspired many ideas used in our current solution.\n\n## Data Handling Challenges\n\nKaggle's data diversity (spacing, modalities, multi-frame DICOMs) required heavy preprocessing. Multi-frame issues were tricky—we had limited DICOM experience, and test sets had deleted fields, breaking dicom2nifti. We switched to pydicom, mapping spacing by slice shapes (e.g., <45 slices → 5mm), and trained a T2-specific orientation classifier for corrections. This ate up time but ensured full test coverage. No try-except fallbacks to 0.5 predictions in final inference—maximized robustness.\n\nWe used flipping TTA and multi-fold model ensembling to increase robustness of the model. Promising results showed that increasing in the public leaderboard led to consistent improvement in the private leaderboard.\n\n## Stage 1: 2D Tri-Axial ROI Extraction\n\nTo crop binary vascular regions efficiently:\n\n- Sample 3 slices per axis (1/4, 1/2, 3/4 positions) from iterative vascular ROI data → 9 slices total.\n- Train nnU-Net 2D config on sliding windows for cropped segmentation.\n- Inference: Merging sliding window patches into the batch dimension enhances both speed and robustness.\n- Local tests: Handled 100% of training cases correctly.\n![stage1](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4254412%2Fc7dd148e42375ae82883854f426db4ff%2Fstage1.png?generation=1760691510047622&alt=media)\nThis stage was key for quick, universal ROI on varied data.\n\n## Stage 2: 3D Multi-Task Learning\n\nUnified all spacing/modalities by resizing to 224x224x224.\n\n**Architecture**: nnU-Net 3D with multi-task heads (vessel/aneurysm seg + classification).\n\n**Enhancements**:\n\n- Cross-attention pooling surpassed average pooling in convergence speed.\n- Modality classification head enhanced training stability.\n- Inference: Only encoder and class head forwarded to reduce latency.\n- Data aug: Left-right flips with swapped labels and masks for classification and segmentation, effectively doubling dataset size.\n![vessel_flip](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4254412%2Ffbfc6b089e512c99a842a91afe975f31%2Fvessel_flip.png?generation=1760637300043403&alt=media)\n*Left: Original vascular anatomy segmentation mask | Right: Augmented mask after horizontal flipping with swapped labels*\n- Aneurysm masks were iteratively generated via model inference and manual refinement based on aneurysm center points.\n- 13-class vascular segmentation refined aneurysm masks using center-point distance heatmaps.\n- Rare classes were oversampled and assigned higher cross-entropy weights for both classification and segmentation.\n- Vessel and aneurysm segmentation stabilized classification convergence.\n- TTA: 8x (flips; swap left/right labels on outputs).\n- Final submissions: Due to platform constraints, we cut model count and ran a conservative two-fold ensemble on 2×T4 GPUs to prevent timeouts (runs were ~9 hours).\n![stage2](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4254412%2F0940c203faf2ab7400461b7f45fa746e%2Fstage2.png?generation=1764323763634081&alt=media)\n## Training Details\n\nBased on nnU-Net defaults, trained from scratch.\n\n**Loss**: CE with heatmap weighting for aneurysm centers.\n\n**Augmentations**: Standard nnU-Net (rotations, scaling, noise) + custom left-right flips and label swaps.\n\n**External data**: \n- TopCoW Training Data and its External Testsets: https://zenodo.org/records/15692630 (Note: The LargeIA dataset was excluded from training.)\n- TopBrain annotations: https://zenodo.org/records/16878417\n\n**Training data annotation and correction**: \n- Aneurysm seg mask: Center position provided in the challenge data csv combined with light manual annotation to iteratively annotate and train our own aneurysm segmentation model.\n- The provided rsna aneurysm annotations contain a few cases that mixed up left vs right side aneurysms, supra- vs infra-clinoid ICA aneurysms,and other position label mix-ups, which were manually corrected if identified.\n- Vessel seg mask customized by merging provided cow-seg masks with TopCoW and TopBrain annotations; a few hard cases, especially of T2 and T1-post (some T1-post even have artery flow void), were manually corrected.\n\n## Results and Insights\n### Ablation Study\n| Experiment Configuration | Public Score | Private Score | Notes |\n| :----------------------- | :----------- | :------------ | :---- |\n| Baseline | 0.84407 | 0.81268 | • First version segmentation annotation data<br>• Epoch 100, Fold 0<br>• resize 224×224×224<br>• cls_modality |\n| Baseline + TTA 4x | 0.87056 | 0.82508 | • First version segmentation annotation data<br>• Epoch 100, Fold 0<br>• resize 224×224×224<br>• cls_modality<br>• Test-Time Augmentation 4x |\n| Improved Data + TTA 4x | 0.88832 | 0.85718 | • Improved segmentation annotation data<br>• Epoch 100, Fold 0<br>• resize 224×224×224<br>• cls_modality<br>• Test-Time Augmentation 4x |\n| Improved Data + TTA 8x | 0.89805 | 0.86228 | • Improved segmentation annotation data<br>• Left-right data augmentation with label swap<br>• Epoch 100, Fold 0<br>• resize 224×224×224<br>• cls_modality<br>• Test-Time Augmentation 8x |\n| Improved Data + TTA 8x + Ensemble | 0.90035 | 0.86727 | • Improved segmentation annotation data<br>• Left-right data augmentation with label swap<br>• Epoch 100, Fold 0 + Epoch 250, Fold 1 (Ensemble)<br>• resize 224×224×224<br>• cls_modality<br>• Test-Time Augmentation 8x |\n\n### Key Technical Insights\n\n- ROI Extraction: Our Stage 1, 2D tri-axial ROI extraction significantly improved inference speed.\n- Vessel Segmentation: Incorporating vessel segmentation clearly enhanced aneurysm segmentation and classification performance.\n- Noise Handling: Our noise-handling techniques performed more effectively on the public dataset.\n- Dataset Shift: Removing the 0.5 prediction fallback revealed a consistent 3–4% performance gap between the public and private leaderboards. This suggests a differing distribution of abnormal cases between the two test sets.\n\n### Performance Evolution\n\n- A consistent performance improvement was observed across all experiment iterations.\n- Top Result: \"Improved Data + TTA 8x + Ensemble\" scored **0.90035** (Public LB, 1st) and **0.86732** (Private LB, 2nd).\n\n## Code Availability\n\n- Training and inference code: https://github.com/PengchengShi1220/RSNA2025_Intracranial-Aneurysm-Detection\n- Inference demo: https://www.kaggle.com/code/pengchengshi/bravecowcow-2nd-place-inference-demo\n- Final submission inference: https://www.kaggle.com/code/pengchengshi/bravecowcow-2nd-place-inference-final-submission\n- Stage 1 model checkpoint: https://www.kaggle.com/models/pengchengshi/dataset180_2d_vessel_box_seg_stable\n- Stage 2 model checkpoint: https://www.kaggle.com/models/pengchengshi/dataset660_26classes_resize224_4661\n\n## Data Availability\n\n- Segmentation labels: https://huggingface.co/datasets/spc819/rsna2025-aneurysm-26class-seg\n\n## 3D Slicer Plugin\n\n- Source code: https://github.com/murong-xu/SlicerBraveCowCow\n- Inference backend: https://github.com/huanghoujing/bravecowcow_inference_docker\n\n## Acknowledgements\n\nThanks to Medical Image Insights and UZH for compute support, Bjoern Menze and the Helmut Horten Foundation for funding support. We are grateful to RSNA/Kaggle hosts, nnU-Net devs, and forum contributors.\n\n## Citation\nPlease cite our work if it is helpful for your research:\n\n```\n@misc{shi2026intracranialaneurysmclassificationsegmentation,\n      title={Intracranial Aneurysm Classification and Segmentation via Tri-Axial ROI and Multi-Task Learning},\n      author={Pengcheng Shi and Kaiyuan Yang and Houjing Huang and Jiawei Chen and Yan Lu and Jiaqi Liu and Murong Xu and Minghui Zhang and Bjoern Menze and Xinglin Zhang},\n      year={2026},\n      eprint={2606.26706},\n      archivePrefix={arXiv},\n      primaryClass={cs.CV},\n      url={https://arxiv.org/abs/2606.26706},\n}\n```\n\n## References\n\n- Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J., & Maier-Hein, K. H. (2021). nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nature Methods, 18(2), 203-211.\n- Yang, Musio, Ma, et al. \"Benchmarking the CoW with the TopCoW Challenge: Topology-Aware Anatomical Segmentation of the Circle of Willis for CTA and MRA.\" arXiv (2025): arXiv-2312",
      "votes": 41
    },
    {
      "id": 3304418,
      "postDate": "2025-10-20T13:47:23.290Z",
      "content": "<p>First of all, congratulations on winning the 2nd place in the RSNA Intracranial Aneurysm Detection competition! Your proposed method is highly innovative and efficient, and I was particularly impressed by the \"Fast 2D tri-axial ROI extraction + 3D Multi-Task Segmentation and Classification\" solution. However, I have a question about one of the key steps: in the Stage 1 (2D Tri-Axial ROI Extraction), you sampled 3 slices from each axis (at the 1/4, 1/2, and 3/4 positions), resulting in a total of 9 slices, and used these to efficiently crop the binary vascular regions. I am very curious: what theoretical basis or preliminary experiments supported your decision that these 9 slices are sufficient to capture the key information needed for accurate ROI extraction? I would greatly appreciate it if you could share your insights to help me resolve this confusion. Thank you very much!</p>",
      "rawMarkdown": "First of all, congratulations on winning the 2nd place in the RSNA Intracranial Aneurysm Detection competition! Your proposed method is highly innovative and efficient, and I was particularly impressed by the \"Fast 2D tri-axial ROI extraction + 3D Multi-Task Segmentation and Classification\" solution. However, I have a question about one of the key steps: in the Stage 1 (2D Tri-Axial ROI Extraction), you sampled 3 slices from each axis (at the 1/4, 1/2, and 3/4 positions), resulting in a total of 9 slices, and used these to efficiently crop the binary vascular regions. I am very curious: what theoretical basis or preliminary experiments supported your decision that these 9 slices are sufficient to capture the key information needed for accurate ROI extraction? I would greatly appreciate it if you could share your insights to help me resolve this confusion. Thank you very much!",
      "votes": 1,
      "replies": [
        {
          "id": 3304443,
          "postDate": "2025-10-20T15:00:27.970Z",
          "content": "<p>Thank you so much!</p>\n<p>That's an excellent question about the Stage 1 ROI extraction. The core principle is that for each axis, our 2D model predicts a segmentation mask from which we extract 4 coordinates. By averaging predictions from different slices along the same axis, we get a robust estimate for that axis's contribution to the final 3D bounding box.</p>\n<p>Theoretically, if the target vascular ROI is sufficiently large, you only need one correctly predicted slice from <em>two</em> different axes to define the full 3D bounding box. However, to enhance robustness—especially since the vascular region isn't always centered—we sampled 3 slices per axis. This provides redundancy against potential outliers in any single prediction.</p>\n<p>In our local experiments, we found that a single slice per axis was often sufficient for an accurate ROI. Using 3 slices per axis was primarily to make the inference process more robust.</p>\n<p>For specific implementation details, please refer to the inference code here: <a href=\"https://github.com/PengchengShi1220/RSNA2025_Intracranial-Aneurysm-Detection/blob/5ee063e04e2e6d3da46fc7ab07e1a6b4e6a0cf60/nnXNet/nnxnet/inference/predict_from_raw_data_2D_orthogonal_planes_fast.py#L149\" target=\"_blank\">predict_from_multi_axial_slices</a></p>",
          "rawMarkdown": "Thank you so much!\n\nThat's an excellent question about the Stage 1 ROI extraction. The core principle is that for each axis, our 2D model predicts a segmentation mask from which we extract 4 coordinates. By averaging predictions from different slices along the same axis, we get a robust estimate for that axis's contribution to the final 3D bounding box.\n\nTheoretically, if the target vascular ROI is sufficiently large, you only need one correctly predicted slice from *two* different axes to define the full 3D bounding box. However, to enhance robustness—especially since the vascular region isn't always centered—we sampled 3 slices per axis. This provides redundancy against potential outliers in any single prediction.\n\nIn our local experiments, we found that a single slice per axis was often sufficient for an accurate ROI. Using 3 slices per axis was primarily to make the inference process more robust.\n\nFor specific implementation details, please refer to the inference code here: [predict_from_multi_axial_slices](https://github.com/PengchengShi1220/RSNA2025_Intracranial-Aneurysm-Detection/blob/5ee063e04e2e6d3da46fc7ab07e1a6b4e6a0cf60/nnXNet/nnxnet/inference/predict_from_raw_data_2D_orthogonal_planes_fast.py#L149)",
          "replies": [
            {
              "id": 3304666,
              "postDate": "2025-10-21T03:57:45.767Z",
              "content": "<p>Hello! Thank you very much for your clear explanation, which has given me a deeper understanding of the design of 3 slices for each of the 3 axes. Regarding the theoretical basis that \"if the vascular ROI is sufficiently large, it is only necessary to correctly predict one slice from two different axes to define a complete 3D bounding box\", I would like to further inquire: Is this conclusion based on specific medical image processing theories, similar methods in relevant literature, or a research summary by your team on the spatial distribution characteristics of vascular anatomical structures (such as the Circle of Willis)? If there is relevant theoretical background or previous research to support it, I hope to receive your sharing, as this will be of great help for me to deeply understand the rationality of this method. Thank you again for your patient explanation!</p>",
              "rawMarkdown": "Hello! Thank you very much for your clear explanation, which has given me a deeper understanding of the design of 3 slices for each of the 3 axes. Regarding the theoretical basis that \"if the vascular ROI is sufficiently large, it is only necessary to correctly predict one slice from two different axes to define a complete 3D bounding box\", I would like to further inquire: Is this conclusion based on specific medical image processing theories, similar methods in relevant literature, or a research summary by your team on the spatial distribution characteristics of vascular anatomical structures (such as the Circle of Willis)? If there is relevant theoretical background or previous research to support it, I hope to receive your sharing, as this will be of great help for me to deeply understand the rationality of this method. Thank you again for your patient explanation!"
            },
            {
              "id": 3304684,
              "postDate": "2025-10-21T04:46:54.070Z",
              "content": "<p>Hello! This conclusion is primarily based on our experimental results. More detailed information will be available in our subsequent paper for your reference.</p>",
              "rawMarkdown": "Hello! This conclusion is primarily based on our experimental results. More detailed information will be available in our subsequent paper for your reference."
            }
          ]
        }
      ]
    },
    {
      "id": 3303133,
      "postDate": "2025-10-17T08:00:32.710Z",
      "content": "<p>Congratualtions! Nice work</p>",
      "rawMarkdown": "Congratualtions! Nice work",
      "votes": 1,
      "replies": [
        {
          "id": 3303149,
          "postDate": "2025-10-17T08:41:58.367Z",
          "content": "<p>Thank you! We truly appreciate your contribution to the community.</p>",
          "rawMarkdown": "Thank you! We truly appreciate your contribution to the community."
        }
      ]
    },
    {
      "id": 3302591,
      "postDate": "2025-10-16T05:50:05.387Z",
      "content": "<p>great work. i always thought monai was useful for medical stuff, but nnUnet seems pretty powerful </p>",
      "rawMarkdown": "great work. i always thought monai was useful for medical stuff, but nnUnet seems pretty powerful ",
      "votes": 1,
      "replies": [
        {
          "id": 3302605,
          "postDate": "2025-10-16T06:30:56.303Z",
          "content": "<p>Thanks! I've really benefited from nnU-Net. Building improvements on top of it makes it easier to achieve optimal performance.</p>",
          "rawMarkdown": "Thanks! I've really benefited from nnU-Net. Building improvements on top of it makes it easier to achieve optimal performance."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3304418,
      "author_name": "shixingkong",
      "author_url": "",
      "post_date": "2025-10-20T13:47:23.290000",
      "content": "<p>First of all, congratulations on winning the 2nd place in the RSNA Intracranial Aneurysm Detection competition! Your proposed method is highly innovative and efficient, and I was particularly impressed by the \"Fast 2D tri-axial ROI extraction + 3D Multi-Task Segmentation and Classification\" solution. However, I have a question about one of the key steps: in the Stage 1 (2D Tri-Axial ROI Extraction), you sampled 3 slices from each axis (at the 1/4, 1/2, and 3/4 positions), resulting in a total of 9 slices, and used these to efficiently crop the binary vascular regions. I am very curious: what theoretical basis or preliminary experiments supported your decision that these 9 slices are sufficient to capture the key information needed for accurate ROI extraction? I would greatly appreciate it if you could share your insights to help me resolve this confusion. Thank you very much!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3304443,
          "author_name": "Pengcheng Shi",
          "author_url": "",
          "post_date": "2025-10-20T15:00:27.970000",
          "content": "<p>Thank you so much!</p>\n<p>That's an excellent question about the Stage 1 ROI extraction. The core principle is that for each axis, our 2D model predicts a segmentation mask from which we extract 4 coordinates. By averaging predictions from different slices along the same axis, we get a robust estimate for that axis's contribution to the final 3D bounding box.</p>\n<p>Theoretically, if the target vascular ROI is sufficiently large, you only need one correctly predicted slice from <em>two</em> different axes to define the full 3D bounding box. However, to enhance robustness—especially since the vascular region isn't always centered—we sampled 3 slices per axis. This provides redundancy against potential outliers in any single prediction.</p>\n<p>In our local experiments, we found that a single slice per axis was often sufficient for an accurate ROI. Using 3 slices per axis was primarily to make the inference process more robust.</p>\n<p>For specific implementation details, please refer to the inference code here: <a href=\"https://github.com/PengchengShi1220/RSNA2025_Intracranial-Aneurysm-Detection/blob/5ee063e04e2e6d3da46fc7ab07e1a6b4e6a0cf60/nnXNet/nnxnet/inference/predict_from_raw_data_2D_orthogonal_planes_fast.py#L149\" target=\"_blank\">predict_from_multi_axial_slices</a></p>",
          "votes": 0,
          "replies": [
            {
              "id": 3304666,
              "author_name": "shixingkong",
              "author_url": "",
              "post_date": "2025-10-21T03:57:45.767000",
              "content": "<p>Hello! Thank you very much for your clear explanation, which has given me a deeper understanding of the design of 3 slices for each of the 3 axes. Regarding the theoretical basis that \"if the vascular ROI is sufficiently large, it is only necessary to correctly predict one slice from two different axes to define a complete 3D bounding box\", I would like to further inquire: Is this conclusion based on specific medical image processing theories, similar methods in relevant literature, or a research summary by your team on the spatial distribution characteristics of vascular anatomical structures (such as the Circle of Willis)? If there is relevant theoretical background or previous research to support it, I hope to receive your sharing, as this will be of great help for me to deeply understand the rationality of this method. Thank you again for your patient explanation!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3304684,
              "author_name": "Pengcheng Shi",
              "author_url": "",
              "post_date": "2025-10-21T04:46:54.070000",
              "content": "<p>Hello! This conclusion is primarily based on our experimental results. More detailed information will be available in our subsequent paper for your reference.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3303133,
      "author_name": "FabianIsensee",
      "author_url": "",
      "post_date": "2025-10-17T08:00:32.710000",
      "content": "<p>Congratualtions! Nice work</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3303149,
          "author_name": "Pengcheng Shi",
          "author_url": "",
          "post_date": "2025-10-17T08:41:58.367000",
          "content": "<p>Thank you! We truly appreciate your contribution to the community.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3302591,
      "author_name": "samu2505",
      "author_url": "",
      "post_date": "2025-10-16T05:50:05.387000",
      "content": "<p>great work. i always thought monai was useful for medical stuff, but nnUnet seems pretty powerful </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3302605,
          "author_name": "Pengcheng Shi",
          "author_url": "",
          "post_date": "2025-10-16T06:30:56.303000",
          "content": "<p>Thanks! I've really benefited from nnU-Net. Building improvements on top of it makes it easier to achieve optimal performance.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3302177": "Thanks to @evancalabrese, RSNA, and Kaggle for organizing this Intracranial Aneurysm Detection competition! A special shoutout to the insightful discussions (especially @honglihang's detailed post on multi-frame DICOM analysis) that helped us gain a deeper understanding of the competition data.\n\nA detailed version of this work is available in the paper at: [https://arxiv.org/abs/2606.26706](https://arxiv.org/abs/2606.26706)\n\n## Overview\n\nOur approach focused on simplicity and generality to handle the diverse data in this classification-focused task. Key elements:\n\n- **Stage 1**: Fast 2D tri-axial ROI extraction using an nnU-Net 2D segmentation model to crop binary vascular regions efficiently, validated on all training cases locally.\n\n- **Stage 2**: 3D multi-task learning (segmentation of vessels and aneurysms + classification) based on nnU-Net, with enhancements like cross-attention pooling, modality 4-class heads, and targeted oversampling for rare classes. All data resized to uniform 224x224x224; heavy TTA (8x) including left-right flips with label swapping.\n\nBuilt entirely on the open-source nnU-Net framework—big thanks to @fabianisensee and team for its robust 3D segmentation baseline.\n\nInference optimized for speed (encoder + classification head only), running on 2x T4 GPUs in ~9 hours for conservative ensembles.\n\n## Who We Are\n\nWe're a mix of algorithm engineers and PhD researchers: Pengcheng Shi, Yan Lu, and Jiawei Chen from Medical Image Insights in Shanghai. Kaiyuan Yang and Houjing Huang from UZH in Zurich. Kaiyuan Yang, Houjing Huang, and Pengcheng Shi are also organizers of the MICCAI TopCoW (https://topcow24.grand-challenge.org/) / TopBrain (https://topbrain2025.grand-challenge.org/) challenges that benchmarked the segmentation of Circle of Willis (CoW) and whole brain vessel anatomy. Most of us are new to Kaggle (our background is in MICCAI events), and we noticed differences: Kaggle emphasizes data wrangling, efficiency under resource limits, and flat metrics (pure classification here, no localization/segmentation scores). In the TopCoW summary pre-print, we have previously explored the potential of TopCoW segmentation model at locating aneurysm with the CoW anatomy, which inspired many ideas used in our current solution.\n\n## Data Handling Challenges\n\nKaggle's data diversity (spacing, modalities, multi-frame DICOMs) required heavy preprocessing. Multi-frame issues were tricky—we had limited DICOM experience, and test sets had deleted fields, breaking dicom2nifti. We switched to pydicom, mapping spacing by slice shapes (e.g., <45 slices → 5mm), and trained a T2-specific orientation classifier for corrections. This ate up time but ensured full test coverage. No try-except fallbacks to 0.5 predictions in final inference—maximized robustness.\n\nWe used flipping TTA and multi-fold model ensembling to increase robustness of the model. Promising results showed that increasing in the public leaderboard led to consistent improvement in the private leaderboard.\n\n## Stage 1: 2D Tri-Axial ROI Extraction\n\nTo crop binary vascular regions efficiently:\n\n- Sample 3 slices per axis (1/4, 1/2, 3/4 positions) from iterative vascular ROI data → 9 slices total.\n- Train nnU-Net 2D config on sliding windows for cropped segmentation.\n- Inference: Merging sliding window patches into the batch dimension enhances both speed and robustness.\n- Local tests: Handled 100% of training cases correctly.\n![stage1](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4254412%2Fc7dd148e42375ae82883854f426db4ff%2Fstage1.png?generation=1760691510047622&alt=media)\nThis stage was key for quick, universal ROI on varied data.\n\n## Stage 2: 3D Multi-Task Learning\n\nUnified all spacing/modalities by resizing to 224x224x224.\n\n**Architecture**: nnU-Net 3D with multi-task heads (vessel/aneurysm seg + classification).\n\n**Enhancements**:\n\n- Cross-attention pooling surpassed average pooling in convergence speed.\n- Modality classification head enhanced training stability.\n- Inference: Only encoder and class head forwarded to reduce latency.\n- Data aug: Left-right flips with swapped labels and masks for classification and segmentation, effectively doubling dataset size.\n![vessel_flip](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4254412%2Ffbfc6b089e512c99a842a91afe975f31%2Fvessel_flip.png?generation=1760637300043403&alt=media)\n*Left: Original vascular anatomy segmentation mask | Right: Augmented mask after horizontal flipping with swapped labels*\n- Aneurysm masks were iteratively generated via model inference and manual refinement based on aneurysm center points.\n- 13-class vascular segmentation refined aneurysm masks using center-point distance heatmaps.\n- Rare classes were oversampled and assigned higher cross-entropy weights for both classification and segmentation.\n- Vessel and aneurysm segmentation stabilized classification convergence.\n- TTA: 8x (flips; swap left/right labels on outputs).\n- Final submissions: Due to platform constraints, we cut model count and ran a conservative two-fold ensemble on 2×T4 GPUs to prevent timeouts (runs were ~9 hours).\n![stage2](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4254412%2F0940c203faf2ab7400461b7f45fa746e%2Fstage2.png?generation=1764323763634081&alt=media)\n## Training Details\n\nBased on nnU-Net defaults, trained from scratch.\n\n**Loss**: CE with heatmap weighting for aneurysm centers.\n\n**Augmentations**: Standard nnU-Net (rotations, scaling, noise) + custom left-right flips and label swaps.\n\n**External data**: \n- TopCoW Training Data and its External Testsets: https://zenodo.org/records/15692630 (Note: The LargeIA dataset was excluded from training.)\n- TopBrain annotations: https://zenodo.org/records/16878417\n\n**Training data annotation and correction**: \n- Aneurysm seg mask: Center position provided in the challenge data csv combined with light manual annotation to iteratively annotate and train our own aneurysm segmentation model.\n- The provided rsna aneurysm annotations contain a few cases that mixed up left vs right side aneurysms, supra- vs infra-clinoid ICA aneurysms,and other position label mix-ups, which were manually corrected if identified.\n- Vessel seg mask customized by merging provided cow-seg masks with TopCoW and TopBrain annotations; a few hard cases, especially of T2 and T1-post (some T1-post even have artery flow void), were manually corrected.\n\n## Results and Insights\n### Ablation Study\n| Experiment Configuration | Public Score | Private Score | Notes |\n| :----------------------- | :----------- | :------------ | :---- |\n| Baseline | 0.84407 | 0.81268 | • First version segmentation annotation data<br>• Epoch 100, Fold 0<br>• resize 224×224×224<br>• cls_modality |\n| Baseline + TTA 4x | 0.87056 | 0.82508 | • First version segmentation annotation data<br>• Epoch 100, Fold 0<br>• resize 224×224×224<br>• cls_modality<br>• Test-Time Augmentation 4x |\n| Improved Data + TTA 4x | 0.88832 | 0.85718 | • Improved segmentation annotation data<br>• Epoch 100, Fold 0<br>• resize 224×224×224<br>• cls_modality<br>• Test-Time Augmentation 4x |\n| Improved Data + TTA 8x | 0.89805 | 0.86228 | • Improved segmentation annotation data<br>• Left-right data augmentation with label swap<br>• Epoch 100, Fold 0<br>• resize 224×224×224<br>• cls_modality<br>• Test-Time Augmentation 8x |\n| Improved Data + TTA 8x + Ensemble | 0.90035 | 0.86727 | • Improved segmentation annotation data<br>• Left-right data augmentation with label swap<br>• Epoch 100, Fold 0 + Epoch 250, Fold 1 (Ensemble)<br>• resize 224×224×224<br>• cls_modality<br>• Test-Time Augmentation 8x |\n\n### Key Technical Insights\n\n- ROI Extraction: Our Stage 1, 2D tri-axial ROI extraction significantly improved inference speed.\n- Vessel Segmentation: Incorporating vessel segmentation clearly enhanced aneurysm segmentation and classification performance.\n- Noise Handling: Our noise-handling techniques performed more effectively on the public dataset.\n- Dataset Shift: Removing the 0.5 prediction fallback revealed a consistent 3–4% performance gap between the public and private leaderboards. This suggests a differing distribution of abnormal cases between the two test sets.\n\n### Performance Evolution\n\n- A consistent performance improvement was observed across all experiment iterations.\n- Top Result: \"Improved Data + TTA 8x + Ensemble\" scored **0.90035** (Public LB, 1st) and **0.86732** (Private LB, 2nd).\n\n## Code Availability\n\n- Training and inference code: https://github.com/PengchengShi1220/RSNA2025_Intracranial-Aneurysm-Detection\n- Inference demo: https://www.kaggle.com/code/pengchengshi/bravecowcow-2nd-place-inference-demo\n- Final submission inference: https://www.kaggle.com/code/pengchengshi/bravecowcow-2nd-place-inference-final-submission\n- Stage 1 model checkpoint: https://www.kaggle.com/models/pengchengshi/dataset180_2d_vessel_box_seg_stable\n- Stage 2 model checkpoint: https://www.kaggle.com/models/pengchengshi/dataset660_26classes_resize224_4661\n\n## Data Availability\n\n- Segmentation labels: https://huggingface.co/datasets/spc819/rsna2025-aneurysm-26class-seg\n\n## 3D Slicer Plugin\n\n- Source code: https://github.com/murong-xu/SlicerBraveCowCow\n- Inference backend: https://github.com/huanghoujing/bravecowcow_inference_docker\n\n## Acknowledgements\n\nThanks to Medical Image Insights and UZH for compute support, Bjoern Menze and the Helmut Horten Foundation for funding support. We are grateful to RSNA/Kaggle hosts, nnU-Net devs, and forum contributors.\n\n## Citation\nPlease cite our work if it is helpful for your research:\n\n```\n@misc{shi2026intracranialaneurysmclassificationsegmentation,\n      title={Intracranial Aneurysm Classification and Segmentation via Tri-Axial ROI and Multi-Task Learning},\n      author={Pengcheng Shi and Kaiyuan Yang and Houjing Huang and Jiawei Chen and Yan Lu and Jiaqi Liu and Murong Xu and Minghui Zhang and Bjoern Menze and Xinglin Zhang},\n      year={2026},\n      eprint={2606.26706},\n      archivePrefix={arXiv},\n      primaryClass={cs.CV},\n      url={https://arxiv.org/abs/2606.26706},\n}\n```\n\n## References\n\n- Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J., & Maier-Hein, K. H. (2021). nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nature Methods, 18(2), 203-211.\n- Yang, Musio, Ma, et al. \"Benchmarking the CoW with the TopCoW Challenge: Topology-Aware Anatomical Segmentation of the Circle of Willis for CTA and MRA.\" arXiv (2025): arXiv-2312",
    "3304418": "First of all, congratulations on winning the 2nd place in the RSNA Intracranial Aneurysm Detection competition! Your proposed method is highly innovative and efficient, and I was particularly impressed by the \"Fast 2D tri-axial ROI extraction + 3D Multi-Task Segmentation and Classification\" solution. However, I have a question about one of the key steps: in the Stage 1 (2D Tri-Axial ROI Extraction), you sampled 3 slices from each axis (at the 1/4, 1/2, and 3/4 positions), resulting in a total of 9 slices, and used these to efficiently crop the binary vascular regions. I am very curious: what theoretical basis or preliminary experiments supported your decision that these 9 slices are sufficient to capture the key information needed for accurate ROI extraction? I would greatly appreciate it if you could share your insights to help me resolve this confusion. Thank you very much!",
    "3303133": "Congratualtions! Nice work",
    "3302591": "great work. i always thought monai was useful for medical stuff, but nnUnet seems pretty powerful "
  }
}