{
  "id": 612039,
  "title": "7th place solution - 3D nnU-Net + blob regression (again)",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/612039",
  "author_name": "Stefan Denner",
  "post_date": "2025-10-16T09:33:32.032000",
  "votes": 24,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Thanks to <a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a>, <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a>, RSNA, and Kaggle for organizing this Intracranial Aneurysm Detection competition! </p>\n<h1>Overview (TLDR)</h1>\n<p>Here’s a brief rundown of our solution — it’s straightforward and easy to implement:</p>\n<ul>\n<li>We formulate the task as multichannel blob regression, optimized using a TopK (20%) BCE loss and then taking the maximum per channel as probability prediction.</li>\n<li>We build on <a href=\"https://github.com/MIC-DKFZ/nnUNet\" target=\"_blank\">nnU-Net</a>, the leading framework for 3D medical image segmentation. We already adapted it for our <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/mic-dkfz-2nd-place-solution-3d-nnu-net-blob-regres\" target=\"_blank\">2nd place solution in the BYU - Locating Bacterial Flagellar Motors 2025</a>.</li>\n<li>Our model is a 3D U-Net with a residual encoder, trained from scratch.</li>\n<li>Inference is done with a single model without test-time augmentation (due to time restrictions).</li>\n<li>Our model achieved a score of 0.83 / 0.83 on the public/private leaderboard.</li>\n</ul>\n<h1>Who are we?</h1>\n<p>We are a team of colleagues (scientists and PhD students) affiliated with the <a href=\"https://www.dkfz.de/en/medical-image-computing\" target=\"_blank\">Divisions of Medical Image Computing</a> at the German Cancer Research Center, as well as <a href=\"https://helmholtz-imaging.de/\" target=\"_blank\">Helmholtz Imaging</a>. Our expertise lies in 3D image analysis — particularly in solving 3D segmentation problems and developing infrastructure to bring algorithms into clinical practice. </p>\n<h1>Method</h1>\n<p>We modeled the task as a heatmap regression and built up on our <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/mic-dkfz-2nd-place-solution-3d-nnu-net-blob-regres\" target=\"_blank\">2nd place solution in the BYU - Locating Bacterial Flagellar Motors 2025</a>. </p>\n<h3>Data used</h3>\n<p>We converted the DICOM series to .nii.gz format using the pydicom library, processing series with 2D DICOM files in parallel. We also used this library to extract information on spacing, origin, and direction. We then converted the resulting image into a SimpleITK image and oriented it in RAS orientation.</p>\n<p>To reduce computing resources, we derived a [200, 160, 160] mm cubic Region of Interest (ROI) on the central superior region of the image. We ensured that all aneurysms in the training set were included in the ROI. Fig. 1 depicts the ROI on top of the image.</p>\n<p>We observed that several series presented defects such as unexpected orientations, empty series, shunt artifacts, movement artifacts, or images with an empty superior space. We decided to keep those series with mild artifacts such as mis-orientations and fixed them via flipping. Series with stronger artifacts such as totally empty images were discarded. In total, ten volumes were discarded.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1460057%2Faa77cf1f2c0d097a505297609efdd7b4%2Fbbox.png?generation=1760605343255626&amp;alt=media\" alt=\"\"><br>\nFigure 1. Axial, coronal and sagittal views of example image series and the ROI box, in red.</p>\n<h3>Preprocessing</h3>\n<p>Data preprocessing was completed with the self-configurable segmentation framework nnU-Net, following a 3D full resolution configuration. The volumes were loaded with a SimpleITK reader that also tries to enforce RAS orientation. All images were resampled to the median spacing found from all training images: <a href=\"mm\" target=\"_blank\">0.70, 0.47, 0.47</a>, being normalized via z-score normalization based on a global mean and a global standard deviation extracted from the training dataset. </p>\n<p>The images were resampled with a special resampler from PyTorch, which is faster than other commonly used functions such as scipy.ndimage.zoom, given the strong time constraints.</p>\n<h3>Network architecture</h3>\n<p>We use nnU-Net’s ResEnc, which is essentially a UNet with a residual encoder and a lightweight convolutional decoder. The architecture included six stages with [32, 64, 128, 256, 320, 320] features in each stage, respectively. </p>\n<h3>Training Procedure</h3>\n<p>We split the provided challenge data into five cross-validation folds, stratifying for modalities across folds, rather than on vessel classes, since we wanted to ensure an adequate performance across all image modalities. Since we joined the challenge relatively late and there were many potential design choices to test, most of the hyperparameter tuning happened exclusively on the first fold of the cross-validation scheme. </p>\n<h3>Blob Regression with nnU-Net</h3>\n<p>nnU-Net is built for semantic segmentation. This also includes its expected data structure. To make it compatible with aneurysm regression we store the ground truth as semantic segmentation maps, where each aneurysm is encoded with a sphere (r=5 voxels) with an integer label representing the vessel class of the ground-truth aneurysm. These spheres are treated by nnU-Net as segmentations and are passed through the data loading and augmentation pipeline as nnU-Net normally would, thus properly applying rotations, mirroring etc, although we did not apply mirroring augmentations in the left/right axis, since several of the labels contained a left/right codification. At the end of the dataloading pipeline we inject a custom transform that converts each aneurysm instance into a blob of the respective channel. We model the 14 classes as separate channels, where the 14th class (Aneurysm Present) is the pixelwise maximum of the 13 anatomical classes.</p>\n<p>We use ‘EDT blobs’, basically 3D spheres that were transformed using the Euclidean Distance Transform (EDT) and rescaled to have a value range of [0, 1]. The optimized sphere size in the first cross-validation fold was 65 voxels. We experimented with sphere sizes from 15 to 95 voxel radii.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1460057%2F2bb5653e29b17b7558cbd3b7ac2b720e%2Fblobs.png?generation=1760607125393264&amp;alt=media\" alt=\"\"><br>\nFigure 2: Blob (EDT, radius 65) at the Right Middle Cerebral Artery (Image not resampled yet)</p>\n<h3>Hyperparameters</h3>\n<p>Our final model was trained with a batch size of 32 and a patch size of 96,x160x128 voxels. Initial learning rate is 0.01 and is decayed over the course of the training using polyLR schedule (same as default nnU-Net). We train with SGD for 3000 epochs (250 iterations per epoch). The loss function utilized is binary cross-entropy, computed only on the 20% worst voxels (the ones with the highest loss value, computed over the entire batch).</p>\n<p>Our final model was trained on 4xA100 40GB using PyTorch's DDP. Training took 4.5 days. </p>\n<h3>Inference</h3>\n<p>We largely use nnU-Net’s inference infrastructure. The input series are dissected into a series of patches. Each patch gets blob regressed. We take the maximum per channel as class probability. We max-aggregate across patches for our final predictions. <br>\nWe used the 2xT4 instances for the prediction and split the patches to predict for each input series evenly across the GPUs. We always use a single model, no ensemble. Inference takes approximately 8 hours.</p>\n<h1>Results</h1>\n<p>Unfortunately, we were hit quite hard by the instabilities of the Kaggle platform. <br>\nWhile our model finished training for 3000 epochs, only submissions until 1500 epochs (three days before the deadline) were successful (private and public LB 0.83). Subsequent submissions timed out, even though only the model weights differed.<br>\nOur internal validation showed that later checkpoints, TTA and Gaussian weighting of the patches would have probably further improved our performance (also previous submissions showed this). <br>\nSurprisingly, our internal performance went up until 0.9 which was not achieved on the leaderboard. We don’t know where this shift comes from. One reason might be that we had to embed our inference in a try/catch block, else an error was thrown after 15min. We don’t know exactly why. <br>\nWe did not exploit the segmentation masks, which could have been added as auxiliary outputs during training to help the model localize better. Due to joining late, we didn't find the time to investigate this. However, other teams showed that this improved their performance.</p>\n<h3>What did not work?</h3>\n<ul>\n<li>We also framed the problem as a detection problem, attempting to solve it with the self-configuration detection framework <a href=\"https://github.com/MIC-DKFZ/nnDetection\" target=\"_blank\">nnDetection</a>. This approach performed better than the solution presented here in the public leaderboard, but it underperformed in the private leaderboard and in our internal validation. <ul>\n<li>nnDetection was trained on instance segmentation label versions, also on cropped data, consisting of a self-configured <a href=\"https://proceedings.mlr.press/v116/jaeger20a/jaeger20a.pdf\" target=\"_blank\">Retina U-Net architecture</a> that learned from the aneurysm positions encoded as boxes and from the aneurysm segmentations . Unlike the solution here, it resampled the input series to an isometric space of [1.0, 1.0, 1.0] mm, training with a batch size of 4 for 100 epochs (2500 iterations per epoch), hybrid loss function combining L1 loss for the regression of box coordinates and focal loss for box class estimation, polynomial learning rate scheduling from an initial value of 0.001, and SGD optimizer with Nesterov momentum. Predicted boxes were postprocessed via non-maximum suppression with a 0.1 intersection over union threshold. We additionally managed to conduct inference with 8 test time augmentations and an inference patch overlap of 0.25.</li></ul></li>\n<li>Isometric space resampling with [1.0, 1.0, 1.0] mm was also implemented, given its potential for faster image processing. However, it substantially worsened our results, so it was discontinued early on.</li>\n<li>We also explored co-training with external aneurysm datasets containing binary classes to better model the Aneurysm Present class (<a href=\"https://adam.isi.uu.nl/data/\" target=\"_blank\">ADAM</a>, <a href=\"https://zenodo.org/records/6801398\" target=\"_blank\">Large IA Segmentation dataset</a>, <a href=\"https://www.codabench.org/competitions/2139/\" target=\"_blank\">INSTED</a>, <a href=\"https://openneuro.org/datasets/ds003949/versions/1.0.1\" target=\"_blank\">Lausanne TOF-MRA Aneurysm Cohort</a>, <a href=\"https://openneuro.org/datasets/ds005096/versions/1.0.3\" target=\"_blank\">Royal Brisbane TOFMRA Intracranial Aneurysm Database</a>, <a href=\"https://github.com/jinxiaokuang/RWS-MT?tab=readme-ov-file\" target=\"_blank\">Jianxiaokuang aneurysm dataset</a>. We realized afterwards that some of these datasets (<a href=\"https://adam.isi.uu.nl/data/\" target=\"_blank\">ADAM</a>,  <a href=\"https://zenodo.org/records/6801398\" target=\"_blank\">Large IA Segmentation dataset</a>) were not allowed, so we discarded them and ran co-training without them. In the end, co-training did not really help, so we resorted back to training only on the challenge cases from scratch. </li>\n<li>We first started with processing the image as a whole but time limitations forced us to crop around the ROI which also resulted in better performance.</li>\n<li>As described above, in our final model we just max-aggregate the patch predictions. However, it is known that the model has some uncertainty close to the edges. A common strategy to mitigate this is gaussian weighting the predictions for each patch (high weight in the center, low weight and the borders). In earlier submissions we saw that this improved our performance. However, in our final model we could not apply this strategy because of platform instabilities. </li>\n<li>We also tried to train with larger patch sizes, which, surprisingly, did not contribute to improve our scores.</li>\n</ul>\n<h3>What would we have wished for?</h3>\n<p>We already stated in the discussion forum that the signature of the predict function was limiting us quite a lot in how we can parallelize processing. More about that <a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/611778\" target=\"_blank\">here</a>. <br>\nThis requirement made it even more difficult for us since we required a specific numpy version which forced us to run our code as a subprocess. We ended up spawning a proxy worker with which we communicated via the std output. This added significant boilerplate and complexity, and felt unnatural.</p>\n<p>We would have also wished for longer submission times, and a less dependent server architecture on the number of submissions sent by different teams, since many of our submissions timed out during the last few days of the challenge due to an increasing workload. <br>\nOn a much broader scope: Kaggle's submission notebook style made it quite hard for us (also the last time). Having the possibility to just use Docker containers would have eased our lives a lot because they allow much higher flexibility.  </p>\n<h1>Acknowledgements</h1>\n<p>We thank RSNA for organizing and Kaggle for hosting this competition. We furthermore want to give a shoutout to our <a href=\"https://www.dkfz.de/en/medical-image-computing\" target=\"_blank\">Divisions of Medical Image Computing</a> at the German Cancer Research Center, as well as <a href=\"https://helmholtz-imaging.de/\" target=\"_blank\">Helmholtz Imaging</a></p>",
  "messages": [
    {
      "id": 3302691,
      "postDate": "2025-10-16T09:33:32.033Z",
      "content": "<p>Thanks to <a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a>, <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a>, RSNA, and Kaggle for organizing this Intracranial Aneurysm Detection competition! </p>\n<h1>Overview (TLDR)</h1>\n<p>Here’s a brief rundown of our solution — it’s straightforward and easy to implement:</p>\n<ul>\n<li>We formulate the task as multichannel blob regression, optimized using a TopK (20%) BCE loss and then taking the maximum per channel as probability prediction.</li>\n<li>We build on <a href=\"https://github.com/MIC-DKFZ/nnUNet\" target=\"_blank\">nnU-Net</a>, the leading framework for 3D medical image segmentation. We already adapted it for our <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/mic-dkfz-2nd-place-solution-3d-nnu-net-blob-regres\" target=\"_blank\">2nd place solution in the BYU - Locating Bacterial Flagellar Motors 2025</a>.</li>\n<li>Our model is a 3D U-Net with a residual encoder, trained from scratch.</li>\n<li>Inference is done with a single model without test-time augmentation (due to time restrictions).</li>\n<li>Our model achieved a score of 0.83 / 0.83 on the public/private leaderboard.</li>\n</ul>\n<h1>Who are we?</h1>\n<p>We are a team of colleagues (scientists and PhD students) affiliated with the <a href=\"https://www.dkfz.de/en/medical-image-computing\" target=\"_blank\">Divisions of Medical Image Computing</a> at the German Cancer Research Center, as well as <a href=\"https://helmholtz-imaging.de/\" target=\"_blank\">Helmholtz Imaging</a>. Our expertise lies in 3D image analysis — particularly in solving 3D segmentation problems and developing infrastructure to bring algorithms into clinical practice. </p>\n<h1>Method</h1>\n<p>We modeled the task as a heatmap regression and built up on our <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/mic-dkfz-2nd-place-solution-3d-nnu-net-blob-regres\" target=\"_blank\">2nd place solution in the BYU - Locating Bacterial Flagellar Motors 2025</a>. </p>\n<h3>Data used</h3>\n<p>We converted the DICOM series to .nii.gz format using the pydicom library, processing series with 2D DICOM files in parallel. We also used this library to extract information on spacing, origin, and direction. We then converted the resulting image into a SimpleITK image and oriented it in RAS orientation.</p>\n<p>To reduce computing resources, we derived a [200, 160, 160] mm cubic Region of Interest (ROI) on the central superior region of the image. We ensured that all aneurysms in the training set were included in the ROI. Fig. 1 depicts the ROI on top of the image.</p>\n<p>We observed that several series presented defects such as unexpected orientations, empty series, shunt artifacts, movement artifacts, or images with an empty superior space. We decided to keep those series with mild artifacts such as mis-orientations and fixed them via flipping. Series with stronger artifacts such as totally empty images were discarded. In total, ten volumes were discarded.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1460057%2Faa77cf1f2c0d097a505297609efdd7b4%2Fbbox.png?generation=1760605343255626&amp;alt=media\" alt=\"\"><br>\nFigure 1. Axial, coronal and sagittal views of example image series and the ROI box, in red.</p>\n<h3>Preprocessing</h3>\n<p>Data preprocessing was completed with the self-configurable segmentation framework nnU-Net, following a 3D full resolution configuration. The volumes were loaded with a SimpleITK reader that also tries to enforce RAS orientation. All images were resampled to the median spacing found from all training images: <a href=\"mm\" target=\"_blank\">0.70, 0.47, 0.47</a>, being normalized via z-score normalization based on a global mean and a global standard deviation extracted from the training dataset. </p>\n<p>The images were resampled with a special resampler from PyTorch, which is faster than other commonly used functions such as scipy.ndimage.zoom, given the strong time constraints.</p>\n<h3>Network architecture</h3>\n<p>We use nnU-Net’s ResEnc, which is essentially a UNet with a residual encoder and a lightweight convolutional decoder. The architecture included six stages with [32, 64, 128, 256, 320, 320] features in each stage, respectively. </p>\n<h3>Training Procedure</h3>\n<p>We split the provided challenge data into five cross-validation folds, stratifying for modalities across folds, rather than on vessel classes, since we wanted to ensure an adequate performance across all image modalities. Since we joined the challenge relatively late and there were many potential design choices to test, most of the hyperparameter tuning happened exclusively on the first fold of the cross-validation scheme. </p>\n<h3>Blob Regression with nnU-Net</h3>\n<p>nnU-Net is built for semantic segmentation. This also includes its expected data structure. To make it compatible with aneurysm regression we store the ground truth as semantic segmentation maps, where each aneurysm is encoded with a sphere (r=5 voxels) with an integer label representing the vessel class of the ground-truth aneurysm. These spheres are treated by nnU-Net as segmentations and are passed through the data loading and augmentation pipeline as nnU-Net normally would, thus properly applying rotations, mirroring etc, although we did not apply mirroring augmentations in the left/right axis, since several of the labels contained a left/right codification. At the end of the dataloading pipeline we inject a custom transform that converts each aneurysm instance into a blob of the respective channel. We model the 14 classes as separate channels, where the 14th class (Aneurysm Present) is the pixelwise maximum of the 13 anatomical classes.</p>\n<p>We use ‘EDT blobs’, basically 3D spheres that were transformed using the Euclidean Distance Transform (EDT) and rescaled to have a value range of [0, 1]. The optimized sphere size in the first cross-validation fold was 65 voxels. We experimented with sphere sizes from 15 to 95 voxel radii.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1460057%2F2bb5653e29b17b7558cbd3b7ac2b720e%2Fblobs.png?generation=1760607125393264&amp;alt=media\" alt=\"\"><br>\nFigure 2: Blob (EDT, radius 65) at the Right Middle Cerebral Artery (Image not resampled yet)</p>\n<h3>Hyperparameters</h3>\n<p>Our final model was trained with a batch size of 32 and a patch size of 96,x160x128 voxels. Initial learning rate is 0.01 and is decayed over the course of the training using polyLR schedule (same as default nnU-Net). We train with SGD for 3000 epochs (250 iterations per epoch). The loss function utilized is binary cross-entropy, computed only on the 20% worst voxels (the ones with the highest loss value, computed over the entire batch).</p>\n<p>Our final model was trained on 4xA100 40GB using PyTorch's DDP. Training took 4.5 days. </p>\n<h3>Inference</h3>\n<p>We largely use nnU-Net’s inference infrastructure. The input series are dissected into a series of patches. Each patch gets blob regressed. We take the maximum per channel as class probability. We max-aggregate across patches for our final predictions. <br>\nWe used the 2xT4 instances for the prediction and split the patches to predict for each input series evenly across the GPUs. We always use a single model, no ensemble. Inference takes approximately 8 hours.</p>\n<h1>Results</h1>\n<p>Unfortunately, we were hit quite hard by the instabilities of the Kaggle platform. <br>\nWhile our model finished training for 3000 epochs, only submissions until 1500 epochs (three days before the deadline) were successful (private and public LB 0.83). Subsequent submissions timed out, even though only the model weights differed.<br>\nOur internal validation showed that later checkpoints, TTA and Gaussian weighting of the patches would have probably further improved our performance (also previous submissions showed this). <br>\nSurprisingly, our internal performance went up until 0.9 which was not achieved on the leaderboard. We don’t know where this shift comes from. One reason might be that we had to embed our inference in a try/catch block, else an error was thrown after 15min. We don’t know exactly why. <br>\nWe did not exploit the segmentation masks, which could have been added as auxiliary outputs during training to help the model localize better. Due to joining late, we didn't find the time to investigate this. However, other teams showed that this improved their performance.</p>\n<h3>What did not work?</h3>\n<ul>\n<li>We also framed the problem as a detection problem, attempting to solve it with the self-configuration detection framework <a href=\"https://github.com/MIC-DKFZ/nnDetection\" target=\"_blank\">nnDetection</a>. This approach performed better than the solution presented here in the public leaderboard, but it underperformed in the private leaderboard and in our internal validation. <ul>\n<li>nnDetection was trained on instance segmentation label versions, also on cropped data, consisting of a self-configured <a href=\"https://proceedings.mlr.press/v116/jaeger20a/jaeger20a.pdf\" target=\"_blank\">Retina U-Net architecture</a> that learned from the aneurysm positions encoded as boxes and from the aneurysm segmentations . Unlike the solution here, it resampled the input series to an isometric space of [1.0, 1.0, 1.0] mm, training with a batch size of 4 for 100 epochs (2500 iterations per epoch), hybrid loss function combining L1 loss for the regression of box coordinates and focal loss for box class estimation, polynomial learning rate scheduling from an initial value of 0.001, and SGD optimizer with Nesterov momentum. Predicted boxes were postprocessed via non-maximum suppression with a 0.1 intersection over union threshold. We additionally managed to conduct inference with 8 test time augmentations and an inference patch overlap of 0.25.</li></ul></li>\n<li>Isometric space resampling with [1.0, 1.0, 1.0] mm was also implemented, given its potential for faster image processing. However, it substantially worsened our results, so it was discontinued early on.</li>\n<li>We also explored co-training with external aneurysm datasets containing binary classes to better model the Aneurysm Present class (<a href=\"https://adam.isi.uu.nl/data/\" target=\"_blank\">ADAM</a>, <a href=\"https://zenodo.org/records/6801398\" target=\"_blank\">Large IA Segmentation dataset</a>, <a href=\"https://www.codabench.org/competitions/2139/\" target=\"_blank\">INSTED</a>, <a href=\"https://openneuro.org/datasets/ds003949/versions/1.0.1\" target=\"_blank\">Lausanne TOF-MRA Aneurysm Cohort</a>, <a href=\"https://openneuro.org/datasets/ds005096/versions/1.0.3\" target=\"_blank\">Royal Brisbane TOFMRA Intracranial Aneurysm Database</a>, <a href=\"https://github.com/jinxiaokuang/RWS-MT?tab=readme-ov-file\" target=\"_blank\">Jianxiaokuang aneurysm dataset</a>. We realized afterwards that some of these datasets (<a href=\"https://adam.isi.uu.nl/data/\" target=\"_blank\">ADAM</a>,  <a href=\"https://zenodo.org/records/6801398\" target=\"_blank\">Large IA Segmentation dataset</a>) were not allowed, so we discarded them and ran co-training without them. In the end, co-training did not really help, so we resorted back to training only on the challenge cases from scratch. </li>\n<li>We first started with processing the image as a whole but time limitations forced us to crop around the ROI which also resulted in better performance.</li>\n<li>As described above, in our final model we just max-aggregate the patch predictions. However, it is known that the model has some uncertainty close to the edges. A common strategy to mitigate this is gaussian weighting the predictions for each patch (high weight in the center, low weight and the borders). In earlier submissions we saw that this improved our performance. However, in our final model we could not apply this strategy because of platform instabilities. </li>\n<li>We also tried to train with larger patch sizes, which, surprisingly, did not contribute to improve our scores.</li>\n</ul>\n<h3>What would we have wished for?</h3>\n<p>We already stated in the discussion forum that the signature of the predict function was limiting us quite a lot in how we can parallelize processing. More about that <a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/611778\" target=\"_blank\">here</a>. <br>\nThis requirement made it even more difficult for us since we required a specific numpy version which forced us to run our code as a subprocess. We ended up spawning a proxy worker with which we communicated via the std output. This added significant boilerplate and complexity, and felt unnatural.</p>\n<p>We would have also wished for longer submission times, and a less dependent server architecture on the number of submissions sent by different teams, since many of our submissions timed out during the last few days of the challenge due to an increasing workload. <br>\nOn a much broader scope: Kaggle's submission notebook style made it quite hard for us (also the last time). Having the possibility to just use Docker containers would have eased our lives a lot because they allow much higher flexibility.  </p>\n<h1>Acknowledgements</h1>\n<p>We thank RSNA for organizing and Kaggle for hosting this competition. We furthermore want to give a shoutout to our <a href=\"https://www.dkfz.de/en/medical-image-computing\" target=\"_blank\">Divisions of Medical Image Computing</a> at the German Cancer Research Center, as well as <a href=\"https://helmholtz-imaging.de/\" target=\"_blank\">Helmholtz Imaging</a></p>",
      "rawMarkdown": "Thanks to @evancalabrese, @ryanholbrook, RSNA, and Kaggle for organizing this Intracranial Aneurysm Detection competition! \n# Overview (TLDR)\nHere’s a brief rundown of our solution — it’s straightforward and easy to implement:\n- We formulate the task as multichannel blob regression, optimized using a TopK (20%) BCE loss and then taking the maximum per channel as probability prediction.\n- We build on [nnU-Net](https://github.com/MIC-DKFZ/nnUNet), the leading framework for 3D medical image segmentation. We already adapted it for our [2nd place solution in the BYU - Locating Bacterial Flagellar Motors 2025](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/mic-dkfz-2nd-place-solution-3d-nnu-net-blob-regres).\n- Our model is a 3D U-Net with a residual encoder, trained from scratch.\n- Inference is done with a single model without test-time augmentation (due to time restrictions).\n- Our model achieved a score of 0.83 / 0.83 on the public/private leaderboard.\n\n# Who are we?\nWe are a team of colleagues (scientists and PhD students) affiliated with the [Divisions of Medical Image Computing](https://www.dkfz.de/en/medical-image-computing) at the German Cancer Research Center, as well as [Helmholtz Imaging](https://helmholtz-imaging.de/). Our expertise lies in 3D image analysis — particularly in solving 3D segmentation problems and developing infrastructure to bring algorithms into clinical practice. \n\n\n# Method\nWe modeled the task as a heatmap regression and built up on our [2nd place solution in the BYU - Locating Bacterial Flagellar Motors 2025](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/mic-dkfz-2nd-place-solution-3d-nnu-net-blob-regres). \n\n### Data used\nWe converted the DICOM series to .nii.gz format using the pydicom library, processing series with 2D DICOM files in parallel. We also used this library to extract information on spacing, origin, and direction. We then converted the resulting image into a SimpleITK image and oriented it in RAS orientation.\n\nTo reduce computing resources, we derived a [200, 160, 160] mm cubic Region of Interest (ROI) on the central superior region of the image. We ensured that all aneurysms in the training set were included in the ROI. Fig. 1 depicts the ROI on top of the image.\n\nWe observed that several series presented defects such as unexpected orientations, empty series, shunt artifacts, movement artifacts, or images with an empty superior space. We decided to keep those series with mild artifacts such as mis-orientations and fixed them via flipping. Series with stronger artifacts such as totally empty images were discarded. In total, ten volumes were discarded.\n\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1460057%2Faa77cf1f2c0d097a505297609efdd7b4%2Fbbox.png?generation=1760605343255626&alt=media)\nFigure 1. Axial, coronal and sagittal views of example image series and the ROI box, in red.\n\n### Preprocessing\nData preprocessing was completed with the self-configurable segmentation framework nnU-Net, following a 3D full resolution configuration. The volumes were loaded with a SimpleITK reader that also tries to enforce RAS orientation. All images were resampled to the median spacing found from all training images: [0.70, 0.47, 0.47] (mm), being normalized via z-score normalization based on a global mean and a global standard deviation extracted from the training dataset. \n\nThe images were resampled with a special resampler from PyTorch, which is faster than other commonly used functions such as scipy.ndimage.zoom, given the strong time constraints.\n\n### Network architecture\nWe use nnU-Net’s ResEnc, which is essentially a UNet with a residual encoder and a lightweight convolutional decoder. The architecture included six stages with [32, 64, 128, 256, 320, 320] features in each stage, respectively. \n\n### Training Procedure\nWe split the provided challenge data into five cross-validation folds, stratifying for modalities across folds, rather than on vessel classes, since we wanted to ensure an adequate performance across all image modalities. Since we joined the challenge relatively late and there were many potential design choices to test, most of the hyperparameter tuning happened exclusively on the first fold of the cross-validation scheme. \n\n### Blob Regression with nnU-Net\n\nnnU-Net is built for semantic segmentation. This also includes its expected data structure. To make it compatible with aneurysm regression we store the ground truth as semantic segmentation maps, where each aneurysm is encoded with a sphere (r=5 voxels) with an integer label representing the vessel class of the ground-truth aneurysm. These spheres are treated by nnU-Net as segmentations and are passed through the data loading and augmentation pipeline as nnU-Net normally would, thus properly applying rotations, mirroring etc, although we did not apply mirroring augmentations in the left/right axis, since several of the labels contained a left/right codification. At the end of the dataloading pipeline we inject a custom transform that converts each aneurysm instance into a blob of the respective channel. We model the 14 classes as separate channels, where the 14th class (Aneurysm Present) is the pixelwise maximum of the 13 anatomical classes.\n\nWe use ‘EDT blobs’, basically 3D spheres that were transformed using the Euclidean Distance Transform (EDT) and rescaled to have a value range of [0, 1]. The optimized sphere size in the first cross-validation fold was 65 voxels. We experimented with sphere sizes from 15 to 95 voxel radii.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1460057%2F2bb5653e29b17b7558cbd3b7ac2b720e%2Fblobs.png?generation=1760607125393264&alt=media)\nFigure 2: Blob (EDT, radius 65) at the Right Middle Cerebral Artery (Image not resampled yet)\n\n### Hyperparameters\nOur final model was trained with a batch size of 32 and a patch size of 96,x160x128 voxels. Initial learning rate is 0.01 and is decayed over the course of the training using polyLR schedule (same as default nnU-Net). We train with SGD for 3000 epochs (250 iterations per epoch). The loss function utilized is binary cross-entropy, computed only on the 20% worst voxels (the ones with the highest loss value, computed over the entire batch).\n\nOur final model was trained on 4xA100 40GB using PyTorch's DDP. Training took 4.5 days. \n\n### Inference\nWe largely use nnU-Net’s inference infrastructure. The input series are dissected into a series of patches. Each patch gets blob regressed. We take the maximum per channel as class probability. We max-aggregate across patches for our final predictions. \nWe used the 2xT4 instances for the prediction and split the patches to predict for each input series evenly across the GPUs. We always use a single model, no ensemble. Inference takes approximately 8 hours.\n\n# Results\nUnfortunately, we were hit quite hard by the instabilities of the Kaggle platform. \nWhile our model finished training for 3000 epochs, only submissions until 1500 epochs (three days before the deadline) were successful (private and public LB 0.83). Subsequent submissions timed out, even though only the model weights differed.\nOur internal validation showed that later checkpoints, TTA and Gaussian weighting of the patches would have probably further improved our performance (also previous submissions showed this). \nSurprisingly, our internal performance went up until 0.9 which was not achieved on the leaderboard. We don’t know where this shift comes from. One reason might be that we had to embed our inference in a try/catch block, else an error was thrown after 15min. We don’t know exactly why. \nWe did not exploit the segmentation masks, which could have been added as auxiliary outputs during training to help the model localize better. Due to joining late, we didn't find the time to investigate this. However, other teams showed that this improved their performance.\n\n### What did not work?\n- We also framed the problem as a detection problem, attempting to solve it with the self-configuration detection framework [nnDetection](https://github.com/MIC-DKFZ/nnDetection). This approach performed better than the solution presented here in the public leaderboard, but it underperformed in the private leaderboard and in our internal validation. \n       - nnDetection was trained on instance segmentation label versions, also on cropped data, consisting of a self-configured [Retina U-Net architecture] (https://proceedings.mlr.press/v116/jaeger20a/jaeger20a.pdf) that learned from the aneurysm positions encoded as boxes and from the aneurysm segmentations . Unlike the solution here, it resampled the input series to an isometric space of [1.0, 1.0, 1.0] mm, training with a batch size of 4 for 100 epochs (2500 iterations per epoch), hybrid loss function combining L1 loss for the regression of box coordinates and focal loss for box class estimation, polynomial learning rate scheduling from an initial value of 0.001, and SGD optimizer with Nesterov momentum. Predicted boxes were postprocessed via non-maximum suppression with a 0.1 intersection over union threshold. We additionally managed to conduct inference with 8 test time augmentations and an inference patch overlap of 0.25.\n- Isometric space resampling with [1.0, 1.0, 1.0] mm was also implemented, given its potential for faster image processing. However, it substantially worsened our results, so it was discontinued early on.\n- We also explored co-training with external aneurysm datasets containing binary classes to better model the Aneurysm Present class ([ADAM](https://adam.isi.uu.nl/data/), [Large IA Segmentation dataset](https://zenodo.org/records/6801398), [INSTED](https://www.codabench.org/competitions/2139/), [Lausanne TOF-MRA Aneurysm Cohort](https://openneuro.org/datasets/ds003949/versions/1.0.1), [Royal Brisbane TOFMRA Intracranial Aneurysm Database](https://openneuro.org/datasets/ds005096/versions/1.0.3), [Jianxiaokuang aneurysm dataset](https://github.com/jinxiaokuang/RWS-MT?tab=readme-ov-file). We realized afterwards that some of these datasets ([ADAM](https://adam.isi.uu.nl/data/),  [Large IA Segmentation dataset](https://zenodo.org/records/6801398)) were not allowed, so we discarded them and ran co-training without them. In the end, co-training did not really help, so we resorted back to training only on the challenge cases from scratch. \n- We first started with processing the image as a whole but time limitations forced us to crop around the ROI which also resulted in better performance.\n- As described above, in our final model we just max-aggregate the patch predictions. However, it is known that the model has some uncertainty close to the edges. A common strategy to mitigate this is gaussian weighting the predictions for each patch (high weight in the center, low weight and the borders). In earlier submissions we saw that this improved our performance. However, in our final model we could not apply this strategy because of platform instabilities. \n- We also tried to train with larger patch sizes, which, surprisingly, did not contribute to improve our scores.\n\n### What would we have wished for?\nWe already stated in the discussion forum that the signature of the predict function was limiting us quite a lot in how we can parallelize processing. More about that [here](https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/611778). \nThis requirement made it even more difficult for us since we required a specific numpy version which forced us to run our code as a subprocess. We ended up spawning a proxy worker with which we communicated via the std output. This added significant boilerplate and complexity, and felt unnatural.\n\nWe would have also wished for longer submission times, and a less dependent server architecture on the number of submissions sent by different teams, since many of our submissions timed out during the last few days of the challenge due to an increasing workload. \nOn a much broader scope: Kaggle's submission notebook style made it quite hard for us (also the last time). Having the possibility to just use Docker containers would have eased our lives a lot because they allow much higher flexibility.  \n\n\n# Acknowledgements\nWe thank RSNA for organizing and Kaggle for hosting this competition. We furthermore want to give a shoutout to our [Divisions of Medical Image Computing](https://www.dkfz.de/en/medical-image-computing) at the German Cancer Research Center, as well as [Helmholtz Imaging](https://helmholtz-imaging.de/)",
      "votes": 24
    },
    {
      "id": 3303326,
      "postDate": "2025-10-17T17:14:51.337Z",
      "content": "<p>Thanks for sharing the informative writeup very much.</p>",
      "rawMarkdown": "Thanks for sharing the informative writeup very much."
    },
    {
      "id": 3302943,
      "postDate": "2025-10-16T20:19:15.397Z",
      "content": "<p>Thanks for sharing the informative writeup very much. Your team is really a symbol of elegant challenge solutions! It's impressive that a single model can achieve such high performance even if it is not fully optimized. </p>\n<p>It would be highly appreciated if you could share answers to the following questions. </p>\n<ol>\n<li>How to automatically derive the [200, 160, 160] ROI?</li>\n<li>How to use the images with AP for training? The corresponding mask will be full zero. </li>\n<li>Is there any design to address the class imbalance?</li>\n<li>TTA was not used in inference, right?</li>\n<li>How different sphere size affect the performance based on the fold-0 split?</li>\n<li>Do you have any plans to release the code? This task is much more challenging than the <a href=\"https://github.com/MIC-DKFZ/kaggle_BYU_Locating_Bacterial-Flagellar_Motors_2025_solution\" target=\"_blank\">BYU task</a>. Looking forward to diving into the implementation details. </li>\n</ol>",
      "rawMarkdown": "Thanks for sharing the informative writeup very much. Your team is really a symbol of elegant challenge solutions! It's impressive that a single model can achieve such high performance even if it is not fully optimized. \n\nIt would be highly appreciated if you could share answers to the following questions. \n1. How to automatically derive the [200, 160, 160] ROI?\n2. How to use the images with AP for training? The corresponding mask will be full zero. \n3. Is there any design to address the class imbalance?\n4. TTA was not used in inference, right?\n5. How different sphere size affect the performance based on the fold-0 split?\n6. Do you have any plans to release the code? This task is much more challenging than the [BYU task](https://github.com/MIC-DKFZ/kaggle_BYU_Locating_Bacterial-Flagellar_Motors_2025_solution). Looking forward to diving into the implementation details. \n\n \n\n",
      "replies": [
        {
          "id": 3302956,
          "postDate": "2025-10-16T20:50:33.427Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/maedward\" target=\"_blank\">@maedward</a>,<br>\nthanks for the kind words. :) </p>\n<ol>\n<li>How to automatically derive the [200, 160, 160] ROI?<br>\n-&gt; It is really as simple as we write it. We just take a center top crop of the image with the respective ROI size in mm. </li>\n<li>How to use the images with AP for training? The corresponding mask will be full zero.<br>\n-&gt; you mean without AP, right? Yes, the target heat map is zero-only. There is no issue with that. </li>\n<li>Is there any design to address the class imbalance?<br>\n-&gt; We thought of it. We also implemented class weighting and low prevalence class oversampling in our nnDetection efforts but there it did not help. For nnUNet we have not tried it out. But is probably worth a shot because we also noticed that one particular class was always achieving rather low performance. However, we also noticed that some other classes which also had low prevalence were achieving quite reasonable performance so it is not only about the class prevalence. </li>\n<li>TTA was not used in inference, right?<br>\n-&gt; Correct. With stepsize 0.5 and median spacing inference took to long with TTA.</li>\n<li>How different sphere size affect the performance based on the fold-0 split?<br>\n-&gt; EDT25  0.888; EDT35  0.876, EDT45  0.893; EDT55  0.883; EDT65  0.896; EDT75  0.871; EDT85  0.871; EDT95  0.866</li>\n<li>Do you have any plans to release the code? This task is much more challenging than the BYU task. Looking forward to diving into the implementation details<br>\n-&gt; Yes, we will release it next week :) </li>\n</ol>",
          "rawMarkdown": "Hi @maedward,\nthanks for the kind words. :) \n1. How to automatically derive the [200, 160, 160] ROI?\n-> It is really as simple as we write it. We just take a center top crop of the image with the respective ROI size in mm. \n2. How to use the images with AP for training? The corresponding mask will be full zero.\n-> you mean without AP, right? Yes, the target heat map is zero-only. There is no issue with that. \n3. Is there any design to address the class imbalance?\n-> We thought of it. We also implemented class weighting and low prevalence class oversampling in our nnDetection efforts but there it did not help. For nnUNet we have not tried it out. But is probably worth a shot because we also noticed that one particular class was always achieving rather low performance. However, we also noticed that some other classes which also had low prevalence were achieving quite reasonable performance so it is not only about the class prevalence. \n4. TTA was not used in inference, right?\n-> Correct. With stepsize 0.5 and median spacing inference took to long with TTA.\n5. How different sphere size affect the performance based on the fold-0 split?\n-> EDT25  0.888; EDT35  0.876, EDT45  0.893; EDT55  0.883; EDT65  0.896; EDT75  0.871; EDT85  0.871; EDT95  0.866\n6. Do you have any plans to release the code? This task is much more challenging than the BYU task. Looking forward to diving into the implementation details\n-> Yes, we will release it next week :) ",
          "votes": 1,
          "replies": [
            {
              "id": 3303321,
              "postDate": "2025-10-17T17:05:37.373Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/st3v3d\" target=\"_blank\">@st3v3d</a> , thank you so much for the detailed and insightful reply. Looking forward to the code next week</p>",
              "rawMarkdown": "Hi @st3v3d , thank you so much for the detailed and insightful reply. Looking forward to the code next week"
            },
            {
              "id": 3304770,
              "postDate": "2025-10-21T10:06:45.407Z",
              "content": "<p>We added the resources to our writeup :) </p>",
              "rawMarkdown": "We added the resources to our writeup :) ",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3303326,
      "author_name": "Khushi Yadav",
      "author_url": "",
      "post_date": "2025-10-17T17:14:51.337000",
      "content": "<p>Thanks for sharing the informative writeup very much.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3302943,
      "author_name": "Ma Edward",
      "author_url": "",
      "post_date": "2025-10-16T20:19:15.397000",
      "content": "<p>Thanks for sharing the informative writeup very much. Your team is really a symbol of elegant challenge solutions! It's impressive that a single model can achieve such high performance even if it is not fully optimized. </p>\n<p>It would be highly appreciated if you could share answers to the following questions. </p>\n<ol>\n<li>How to automatically derive the [200, 160, 160] ROI?</li>\n<li>How to use the images with AP for training? The corresponding mask will be full zero. </li>\n<li>Is there any design to address the class imbalance?</li>\n<li>TTA was not used in inference, right?</li>\n<li>How different sphere size affect the performance based on the fold-0 split?</li>\n<li>Do you have any plans to release the code? This task is much more challenging than the <a href=\"https://github.com/MIC-DKFZ/kaggle_BYU_Locating_Bacterial-Flagellar_Motors_2025_solution\" target=\"_blank\">BYU task</a>. Looking forward to diving into the implementation details. </li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 3302956,
          "author_name": "Stefan Denner",
          "author_url": "",
          "post_date": "2025-10-16T20:50:33.427000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/maedward\" target=\"_blank\">@maedward</a>,<br>\nthanks for the kind words. :) </p>\n<ol>\n<li>How to automatically derive the [200, 160, 160] ROI?<br>\n-&gt; It is really as simple as we write it. We just take a center top crop of the image with the respective ROI size in mm. </li>\n<li>How to use the images with AP for training? The corresponding mask will be full zero.<br>\n-&gt; you mean without AP, right? Yes, the target heat map is zero-only. There is no issue with that. </li>\n<li>Is there any design to address the class imbalance?<br>\n-&gt; We thought of it. We also implemented class weighting and low prevalence class oversampling in our nnDetection efforts but there it did not help. For nnUNet we have not tried it out. But is probably worth a shot because we also noticed that one particular class was always achieving rather low performance. However, we also noticed that some other classes which also had low prevalence were achieving quite reasonable performance so it is not only about the class prevalence. </li>\n<li>TTA was not used in inference, right?<br>\n-&gt; Correct. With stepsize 0.5 and median spacing inference took to long with TTA.</li>\n<li>How different sphere size affect the performance based on the fold-0 split?<br>\n-&gt; EDT25  0.888; EDT35  0.876, EDT45  0.893; EDT55  0.883; EDT65  0.896; EDT75  0.871; EDT85  0.871; EDT95  0.866</li>\n<li>Do you have any plans to release the code? This task is much more challenging than the BYU task. Looking forward to diving into the implementation details<br>\n-&gt; Yes, we will release it next week :) </li>\n</ol>",
          "votes": 1,
          "replies": [
            {
              "id": 3303321,
              "author_name": "Ma Edward",
              "author_url": "",
              "post_date": "2025-10-17T17:05:37.373000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/st3v3d\" target=\"_blank\">@st3v3d</a> , thank you so much for the detailed and insightful reply. Looking forward to the code next week</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3304770,
              "author_name": "Stefan Denner",
              "author_url": "",
              "post_date": "2025-10-21T10:06:45.407000",
              "content": "<p>We added the resources to our writeup :) </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3302691": "Thanks to @evancalabrese, @ryanholbrook, RSNA, and Kaggle for organizing this Intracranial Aneurysm Detection competition! \n# Overview (TLDR)\nHere’s a brief rundown of our solution — it’s straightforward and easy to implement:\n- We formulate the task as multichannel blob regression, optimized using a TopK (20%) BCE loss and then taking the maximum per channel as probability prediction.\n- We build on [nnU-Net](https://github.com/MIC-DKFZ/nnUNet), the leading framework for 3D medical image segmentation. We already adapted it for our [2nd place solution in the BYU - Locating Bacterial Flagellar Motors 2025](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/mic-dkfz-2nd-place-solution-3d-nnu-net-blob-regres).\n- Our model is a 3D U-Net with a residual encoder, trained from scratch.\n- Inference is done with a single model without test-time augmentation (due to time restrictions).\n- Our model achieved a score of 0.83 / 0.83 on the public/private leaderboard.\n\n# Who are we?\nWe are a team of colleagues (scientists and PhD students) affiliated with the [Divisions of Medical Image Computing](https://www.dkfz.de/en/medical-image-computing) at the German Cancer Research Center, as well as [Helmholtz Imaging](https://helmholtz-imaging.de/). Our expertise lies in 3D image analysis — particularly in solving 3D segmentation problems and developing infrastructure to bring algorithms into clinical practice. \n\n\n# Method\nWe modeled the task as a heatmap regression and built up on our [2nd place solution in the BYU - Locating Bacterial Flagellar Motors 2025](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/mic-dkfz-2nd-place-solution-3d-nnu-net-blob-regres). \n\n### Data used\nWe converted the DICOM series to .nii.gz format using the pydicom library, processing series with 2D DICOM files in parallel. We also used this library to extract information on spacing, origin, and direction. We then converted the resulting image into a SimpleITK image and oriented it in RAS orientation.\n\nTo reduce computing resources, we derived a [200, 160, 160] mm cubic Region of Interest (ROI) on the central superior region of the image. We ensured that all aneurysms in the training set were included in the ROI. Fig. 1 depicts the ROI on top of the image.\n\nWe observed that several series presented defects such as unexpected orientations, empty series, shunt artifacts, movement artifacts, or images with an empty superior space. We decided to keep those series with mild artifacts such as mis-orientations and fixed them via flipping. Series with stronger artifacts such as totally empty images were discarded. In total, ten volumes were discarded.\n\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1460057%2Faa77cf1f2c0d097a505297609efdd7b4%2Fbbox.png?generation=1760605343255626&alt=media)\nFigure 1. Axial, coronal and sagittal views of example image series and the ROI box, in red.\n\n### Preprocessing\nData preprocessing was completed with the self-configurable segmentation framework nnU-Net, following a 3D full resolution configuration. The volumes were loaded with a SimpleITK reader that also tries to enforce RAS orientation. All images were resampled to the median spacing found from all training images: [0.70, 0.47, 0.47] (mm), being normalized via z-score normalization based on a global mean and a global standard deviation extracted from the training dataset. \n\nThe images were resampled with a special resampler from PyTorch, which is faster than other commonly used functions such as scipy.ndimage.zoom, given the strong time constraints.\n\n### Network architecture\nWe use nnU-Net’s ResEnc, which is essentially a UNet with a residual encoder and a lightweight convolutional decoder. The architecture included six stages with [32, 64, 128, 256, 320, 320] features in each stage, respectively. \n\n### Training Procedure\nWe split the provided challenge data into five cross-validation folds, stratifying for modalities across folds, rather than on vessel classes, since we wanted to ensure an adequate performance across all image modalities. Since we joined the challenge relatively late and there were many potential design choices to test, most of the hyperparameter tuning happened exclusively on the first fold of the cross-validation scheme. \n\n### Blob Regression with nnU-Net\n\nnnU-Net is built for semantic segmentation. This also includes its expected data structure. To make it compatible with aneurysm regression we store the ground truth as semantic segmentation maps, where each aneurysm is encoded with a sphere (r=5 voxels) with an integer label representing the vessel class of the ground-truth aneurysm. These spheres are treated by nnU-Net as segmentations and are passed through the data loading and augmentation pipeline as nnU-Net normally would, thus properly applying rotations, mirroring etc, although we did not apply mirroring augmentations in the left/right axis, since several of the labels contained a left/right codification. At the end of the dataloading pipeline we inject a custom transform that converts each aneurysm instance into a blob of the respective channel. We model the 14 classes as separate channels, where the 14th class (Aneurysm Present) is the pixelwise maximum of the 13 anatomical classes.\n\nWe use ‘EDT blobs’, basically 3D spheres that were transformed using the Euclidean Distance Transform (EDT) and rescaled to have a value range of [0, 1]. The optimized sphere size in the first cross-validation fold was 65 voxels. We experimented with sphere sizes from 15 to 95 voxel radii.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1460057%2F2bb5653e29b17b7558cbd3b7ac2b720e%2Fblobs.png?generation=1760607125393264&alt=media)\nFigure 2: Blob (EDT, radius 65) at the Right Middle Cerebral Artery (Image not resampled yet)\n\n### Hyperparameters\nOur final model was trained with a batch size of 32 and a patch size of 96,x160x128 voxels. Initial learning rate is 0.01 and is decayed over the course of the training using polyLR schedule (same as default nnU-Net). We train with SGD for 3000 epochs (250 iterations per epoch). The loss function utilized is binary cross-entropy, computed only on the 20% worst voxels (the ones with the highest loss value, computed over the entire batch).\n\nOur final model was trained on 4xA100 40GB using PyTorch's DDP. Training took 4.5 days. \n\n### Inference\nWe largely use nnU-Net’s inference infrastructure. The input series are dissected into a series of patches. Each patch gets blob regressed. We take the maximum per channel as class probability. We max-aggregate across patches for our final predictions. \nWe used the 2xT4 instances for the prediction and split the patches to predict for each input series evenly across the GPUs. We always use a single model, no ensemble. Inference takes approximately 8 hours.\n\n# Results\nUnfortunately, we were hit quite hard by the instabilities of the Kaggle platform. \nWhile our model finished training for 3000 epochs, only submissions until 1500 epochs (three days before the deadline) were successful (private and public LB 0.83). Subsequent submissions timed out, even though only the model weights differed.\nOur internal validation showed that later checkpoints, TTA and Gaussian weighting of the patches would have probably further improved our performance (also previous submissions showed this). \nSurprisingly, our internal performance went up until 0.9 which was not achieved on the leaderboard. We don’t know where this shift comes from. One reason might be that we had to embed our inference in a try/catch block, else an error was thrown after 15min. We don’t know exactly why. \nWe did not exploit the segmentation masks, which could have been added as auxiliary outputs during training to help the model localize better. Due to joining late, we didn't find the time to investigate this. However, other teams showed that this improved their performance.\n\n### What did not work?\n- We also framed the problem as a detection problem, attempting to solve it with the self-configuration detection framework [nnDetection](https://github.com/MIC-DKFZ/nnDetection). This approach performed better than the solution presented here in the public leaderboard, but it underperformed in the private leaderboard and in our internal validation. \n       - nnDetection was trained on instance segmentation label versions, also on cropped data, consisting of a self-configured [Retina U-Net architecture] (https://proceedings.mlr.press/v116/jaeger20a/jaeger20a.pdf) that learned from the aneurysm positions encoded as boxes and from the aneurysm segmentations . Unlike the solution here, it resampled the input series to an isometric space of [1.0, 1.0, 1.0] mm, training with a batch size of 4 for 100 epochs (2500 iterations per epoch), hybrid loss function combining L1 loss for the regression of box coordinates and focal loss for box class estimation, polynomial learning rate scheduling from an initial value of 0.001, and SGD optimizer with Nesterov momentum. Predicted boxes were postprocessed via non-maximum suppression with a 0.1 intersection over union threshold. We additionally managed to conduct inference with 8 test time augmentations and an inference patch overlap of 0.25.\n- Isometric space resampling with [1.0, 1.0, 1.0] mm was also implemented, given its potential for faster image processing. However, it substantially worsened our results, so it was discontinued early on.\n- We also explored co-training with external aneurysm datasets containing binary classes to better model the Aneurysm Present class ([ADAM](https://adam.isi.uu.nl/data/), [Large IA Segmentation dataset](https://zenodo.org/records/6801398), [INSTED](https://www.codabench.org/competitions/2139/), [Lausanne TOF-MRA Aneurysm Cohort](https://openneuro.org/datasets/ds003949/versions/1.0.1), [Royal Brisbane TOFMRA Intracranial Aneurysm Database](https://openneuro.org/datasets/ds005096/versions/1.0.3), [Jianxiaokuang aneurysm dataset](https://github.com/jinxiaokuang/RWS-MT?tab=readme-ov-file). We realized afterwards that some of these datasets ([ADAM](https://adam.isi.uu.nl/data/),  [Large IA Segmentation dataset](https://zenodo.org/records/6801398)) were not allowed, so we discarded them and ran co-training without them. In the end, co-training did not really help, so we resorted back to training only on the challenge cases from scratch. \n- We first started with processing the image as a whole but time limitations forced us to crop around the ROI which also resulted in better performance.\n- As described above, in our final model we just max-aggregate the patch predictions. However, it is known that the model has some uncertainty close to the edges. A common strategy to mitigate this is gaussian weighting the predictions for each patch (high weight in the center, low weight and the borders). In earlier submissions we saw that this improved our performance. However, in our final model we could not apply this strategy because of platform instabilities. \n- We also tried to train with larger patch sizes, which, surprisingly, did not contribute to improve our scores.\n\n### What would we have wished for?\nWe already stated in the discussion forum that the signature of the predict function was limiting us quite a lot in how we can parallelize processing. More about that [here](https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/611778). \nThis requirement made it even more difficult for us since we required a specific numpy version which forced us to run our code as a subprocess. We ended up spawning a proxy worker with which we communicated via the std output. This added significant boilerplate and complexity, and felt unnatural.\n\nWe would have also wished for longer submission times, and a less dependent server architecture on the number of submissions sent by different teams, since many of our submissions timed out during the last few days of the challenge due to an increasing workload. \nOn a much broader scope: Kaggle's submission notebook style made it quite hard for us (also the last time). Having the possibility to just use Docker containers would have eased our lives a lot because they allow much higher flexibility.  \n\n\n# Acknowledgements\nWe thank RSNA for organizing and Kaggle for hosting this competition. We furthermore want to give a shoutout to our [Divisions of Medical Image Computing](https://www.dkfz.de/en/medical-image-computing) at the German Cancer Research Center, as well as [Helmholtz Imaging](https://helmholtz-imaging.de/)",
    "3303326": "Thanks for sharing the informative writeup very much.",
    "3302943": "Thanks for sharing the informative writeup very much. Your team is really a symbol of elegant challenge solutions! It's impressive that a single model can achieve such high performance even if it is not fully optimized. \n\nIt would be highly appreciated if you could share answers to the following questions. \n1. How to automatically derive the [200, 160, 160] ROI?\n2. How to use the images with AP for training? The corresponding mask will be full zero. \n3. Is there any design to address the class imbalance?\n4. TTA was not used in inference, right?\n5. How different sphere size affect the performance based on the fold-0 split?\n6. Do you have any plans to release the code? This task is much more challenging than the [BYU task](https://github.com/MIC-DKFZ/kaggle_BYU_Locating_Bacterial-Flagellar_Motors_2025_solution). Looking forward to diving into the implementation details. \n\n \n\n"
  }
}