{
  "id": 611856,
  "title": "3rd place solution",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/611856",
  "author_name": "TmT",
  "post_date": "2025-10-15T06:45:49.074000",
  "votes": 46,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Thanks to RSNA and Kaggle for hosting this competition—it was a great opportunity to work with real-world medical data. I’m also grateful to my teammates <a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a>, <a href=\"https://www.kaggle.com/nmstopen\" target=\"_blank\">@nmstopen</a>, <a href=\"https://www.kaggle.com/heroalchem\" target=\"_blank\">@heroalchem</a>, and <a href=\"https://www.kaggle.com/dainagao\" target=\"_blank\">@dainagao</a>; our collaboration directly contributed to the results we achieved.</p>\n<h3>Overview</h3>\n<p>Our main solution consists of two stages:</p>\n<ol>\n<li>3D vessel region detection</li>\n<li>3D ROI classification</li>\n</ol>\n<h3>Stage 1: Vessel Region Detection</h3>\n<p>Because the target vessels occupy only a limited portion of the field of view, we first detect whole vessel regions before downstream analysis. We also experimented with vessel segmentation, but it was not sufficiently robust across cases.<br>\nAneurysm locations were relatively consistent in XY coordinates on the axial plane across cases, so we used the middle slice from the sagittal and coronal planes as input images for detection. We computed MIPs of the segmentation masks along the sagittal and coronal directions, then constructed 2D, axis-aligned bounding boxes by taking the minimum and maximum mask coordinates in each view. We used YOLOv8n and YOLOv8m for detection, achieving over 0.95 mAP@0.5 on the validation set. After detection, we reconstructed a 3D ROI by combining results from the sagittal and coronal views, and then cropped a fixed-size 3D bounding box of 90×90×90 mm (or 120×120×120 mm) centered on each detection to generate analysis patches.<br>\nExamples of detection results. Green = prediction; red = ground truth.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1408427%2F979857f3afb1fd7d49b4425c121b4ced%2Fyolo.png?generation=1760510088174695&amp;alt=media\" alt=\"\"></p>\n<h3>Stage 2: 3D ROI Classification</h3>\n<h4>Training</h4>\n<p>Using the 3D vessel ROIs from Stage 1, we trained 3D ResNet-18 backbones, implemented with the <code>timm-3d</code> library.</p>\n<p>Our models are inspired by <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/daddies-4th-place-simple-resnet18-classification\" target=\"_blank\">BYU 4th place solution</a>. <br>\nWe attached a 14-class classification head to <strong>each feature map (not whole volume)</strong> and optimized with weighted BCE loss (multi-label setting). In the default 3D configuration (input 128×128×128), the network produced a 4×4×4 feature map, which did not yield good results. Increasing spatial resolution helped: we changed the stride from 2 to 1 in selected convolution layers to obtain larger feature maps.</p>\n<p>We explored the input volume from 128×128×128 up to 224×224×224, and feature-map sizes from 8×8×8 to 48×48×48. Notably, increasing the feature map from 8×8×8 to 25×25×25 and the image size from 128×128×128 to 196×196×196 improved the LB score from 0.77 to 0.81 in a single-fold model.</p>\n<h4>Inference &amp; Aggregation</h4>\n<p>For inference, we aggregated feature-map predictions into a per-case prediction. Specifically, for each class we sorted the <code>Aneurysm Present</code> scores across spatial positions and averaged the top-N scores (Top-N mean) to produce the final class prediction. N depends on model configuration.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1408427%2F886c5768338572f3a261ab74e3881e9f%2Fpipeline.png?generation=1760510113790199&amp;alt=media\" alt=\"\"></p>\n<h4>Model ensemble</h4>\n<p>We built an ensemble of 11 models with different crop sizes, image resolutions, and variants with reduced stride in selected convolution layers. All models used a 3D ResNet-18 backbone. In this ensemble, the public/private LB scores were 0.86/0.84.</p>\n<h3>Missing DICOM tags &amp; Fallbacks</h3>\n<p>There were many missing DICOM tags in the test data, so we built a model to estimate voxel spacing along the X, Y, and Z axes. To preserve XY spacing, each slice was padded and center-cropped to 512×512 pixels, and 10 central slices were sampled along the Z-axis. Each slice was stacked with its adjacent slices to form 2.5D RGB inputs for capturing Z-axis information. Using these inputs, an EfficientNet V2 S regression model predicted voxel spacing from image appearance, with final values obtained by averaging outputs across the 10 slices. On internal validation, the model achieved MAE = 0.015 mm (X), 0.020 mm (Y), and 0.071 mm (Z), improving our private LB from 0.84 to 0.85.</p>\n<h3>What didn't work for us</h3>\n<ul>\n<li>Vessel segmentation</li>\n<li>Modality and Plane heads for auxiliary loss</li>\n<li>Complicated model like MIL, LSTMs and decorder heads</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/code/tamotamo/rsna2025-3rd-place-inference\" target=\"_blank\">Inference code</a><br>\n<a href=\"https://github.com/tamotamo17/RSNA2025-3rd-place-solution\" target=\"_blank\">Training code</a></p>",
  "messages": [
    {
      "id": 3302141,
      "postDate": "2025-10-15T06:45:49.073Z",
      "content": "<p>Thanks to RSNA and Kaggle for hosting this competition—it was a great opportunity to work with real-world medical data. I’m also grateful to my teammates <a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a>, <a href=\"https://www.kaggle.com/nmstopen\" target=\"_blank\">@nmstopen</a>, <a href=\"https://www.kaggle.com/heroalchem\" target=\"_blank\">@heroalchem</a>, and <a href=\"https://www.kaggle.com/dainagao\" target=\"_blank\">@dainagao</a>; our collaboration directly contributed to the results we achieved.</p>\n<h3>Overview</h3>\n<p>Our main solution consists of two stages:</p>\n<ol>\n<li>3D vessel region detection</li>\n<li>3D ROI classification</li>\n</ol>\n<h3>Stage 1: Vessel Region Detection</h3>\n<p>Because the target vessels occupy only a limited portion of the field of view, we first detect whole vessel regions before downstream analysis. We also experimented with vessel segmentation, but it was not sufficiently robust across cases.<br>\nAneurysm locations were relatively consistent in XY coordinates on the axial plane across cases, so we used the middle slice from the sagittal and coronal planes as input images for detection. We computed MIPs of the segmentation masks along the sagittal and coronal directions, then constructed 2D, axis-aligned bounding boxes by taking the minimum and maximum mask coordinates in each view. We used YOLOv8n and YOLOv8m for detection, achieving over 0.95 mAP@0.5 on the validation set. After detection, we reconstructed a 3D ROI by combining results from the sagittal and coronal views, and then cropped a fixed-size 3D bounding box of 90×90×90 mm (or 120×120×120 mm) centered on each detection to generate analysis patches.<br>\nExamples of detection results. Green = prediction; red = ground truth.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1408427%2F979857f3afb1fd7d49b4425c121b4ced%2Fyolo.png?generation=1760510088174695&amp;alt=media\" alt=\"\"></p>\n<h3>Stage 2: 3D ROI Classification</h3>\n<h4>Training</h4>\n<p>Using the 3D vessel ROIs from Stage 1, we trained 3D ResNet-18 backbones, implemented with the <code>timm-3d</code> library.</p>\n<p>Our models are inspired by <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/daddies-4th-place-simple-resnet18-classification\" target=\"_blank\">BYU 4th place solution</a>. <br>\nWe attached a 14-class classification head to <strong>each feature map (not whole volume)</strong> and optimized with weighted BCE loss (multi-label setting). In the default 3D configuration (input 128×128×128), the network produced a 4×4×4 feature map, which did not yield good results. Increasing spatial resolution helped: we changed the stride from 2 to 1 in selected convolution layers to obtain larger feature maps.</p>\n<p>We explored the input volume from 128×128×128 up to 224×224×224, and feature-map sizes from 8×8×8 to 48×48×48. Notably, increasing the feature map from 8×8×8 to 25×25×25 and the image size from 128×128×128 to 196×196×196 improved the LB score from 0.77 to 0.81 in a single-fold model.</p>\n<h4>Inference &amp; Aggregation</h4>\n<p>For inference, we aggregated feature-map predictions into a per-case prediction. Specifically, for each class we sorted the <code>Aneurysm Present</code> scores across spatial positions and averaged the top-N scores (Top-N mean) to produce the final class prediction. N depends on model configuration.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1408427%2F886c5768338572f3a261ab74e3881e9f%2Fpipeline.png?generation=1760510113790199&amp;alt=media\" alt=\"\"></p>\n<h4>Model ensemble</h4>\n<p>We built an ensemble of 11 models with different crop sizes, image resolutions, and variants with reduced stride in selected convolution layers. All models used a 3D ResNet-18 backbone. In this ensemble, the public/private LB scores were 0.86/0.84.</p>\n<h3>Missing DICOM tags &amp; Fallbacks</h3>\n<p>There were many missing DICOM tags in the test data, so we built a model to estimate voxel spacing along the X, Y, and Z axes. To preserve XY spacing, each slice was padded and center-cropped to 512×512 pixels, and 10 central slices were sampled along the Z-axis. Each slice was stacked with its adjacent slices to form 2.5D RGB inputs for capturing Z-axis information. Using these inputs, an EfficientNet V2 S regression model predicted voxel spacing from image appearance, with final values obtained by averaging outputs across the 10 slices. On internal validation, the model achieved MAE = 0.015 mm (X), 0.020 mm (Y), and 0.071 mm (Z), improving our private LB from 0.84 to 0.85.</p>\n<h3>What didn't work for us</h3>\n<ul>\n<li>Vessel segmentation</li>\n<li>Modality and Plane heads for auxiliary loss</li>\n<li>Complicated model like MIL, LSTMs and decorder heads</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/code/tamotamo/rsna2025-3rd-place-inference\" target=\"_blank\">Inference code</a><br>\n<a href=\"https://github.com/tamotamo17/RSNA2025-3rd-place-solution\" target=\"_blank\">Training code</a></p>",
      "rawMarkdown": "Thanks to RSNA and Kaggle for hosting this competition—it was a great opportunity to work with real-world medical data. I’m also grateful to my teammates @yosukeyama, @nmstopen, @heroalchem, and @dainagao; our collaboration directly contributed to the results we achieved.\n\n### Overview\n\nOur main solution consists of two stages:\n\n1. 3D vessel region detection\n1. 3D ROI classification\n\n### Stage 1: Vessel Region Detection\n\nBecause the target vessels occupy only a limited portion of the field of view, we first detect whole vessel regions before downstream analysis. We also experimented with vessel segmentation, but it was not sufficiently robust across cases.\nAneurysm locations were relatively consistent in XY coordinates on the axial plane across cases, so we used the middle slice from the sagittal and coronal planes as input images for detection. We computed MIPs of the segmentation masks along the sagittal and coronal directions, then constructed 2D, axis-aligned bounding boxes by taking the minimum and maximum mask coordinates in each view. We used YOLOv8n and YOLOv8m for detection, achieving over 0.95 mAP@0.5 on the validation set. After detection, we reconstructed a 3D ROI by combining results from the sagittal and coronal views, and then cropped a fixed-size 3D bounding box of 90×90×90 mm (or 120×120×120 mm) centered on each detection to generate analysis patches.\nExamples of detection results. Green = prediction; red = ground truth.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1408427%2F979857f3afb1fd7d49b4425c121b4ced%2Fyolo.png?generation=1760510088174695&alt=media)\n\n\n### Stage 2: 3D ROI Classification\n#### Training\nUsing the 3D vessel ROIs from Stage 1, we trained 3D ResNet-18 backbones, implemented with the `timm-3d` library.\n\nOur models are inspired by [BYU 4th place solution](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/daddies-4th-place-simple-resnet18-classification). \nWe attached a 14-class classification head to **each feature map (not whole volume)** and optimized with weighted BCE loss (multi-label setting). In the default 3D configuration (input 128×128×128), the network produced a 4×4×4 feature map, which did not yield good results. Increasing spatial resolution helped: we changed the stride from 2 to 1 in selected convolution layers to obtain larger feature maps.\n\nWe explored the input volume from 128×128×128 up to 224×224×224, and feature-map sizes from 8×8×8 to 48×48×48. Notably, increasing the feature map from 8×8×8 to 25×25×25 and the image size from 128×128×128 to 196×196×196 improved the LB score from 0.77 to 0.81 in a single-fold model.\n\n#### Inference & Aggregation\n\nFor inference, we aggregated feature-map predictions into a per-case prediction. Specifically, for each class we sorted the `Aneurysm Present` scores across spatial positions and averaged the top-N scores (Top-N mean) to produce the final class prediction. N depends on model configuration.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1408427%2F886c5768338572f3a261ab74e3881e9f%2Fpipeline.png?generation=1760510113790199&alt=media)\n#### Model ensemble\nWe built an ensemble of 11 models with different crop sizes, image resolutions, and variants with reduced stride in selected convolution layers. All models used a 3D ResNet-18 backbone. In this ensemble, the public/private LB scores were 0.86/0.84.\n\n### Missing DICOM tags & Fallbacks\nThere were many missing DICOM tags in the test data, so we built a model to estimate voxel spacing along the X, Y, and Z axes. To preserve XY spacing, each slice was padded and center-cropped to 512×512 pixels, and 10 central slices were sampled along the Z-axis. Each slice was stacked with its adjacent slices to form 2.5D RGB inputs for capturing Z-axis information. Using these inputs, an EfficientNet V2 S regression model predicted voxel spacing from image appearance, with final values obtained by averaging outputs across the 10 slices. On internal validation, the model achieved MAE = 0.015 mm (X), 0.020 mm (Y), and 0.071 mm (Z), improving our private LB from 0.84 to 0.85.\n\n### What didn't work for us\n- Vessel segmentation\n- Modality and Plane heads for auxiliary loss\n- Complicated model like MIL, LSTMs and decorder heads\n\n[Inference code](https://www.kaggle.com/code/tamotamo/rsna2025-3rd-place-inference)\n[Training code](https://github.com/tamotamo17/RSNA2025-3rd-place-solution)",
      "votes": 46
    },
    {
      "id": 3304944,
      "postDate": "2025-10-21T16:01:20.963Z",
      "content": "<p>Congrats!<br>\nThank you for sharing such excellent insights!<br>\nI was surprised that the ROI region could be made much smaller than I expected, and I was also impressed by the smart approach to creating training data for YOLO. This apploach is so helpful!</p>\n<p>I have some questions regarding training:</p>\n<ul>\n<li>I was training using the  32-channel Efficienet you shapred as a reference, but no matter how much I experimented with the scheduler, class weights, and loss functions, the training loss never stabilized and kept oscillating. Did you experience similar phenomena when you train the 32-channel model.</li>\n<li>In the final pipeline you used, was the loss function stable?</li>\n</ul>",
      "rawMarkdown": "Congrats!\nThank you for sharing such excellent insights!\nI was surprised that the ROI region could be made much smaller than I expected, and I was also impressed by the smart approach to creating training data for YOLO. This apploach is so helpful!\n\nI have some questions regarding training:\n\n- I was training using the  32-channel Efficienet you shapred as a reference, but no matter how much I experimented with the scheduler, class weights, and loss functions, the training loss never stabilized and kept oscillating. Did you experience similar phenomena when you train the 32-channel model.\n- In the final pipeline you used, was the loss function stable?",
      "replies": [
        {
          "id": 3305033,
          "postDate": "2025-10-21T22:04:49.987Z",
          "content": "<p>Thanks!<br>\nWhen I trained on whole volumes, the loss remained unstable and the score never surpassed 0.7.<br>\nAneurysms are highly localized in 3D, so training on whole volumes turns most voxels into noise.   <br>\nThe model then learns accidental, non-causal patterns, causing unstable loss. Adding localization (tight ROIs and location hints) made it focus on true aneurysm features and stabilized training.  <br>\nIn our pipeline, ROI cropping and higher-resolution feature maps produced consistently stable loss.</p>",
          "rawMarkdown": "Thanks!\nWhen I trained on whole volumes, the loss remained unstable and the score never surpassed 0.7.\nAneurysms are highly localized in 3D, so training on whole volumes turns most voxels into noise.   \nThe model then learns accidental, non-causal patterns, causing unstable loss. Adding localization (tight ROIs and location hints) made it focus on true aneurysm features and stabilized training.  \nIn our pipeline, ROI cropping and higher-resolution feature maps produced consistently stable loss.",
          "replies": [
            {
              "id": 3305080,
              "postDate": "2025-10-22T01:06:00.610Z",
              "content": "<p>Thank you for some infoemation<br>\nas you say, my pipeline ,which crop only skull and resize to 384, may be bad…</p>",
              "rawMarkdown": "Thank you for some infoemation\nas you say, my pipeline ,which crop only skull and resize to 384, may be bad...\n"
            }
          ]
        },
        {
          "id": 3305059,
          "postDate": "2025-10-21T23:41:27.387Z",
          "content": "<p>Regarding the 32-channel model loss, I observed that the loss decreased steadily during training.<br>\nYou can refer to the training log included in the dataset below:<br>\n<a href=\"https://www.kaggle.com/datasets/yosukeyama/rsna2025-effnetv2-32ch\" target=\"_blank\">https://www.kaggle.com/datasets/yosukeyama/rsna2025-effnetv2-32ch</a></p>\n<p>I did not experience unstable or oscillating losses. However, such instability can sometimes occur in medical imaging tasks because the models are highly sensitive to preprocessing pipelines and hyperparameter settings. Even small differences in normalization, clipping, or augmentation can cause large variations in training stability.</p>\n<p>As TmT pointed out, compressing the entire volume inevitably causes a substantial loss of information, which limits performance.<br>\nThrough manual EDA, I found that large aneurysms were well preserved, but very small aneurysms and cases with long z-axis volumes (such as CTA) became difficult to visualize, which I believe contributed to the performance ceiling.</p>",
          "rawMarkdown": "Regarding the 32-channel model loss, I observed that the loss decreased steadily during training.\nYou can refer to the training log included in the dataset below:\nhttps://www.kaggle.com/datasets/yosukeyama/rsna2025-effnetv2-32ch\n\nI did not experience unstable or oscillating losses. However, such instability can sometimes occur in medical imaging tasks because the models are highly sensitive to preprocessing pipelines and hyperparameter settings. Even small differences in normalization, clipping, or augmentation can cause large variations in training stability.\n\nAs TmT pointed out, compressing the entire volume inevitably causes a substantial loss of information, which limits performance.\nThrough manual EDA, I found that large aneurysms were well preserved, but very small aneurysms and cases with long z-axis volumes (such as CTA) became difficult to visualize, which I believe contributed to the performance ceiling.",
          "replies": [
            {
              "id": 3305083,
              "postDate": "2025-10-22T01:15:18.533Z",
              "content": "<p>Got it, thank you!<br>\nThis really highlights the challenges in medical imaging.<br>\nAlso, I appreciate you sharing details about your custom EDA!</p>",
              "rawMarkdown": "Got it, thank you!\nThis really highlights the challenges in medical imaging.\nAlso, I appreciate you sharing details about your custom EDA!"
            }
          ]
        }
      ]
    },
    {
      "id": 3302694,
      "postDate": "2025-10-16T09:38:28.813Z",
      "content": "<p>Great approach. Ensemble different crop size is smart. Can I ask when did your team give up public 32ch approach?</p>",
      "rawMarkdown": "Great approach. Ensemble different crop size is smart. Can I ask when did your team give up public 32ch approach?",
      "replies": [
        {
          "id": 3302710,
          "postDate": "2025-10-16T10:33:56.710Z",
          "content": "<p>Thanks for your question!<br>\nThe 32-channel public notebook was actually my personal baseline code rather than a team project, so I’ll answer that one.</p>\n<p>It wasn’t really something we “gave up” on — it was meant as a simple starting baseline from the beginning. In practice, compressing full volumes into 32 slices turned out to be too destructive for performance. I also tried several MIL and 3D models using 32-channel inputs (and 62-ch), but their results were similar to standard 2D models, which confirmed that limitation.</p>\n<p>The stronger public notebooks at that stage were those using MIP representations, so I decided to share a 3D-volume-based baseline approach instead, hoping it would contribute more to the community.</p>\n<p>Within our team, we all agreed early on that reliable ROI detection would be key. Building on that shared view, TmT designed and completed this brilliant pipeline. Personally, I also experimented with a 3D EfficientNet-V2 model (it actually improved our private score, though not the public one, so we didn’t select it for the final submission) and with several MIL variants, but none gave consistent score improvements.</p>",
          "rawMarkdown": "Thanks for your question!\nThe 32-channel public notebook was actually my personal baseline code rather than a team project, so I’ll answer that one.\n\nIt wasn’t really something we “gave up” on — it was meant as a simple starting baseline from the beginning. In practice, compressing full volumes into 32 slices turned out to be too destructive for performance. I also tried several MIL and 3D models using 32-channel inputs (and 62-ch), but their results were similar to standard 2D models, which confirmed that limitation.\n\nThe stronger public notebooks at that stage were those using MIP representations, so I decided to share a 3D-volume-based baseline approach instead, hoping it would contribute more to the community.\n\nWithin our team, we all agreed early on that reliable ROI detection would be key. Building on that shared view, TmT designed and completed this brilliant pipeline. Personally, I also experimented with a 3D EfficientNet-V2 model (it actually improved our private score, though not the public one, so we didn’t select it for the final submission) and with several MIL variants, but none gave consistent score improvements.",
          "votes": 2,
          "replies": [
            {
              "id": 3302717,
              "postDate": "2025-10-16T10:59:57.250Z",
              "content": "<p><a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a> Thanks for your reply! I also want to ask what is the CV performance of the 3D EfficientNet-V2? Is that the highest single model CV your team achieved?</p>",
              "rawMarkdown": "@yosukeyama Thanks for your reply! I also want to ask what is the CV performance of the 3D EfficientNet-V2? Is that the highest single model CV your team achieved?"
            },
            {
              "id": 3302763,
              "postDate": "2025-10-16T13:05:39.843Z",
              "content": "<p>Thanks! It wasn’t our best model — CV was just under 0.79 on one of the 5 folds.<br>\nTmT’s best model exceeded 0.8. I didn’t submit the single EfficientNet-V2 result to LB, but the ensemble with TmT’s model slightly improved the private score (likely due to architectural diversity), though not the public one.</p>",
              "rawMarkdown": "Thanks! It wasn’t our best model — CV was just under 0.79 on one of the 5 folds.\nTmT’s best model exceeded 0.8. I didn’t submit the single EfficientNet-V2 result to LB, but the ensemble with TmT’s model slightly improved the private score (likely due to architectural diversity), though not the public one.",
              "votes": 1
            },
            {
              "id": 3305081,
              "postDate": "2025-10-22T01:14:48.977Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3302601,
      "postDate": "2025-10-16T06:18:28.230Z",
      "content": "<p>Nice work! Impressive results with such a concise method — congratulations. For 3D CNN direct classification, I noticed you didn't even use vessel or aneurysm segmentation labels as auxiliary supervision. How did you avoid overfitting? Did you use a pre-trained model?</p>",
      "rawMarkdown": "Nice work! Impressive results with such a concise method — congratulations. For 3D CNN direct classification, I noticed you didn't even use vessel or aneurysm segmentation labels as auxiliary supervision. How did you avoid overfitting? Did you use a pre-trained model?",
      "replies": [
        {
          "id": 3302714,
          "postDate": "2025-10-16T10:48:31.477Z",
          "content": "<p>Thanks! We did not use a vessel mask, but used aneurysm coordinates. Training used a voxel-wise BCE loss on the feature maps (not volume-level). We did not use pretrained weights and observed no overfitting, possibly due to the voxel-level supervision.</p>",
          "rawMarkdown": "Thanks! We did not use a vessel mask, but used aneurysm coordinates. Training used a voxel-wise BCE loss on the feature maps (not volume-level). We did not use pretrained weights and observed no overfitting, possibly due to the voxel-level supervision."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3304944,
      "author_name": "shiba-inu",
      "author_url": "",
      "post_date": "2025-10-21T16:01:20.963000",
      "content": "<p>Congrats!<br>\nThank you for sharing such excellent insights!<br>\nI was surprised that the ROI region could be made much smaller than I expected, and I was also impressed by the smart approach to creating training data for YOLO. This apploach is so helpful!</p>\n<p>I have some questions regarding training:</p>\n<ul>\n<li>I was training using the  32-channel Efficienet you shapred as a reference, but no matter how much I experimented with the scheduler, class weights, and loss functions, the training loss never stabilized and kept oscillating. Did you experience similar phenomena when you train the 32-channel model.</li>\n<li>In the final pipeline you used, was the loss function stable?</li>\n</ul>",
      "votes": 0,
      "replies": [
        {
          "id": 3305033,
          "author_name": "TmT",
          "author_url": "",
          "post_date": "2025-10-21T22:04:49.987000",
          "content": "<p>Thanks!<br>\nWhen I trained on whole volumes, the loss remained unstable and the score never surpassed 0.7.<br>\nAneurysms are highly localized in 3D, so training on whole volumes turns most voxels into noise.   <br>\nThe model then learns accidental, non-causal patterns, causing unstable loss. Adding localization (tight ROIs and location hints) made it focus on true aneurysm features and stabilized training.  <br>\nIn our pipeline, ROI cropping and higher-resolution feature maps produced consistently stable loss.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3305080,
              "author_name": "shiba-inu",
              "author_url": "",
              "post_date": "2025-10-22T01:06:00.610000",
              "content": "<p>Thank you for some infoemation<br>\nas you say, my pipeline ,which crop only skull and resize to 384, may be bad…</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3305059,
          "author_name": "YYama",
          "author_url": "",
          "post_date": "2025-10-21T23:41:27.387000",
          "content": "<p>Regarding the 32-channel model loss, I observed that the loss decreased steadily during training.<br>\nYou can refer to the training log included in the dataset below:<br>\n<a href=\"https://www.kaggle.com/datasets/yosukeyama/rsna2025-effnetv2-32ch\" target=\"_blank\">https://www.kaggle.com/datasets/yosukeyama/rsna2025-effnetv2-32ch</a></p>\n<p>I did not experience unstable or oscillating losses. However, such instability can sometimes occur in medical imaging tasks because the models are highly sensitive to preprocessing pipelines and hyperparameter settings. Even small differences in normalization, clipping, or augmentation can cause large variations in training stability.</p>\n<p>As TmT pointed out, compressing the entire volume inevitably causes a substantial loss of information, which limits performance.<br>\nThrough manual EDA, I found that large aneurysms were well preserved, but very small aneurysms and cases with long z-axis volumes (such as CTA) became difficult to visualize, which I believe contributed to the performance ceiling.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3305083,
              "author_name": "shiba-inu",
              "author_url": "",
              "post_date": "2025-10-22T01:15:18.533000",
              "content": "<p>Got it, thank you!<br>\nThis really highlights the challenges in medical imaging.<br>\nAlso, I appreciate you sharing details about your custom EDA!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3302694,
      "author_name": "Tom",
      "author_url": "",
      "post_date": "2025-10-16T09:38:28.813000",
      "content": "<p>Great approach. Ensemble different crop size is smart. Can I ask when did your team give up public 32ch approach?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3302710,
          "author_name": "YYama",
          "author_url": "",
          "post_date": "2025-10-16T10:33:56.710000",
          "content": "<p>Thanks for your question!<br>\nThe 32-channel public notebook was actually my personal baseline code rather than a team project, so I’ll answer that one.</p>\n<p>It wasn’t really something we “gave up” on — it was meant as a simple starting baseline from the beginning. In practice, compressing full volumes into 32 slices turned out to be too destructive for performance. I also tried several MIL and 3D models using 32-channel inputs (and 62-ch), but their results were similar to standard 2D models, which confirmed that limitation.</p>\n<p>The stronger public notebooks at that stage were those using MIP representations, so I decided to share a 3D-volume-based baseline approach instead, hoping it would contribute more to the community.</p>\n<p>Within our team, we all agreed early on that reliable ROI detection would be key. Building on that shared view, TmT designed and completed this brilliant pipeline. Personally, I also experimented with a 3D EfficientNet-V2 model (it actually improved our private score, though not the public one, so we didn’t select it for the final submission) and with several MIL variants, but none gave consistent score improvements.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3302717,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2025-10-16T10:59:57.250000",
              "content": "<p><a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a> Thanks for your reply! I also want to ask what is the CV performance of the 3D EfficientNet-V2? Is that the highest single model CV your team achieved?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3302763,
              "author_name": "YYama",
              "author_url": "",
              "post_date": "2025-10-16T13:05:39.843000",
              "content": "<p>Thanks! It wasn’t our best model — CV was just under 0.79 on one of the 5 folds.<br>\nTmT’s best model exceeded 0.8. I didn’t submit the single EfficientNet-V2 result to LB, but the ensemble with TmT’s model slightly improved the private score (likely due to architectural diversity), though not the public one.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3305081,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-10-22T01:14:48.977000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3302601,
      "author_name": "Shuolin Liu",
      "author_url": "",
      "post_date": "2025-10-16T06:18:28.230000",
      "content": "<p>Nice work! Impressive results with such a concise method — congratulations. For 3D CNN direct classification, I noticed you didn't even use vessel or aneurysm segmentation labels as auxiliary supervision. How did you avoid overfitting? Did you use a pre-trained model?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3302714,
          "author_name": "TmT",
          "author_url": "",
          "post_date": "2025-10-16T10:48:31.477000",
          "content": "<p>Thanks! We did not use a vessel mask, but used aneurysm coordinates. Training used a voxel-wise BCE loss on the feature maps (not volume-level). We did not use pretrained weights and observed no overfitting, possibly due to the voxel-level supervision.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3302141": "Thanks to RSNA and Kaggle for hosting this competition—it was a great opportunity to work with real-world medical data. I’m also grateful to my teammates @yosukeyama, @nmstopen, @heroalchem, and @dainagao; our collaboration directly contributed to the results we achieved.\n\n### Overview\n\nOur main solution consists of two stages:\n\n1. 3D vessel region detection\n1. 3D ROI classification\n\n### Stage 1: Vessel Region Detection\n\nBecause the target vessels occupy only a limited portion of the field of view, we first detect whole vessel regions before downstream analysis. We also experimented with vessel segmentation, but it was not sufficiently robust across cases.\nAneurysm locations were relatively consistent in XY coordinates on the axial plane across cases, so we used the middle slice from the sagittal and coronal planes as input images for detection. We computed MIPs of the segmentation masks along the sagittal and coronal directions, then constructed 2D, axis-aligned bounding boxes by taking the minimum and maximum mask coordinates in each view. We used YOLOv8n and YOLOv8m for detection, achieving over 0.95 mAP@0.5 on the validation set. After detection, we reconstructed a 3D ROI by combining results from the sagittal and coronal views, and then cropped a fixed-size 3D bounding box of 90×90×90 mm (or 120×120×120 mm) centered on each detection to generate analysis patches.\nExamples of detection results. Green = prediction; red = ground truth.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1408427%2F979857f3afb1fd7d49b4425c121b4ced%2Fyolo.png?generation=1760510088174695&alt=media)\n\n\n### Stage 2: 3D ROI Classification\n#### Training\nUsing the 3D vessel ROIs from Stage 1, we trained 3D ResNet-18 backbones, implemented with the `timm-3d` library.\n\nOur models are inspired by [BYU 4th place solution](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/daddies-4th-place-simple-resnet18-classification). \nWe attached a 14-class classification head to **each feature map (not whole volume)** and optimized with weighted BCE loss (multi-label setting). In the default 3D configuration (input 128×128×128), the network produced a 4×4×4 feature map, which did not yield good results. Increasing spatial resolution helped: we changed the stride from 2 to 1 in selected convolution layers to obtain larger feature maps.\n\nWe explored the input volume from 128×128×128 up to 224×224×224, and feature-map sizes from 8×8×8 to 48×48×48. Notably, increasing the feature map from 8×8×8 to 25×25×25 and the image size from 128×128×128 to 196×196×196 improved the LB score from 0.77 to 0.81 in a single-fold model.\n\n#### Inference & Aggregation\n\nFor inference, we aggregated feature-map predictions into a per-case prediction. Specifically, for each class we sorted the `Aneurysm Present` scores across spatial positions and averaged the top-N scores (Top-N mean) to produce the final class prediction. N depends on model configuration.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1408427%2F886c5768338572f3a261ab74e3881e9f%2Fpipeline.png?generation=1760510113790199&alt=media)\n#### Model ensemble\nWe built an ensemble of 11 models with different crop sizes, image resolutions, and variants with reduced stride in selected convolution layers. All models used a 3D ResNet-18 backbone. In this ensemble, the public/private LB scores were 0.86/0.84.\n\n### Missing DICOM tags & Fallbacks\nThere were many missing DICOM tags in the test data, so we built a model to estimate voxel spacing along the X, Y, and Z axes. To preserve XY spacing, each slice was padded and center-cropped to 512×512 pixels, and 10 central slices were sampled along the Z-axis. Each slice was stacked with its adjacent slices to form 2.5D RGB inputs for capturing Z-axis information. Using these inputs, an EfficientNet V2 S regression model predicted voxel spacing from image appearance, with final values obtained by averaging outputs across the 10 slices. On internal validation, the model achieved MAE = 0.015 mm (X), 0.020 mm (Y), and 0.071 mm (Z), improving our private LB from 0.84 to 0.85.\n\n### What didn't work for us\n- Vessel segmentation\n- Modality and Plane heads for auxiliary loss\n- Complicated model like MIL, LSTMs and decorder heads\n\n[Inference code](https://www.kaggle.com/code/tamotamo/rsna2025-3rd-place-inference)\n[Training code](https://github.com/tamotamo17/RSNA2025-3rd-place-solution)",
    "3304944": "Congrats!\nThank you for sharing such excellent insights!\nI was surprised that the ROI region could be made much smaller than I expected, and I was also impressed by the smart approach to creating training data for YOLO. This apploach is so helpful!\n\nI have some questions regarding training:\n\n- I was training using the  32-channel Efficienet you shapred as a reference, but no matter how much I experimented with the scheduler, class weights, and loss functions, the training loss never stabilized and kept oscillating. Did you experience similar phenomena when you train the 32-channel model.\n- In the final pipeline you used, was the loss function stable?",
    "3302694": "Great approach. Ensemble different crop size is smart. Can I ask when did your team give up public 32ch approach?",
    "3302601": "Nice work! Impressive results with such a concise method — congratulations. For 3D CNN direct classification, I noticed you didn't even use vessel or aneurysm segmentation labels as auxiliary supervision. How did you avoid overfitting? Did you use a pre-trained model?"
  }
}