{
  "id": 611925,
  "title": "6th Place Solution",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/611925",
  "author_name": "Theo Viel",
  "post_date": "2025-10-15T15:43:14.285000",
  "votes": 38,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Thanks to the hosts for once again a nice challenge. We always enjoy joining RSNA competitions :)</p>\n<h2>Overview</h2>\n<p>Our solution is a 2.5D pipeline which consists of 4 steps:</p>\n<ul>\n<li>Skull cropping</li>\n<li>Vessel segmentation (only at train time)</li>\n<li>2D Aneurysm frame-level classification</li>\n<li>Aggregation using a sequence model</li>\n</ul>\n<p>It is inspired by Theo’s <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/writeups/on-strike-2nd-place-solution\" target=\"_blank\">2nd place solution</a> from 2 years ago - although the pipeline required a lot of adjustments to achieve decent performance.</p>\n<p>It achieves <strong>CV 0.895 - Public 0.84 - Private 0.84</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2Ff8a9a0b2fe3c5ecc2f1205dcce497097%2FRSNA-IMG2.png?generation=1761605329379498&amp;alt=media\" alt=\"Image\"></p>\n<p><em>Pipeline Overview. Click <a href=\"https://ibb.co/Lzqg59YK\" target=\"_blank\">here</a> for the full size the image.</em></p>\n<h2>Models</h2>\n<h3>ROI Skull cropping</h3>\n<p>Nothing fancy, the model is a simple 3D-Unet trained with some open-source data found online. The task is quite easy and the model works on all the orientations. This allows for more uniformity across the dataset.</p>\n<h3>Vessel segmentation</h3>\n<p>The task is very similar to the skull segmentation task, but way harder. This model only supports axial data, and requires some tricks to achieve good performance:</p>\n<ul>\n<li>Position encoding to feed left / right information to the CNN</li>\n<li>Mask dilation + dice to fight class imbalance</li>\n</ul>\n<p>This model is only used to sample good negatives to feed the 2D models. The initial goal was also to train a vessel crop classification model - but this approach did quite poorly.</p>\n<h3>2D classification</h3>\n<h4>Main points</h4>\n<p>This is where most of the heavy lifting comes from. We heavily optimized 2D-CNN models by training them with cleverly sampled frames. </p>\n<p>Architectures used are <code>coatnet_rmlp_2_rw_384</code> and <code>maxvit_rmlp_base_rw_384</code>. The Relative Position Encoding MLP(rmlp) helps a lot since position information is very important to distinguish the 13 classes. On top of it, a custom pooling is used not to lose the left/right information before the logits layer.</p>\n<p>This model works on coronal and axial data: sagittal stacks are converted to axial during inference, and ignored during training.</p>\n<h4>More Ideas</h4>\n<p>Each of the Following approximately brought around +0.01 CV</p>\n<ul>\n<li>Further cropping the ROI by using fixed ratios to restrict the skull to areas that actually contain aneurysm.</li>\n<li>Strong ShiftScaleRotate and color augmentations</li>\n<li>Cutmix (or Mixup) to prevent overfitting</li>\n<li>Horizontal flip augmentation that also flips the targets</li>\n<li>Use 2 adjacent frames as 3 channels</li>\n<li>Use SliceSpacing information to sample the adjacent frames further if the spacing is small.</li>\n<li>Ian manually refined the localizers labels to obtain segmentation masks. This allowed for more accurate frame sampling for positives</li>\n</ul>\n<p>Other things that did not really help CV but were kept for robustness</p>\n<ul>\n<li>External data from OpenNeuro.org</li>\n<li>Axial -&gt; Coronal augmentation</li>\n</ul>\n<h3>Sequence models</h3>\n<p>Using a simple max aggregation using predictions of the 2D models on all the frames already gave 0.87 CV. Adding a custom sequence model further improved results to 0.88. This also allowed for easy ensembling, with a 3-model ensemble reaching <strong>0.895 CV</strong>.</p>\n<p>To reduce inference time, the maximum number of frames is limited to a fixed number depending on the modality. Furthermore, only half of the frames are inferred for stacks of size &gt; 64, and a quarter for stacks &gt; 128. The model uses 3 adjacent frames as input which means the information is not lost anyways.</p>\n<h2>Final words</h2>\n<p>Using 2 models instead of 3 in the pipeline would make the pipeline run in 7 hours, which would probably have been enough to blend-in a 3D pipeline. We also had strong results with such approaches (LB 0.79).\nHowever CV improvements were small (&lt;0.01) and due to submission runtime instability and our CV being concerningly high compared to LB, we decided not to invest more time there.</p>\n<p><em>Thanks for reading !</em></p>",
  "messages": [
    {
      "id": 3302339,
      "postDate": "2025-10-15T15:43:14.287Z",
      "content": "<p>Thanks to the hosts for once again a nice challenge. We always enjoy joining RSNA competitions :)</p>\n<h2>Overview</h2>\n<p>Our solution is a 2.5D pipeline which consists of 4 steps:</p>\n<ul>\n<li>Skull cropping</li>\n<li>Vessel segmentation (only at train time)</li>\n<li>2D Aneurysm frame-level classification</li>\n<li>Aggregation using a sequence model</li>\n</ul>\n<p>It is inspired by Theo’s <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/writeups/on-strike-2nd-place-solution\" target=\"_blank\">2nd place solution</a> from 2 years ago - although the pipeline required a lot of adjustments to achieve decent performance.</p>\n<p>It achieves <strong>CV 0.895 - Public 0.84 - Private 0.84</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2Ff8a9a0b2fe3c5ecc2f1205dcce497097%2FRSNA-IMG2.png?generation=1761605329379498&amp;alt=media\" alt=\"Image\"></p>\n<p><em>Pipeline Overview. Click <a href=\"https://ibb.co/Lzqg59YK\" target=\"_blank\">here</a> for the full size the image.</em></p>\n<h2>Models</h2>\n<h3>ROI Skull cropping</h3>\n<p>Nothing fancy, the model is a simple 3D-Unet trained with some open-source data found online. The task is quite easy and the model works on all the orientations. This allows for more uniformity across the dataset.</p>\n<h3>Vessel segmentation</h3>\n<p>The task is very similar to the skull segmentation task, but way harder. This model only supports axial data, and requires some tricks to achieve good performance:</p>\n<ul>\n<li>Position encoding to feed left / right information to the CNN</li>\n<li>Mask dilation + dice to fight class imbalance</li>\n</ul>\n<p>This model is only used to sample good negatives to feed the 2D models. The initial goal was also to train a vessel crop classification model - but this approach did quite poorly.</p>\n<h3>2D classification</h3>\n<h4>Main points</h4>\n<p>This is where most of the heavy lifting comes from. We heavily optimized 2D-CNN models by training them with cleverly sampled frames. </p>\n<p>Architectures used are <code>coatnet_rmlp_2_rw_384</code> and <code>maxvit_rmlp_base_rw_384</code>. The Relative Position Encoding MLP(rmlp) helps a lot since position information is very important to distinguish the 13 classes. On top of it, a custom pooling is used not to lose the left/right information before the logits layer.</p>\n<p>This model works on coronal and axial data: sagittal stacks are converted to axial during inference, and ignored during training.</p>\n<h4>More Ideas</h4>\n<p>Each of the Following approximately brought around +0.01 CV</p>\n<ul>\n<li>Further cropping the ROI by using fixed ratios to restrict the skull to areas that actually contain aneurysm.</li>\n<li>Strong ShiftScaleRotate and color augmentations</li>\n<li>Cutmix (or Mixup) to prevent overfitting</li>\n<li>Horizontal flip augmentation that also flips the targets</li>\n<li>Use 2 adjacent frames as 3 channels</li>\n<li>Use SliceSpacing information to sample the adjacent frames further if the spacing is small.</li>\n<li>Ian manually refined the localizers labels to obtain segmentation masks. This allowed for more accurate frame sampling for positives</li>\n</ul>\n<p>Other things that did not really help CV but were kept for robustness</p>\n<ul>\n<li>External data from OpenNeuro.org</li>\n<li>Axial -&gt; Coronal augmentation</li>\n</ul>\n<h3>Sequence models</h3>\n<p>Using a simple max aggregation using predictions of the 2D models on all the frames already gave 0.87 CV. Adding a custom sequence model further improved results to 0.88. This also allowed for easy ensembling, with a 3-model ensemble reaching <strong>0.895 CV</strong>.</p>\n<p>To reduce inference time, the maximum number of frames is limited to a fixed number depending on the modality. Furthermore, only half of the frames are inferred for stacks of size &gt; 64, and a quarter for stacks &gt; 128. The model uses 3 adjacent frames as input which means the information is not lost anyways.</p>\n<h2>Final words</h2>\n<p>Using 2 models instead of 3 in the pipeline would make the pipeline run in 7 hours, which would probably have been enough to blend-in a 3D pipeline. We also had strong results with such approaches (LB 0.79).\nHowever CV improvements were small (&lt;0.01) and due to submission runtime instability and our CV being concerningly high compared to LB, we decided not to invest more time there.</p>\n<p><em>Thanks for reading !</em></p>",
      "rawMarkdown": "Thanks to the hosts for once again a nice challenge. We always enjoy joining RSNA competitions :)\n\n## Overview\n\nOur solution is a 2.5D pipeline which consists of 4 steps:\n- Skull cropping\n- Vessel segmentation (only at train time)\n- 2D Aneurysm frame-level classification\n- Aggregation using a sequence model\n\nIt is inspired by Theo’s [2nd place solution](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/writeups/on-strike-2nd-place-solution) from 2 years ago - although the pipeline required a lot of adjustments to achieve decent performance.\n\nIt achieves **CV 0.895 - Public 0.84 - Private 0.84**\n\n![Image](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2Ff8a9a0b2fe3c5ecc2f1205dcce497097%2FRSNA-IMG2.png?generation=1761605329379498&alt=media)\n\n*Pipeline Overview. Click [here](https://ibb.co/Lzqg59YK) for the full size the image.*\n\n## Models\n\n### ROI Skull cropping\n\nNothing fancy, the model is a simple 3D-Unet trained with some open-source data found online. The task is quite easy and the model works on all the orientations. This allows for more uniformity across the dataset.\n\n### Vessel segmentation\n\nThe task is very similar to the skull segmentation task, but way harder. This model only supports axial data, and requires some tricks to achieve good performance:\n- Position encoding to feed left / right information to the CNN\n- Mask dilation + dice to fight class imbalance\n\nThis model is only used to sample good negatives to feed the 2D models. The initial goal was also to train a vessel crop classification model - but this approach did quite poorly.\n\n### 2D classification\n\n#### Main points\n\nThis is where most of the heavy lifting comes from. We heavily optimized 2D-CNN models by training them with cleverly sampled frames. \n\nArchitectures used are `coatnet_rmlp_2_rw_384` and `maxvit_rmlp_base_rw_384`. The Relative Position Encoding MLP(rmlp) helps a lot since position information is very important to distinguish the 13 classes. On top of it, a custom pooling is used not to lose the left/right information before the logits layer.\n\nThis model works on coronal and axial data: sagittal stacks are converted to axial during inference, and ignored during training.\n\n#### More Ideas\n\nEach of the Following approximately brought around +0.01 CV\n\n- Further cropping the ROI by using fixed ratios to restrict the skull to areas that actually contain aneurysm.\n- Strong ShiftScaleRotate and color augmentations\n- Cutmix (or Mixup) to prevent overfitting\n- Horizontal flip augmentation that also flips the targets\n- Use 2 adjacent frames as 3 channels\n- Use SliceSpacing information to sample the adjacent frames further if the spacing is small.\n- Ian manually refined the localizers labels to obtain segmentation masks. This allowed for more accurate frame sampling for positives\n\nOther things that did not really help CV but were kept for robustness\n- External data from OpenNeuro.org\n- Axial -> Coronal augmentation\n\n### Sequence models\n\nUsing a simple max aggregation using predictions of the 2D models on all the frames already gave 0.87 CV. Adding a custom sequence model further improved results to 0.88. This also allowed for easy ensembling, with a 3-model ensemble reaching **0.895 CV**.\n\nTo reduce inference time, the maximum number of frames is limited to a fixed number depending on the modality. Furthermore, only half of the frames are inferred for stacks of size > 64, and a quarter for stacks > 128. The model uses 3 adjacent frames as input which means the information is not lost anyways.\n\n## Final words\n\nUsing 2 models instead of 3 in the pipeline would make the pipeline run in 7 hours, which would probably have been enough to blend-in a 3D pipeline. We also had strong results with such approaches (LB 0.79).\nHowever CV improvements were small (<0.01) and due to submission runtime instability and our CV being concerningly high compared to LB, we decided not to invest more time there.\n\n*Thanks for reading !*\n",
      "votes": 38
    },
    {
      "id": 3302457,
      "postDate": "2025-10-15T22:44:43.307Z",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> Nice solution! Have you considered using a 3D Spatial Transformer Network (STN) to automatically learn and focus on the regions of interest during the initial stage?</p>",
      "rawMarkdown": "@theoviel Nice solution! Have you considered using a 3D Spatial Transformer Network (STN) to automatically learn and focus on the regions of interest during the initial stage?",
      "votes": 1,
      "replies": [
        {
          "id": 3303161,
          "postDate": "2025-10-17T09:00:49.937Z",
          "content": "<p>Thanks ! We have not. Do you have any reference for that ?</p>",
          "rawMarkdown": "Thanks ! We have not. Do you have any reference for that ?",
          "replies": [
            {
              "id": 3303548,
              "postDate": "2025-10-18T11:29:42.397Z",
              "content": "<p><a href=\"https://medium.com/data-science/review-stn-spatial-transformer-network-image-classification-d3cbd98a70aa\" target=\"_blank\">https://medium.com/data-science/review-stn-spatial-transformer-network-image-classification-d3cbd98a70aa</a></p>",
              "rawMarkdown": "https://medium.com/data-science/review-stn-spatial-transformer-network-image-classification-d3cbd98a70aa",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3302361,
      "postDate": "2025-10-15T16:35:00.457Z",
      "content": "<p>Congratulations to the entire team!<br>\nAll images are simply minmaxed normalized? No windowing?</p>\n<p>Did you try training from scratch the 2.5D model?</p>",
      "rawMarkdown": "Congratulations to the entire team!\nAll images are simply minmaxed normalized? No windowing?\n\nDid you try training from scratch the 2.5D model?",
      "votes": 1,
      "replies": [
        {
          "id": 3302375,
          "postDate": "2025-10-15T17:09:28.013Z",
          "content": "<p>Normalization is the following:</p>\n<pre><code>min_ = .percentile(.(), )\nmax_ = .percentile(.(), )\n = .clip(, min_, max_)\n\neps =   min_ == max_  \n = (.astype(.float32) - min_) / (max_ - min_ + eps)\n</code></pre>\n<p><br>\nNo luck with CT windowing, at least for the 2D models.</p>\n<blockquote>\n  <p>Did you try training from scratch the 2.5D model?</p>\n</blockquote>\n<p>Do you mean end-to-end CNN + LSTM head ? </p>",
          "rawMarkdown": "Normalization is the following:\n```            \nmin_ = np.percentile(image.flatten(), 0.1)\nmax_ = np.percentile(image.flatten(), 99.9)\nimage = np.clip(image, min_, max_)\n\neps = 1e-5 if min_ == max_ else 0\nimage = (image.astype(np.float32) - min_) / (max_ - min_ + eps)\n``` \nNo luck with CT windowing, at least for the 2D models.\n\n> Did you try training from scratch the 2.5D model?\n\nDo you mean end-to-end CNN + LSTM head ? ",
          "votes": 2,
          "replies": [
            {
              "id": 3302400,
              "postDate": "2025-10-15T18:27:54.347Z",
              "content": "<p>Yes end to end training </p>",
              "rawMarkdown": "Yes end to end training "
            },
            {
              "id": 3302438,
              "postDate": "2025-10-15T21:13:30.733Z",
              "content": "<p>Multiple frames + RNN head performed quite poorly.<br>\nI did not try to feed all the frames though. Seems quite heavy computationally and I think models will overfit very fast.</p>",
              "rawMarkdown": "Multiple frames + RNN head performed quite poorly.\nI did not try to feed all the frames though. Seems quite heavy computationally and I think models will overfit very fast.",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3302457,
      "author_name": "Tom",
      "author_url": "",
      "post_date": "2025-10-15T22:44:43.307000",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> Nice solution! Have you considered using a 3D Spatial Transformer Network (STN) to automatically learn and focus on the regions of interest during the initial stage?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3303161,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2025-10-17T09:00:49.937000",
          "content": "<p>Thanks ! We have not. Do you have any reference for that ?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3303548,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2025-10-18T11:29:42.397000",
              "content": "<p><a href=\"https://medium.com/data-science/review-stn-spatial-transformer-network-image-classification-d3cbd98a70aa\" target=\"_blank\">https://medium.com/data-science/review-stn-spatial-transformer-network-image-classification-d3cbd98a70aa</a></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3302361,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2025-10-15T16:35:00.457000",
      "content": "<p>Congratulations to the entire team!<br>\nAll images are simply minmaxed normalized? No windowing?</p>\n<p>Did you try training from scratch the 2.5D model?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3302375,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2025-10-15T17:09:28.013000",
          "content": "<p>Normalization is the following:</p>\n<pre><code>min_ = .percentile(.(), )\nmax_ = .percentile(.(), )\n = .clip(, min_, max_)\n\neps =   min_ == max_  \n = (.astype(.float32) - min_) / (max_ - min_ + eps)\n</code></pre>\n<p><br>\nNo luck with CT windowing, at least for the 2D models.</p>\n<blockquote>\n  <p>Did you try training from scratch the 2.5D model?</p>\n</blockquote>\n<p>Do you mean end-to-end CNN + LSTM head ? </p>",
          "votes": 2,
          "replies": [
            {
              "id": 3302400,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2025-10-15T18:27:54.347000",
              "content": "<p>Yes end to end training </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3302438,
              "author_name": "Theo Viel",
              "author_url": "",
              "post_date": "2025-10-15T21:13:30.733000",
              "content": "<p>Multiple frames + RNN head performed quite poorly.<br>\nI did not try to feed all the frames though. Seems quite heavy computationally and I think models will overfit very fast.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3302339": "Thanks to the hosts for once again a nice challenge. We always enjoy joining RSNA competitions :)\n\n## Overview\n\nOur solution is a 2.5D pipeline which consists of 4 steps:\n- Skull cropping\n- Vessel segmentation (only at train time)\n- 2D Aneurysm frame-level classification\n- Aggregation using a sequence model\n\nIt is inspired by Theo’s [2nd place solution](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/writeups/on-strike-2nd-place-solution) from 2 years ago - although the pipeline required a lot of adjustments to achieve decent performance.\n\nIt achieves **CV 0.895 - Public 0.84 - Private 0.84**\n\n![Image](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F5570735%2Ff8a9a0b2fe3c5ecc2f1205dcce497097%2FRSNA-IMG2.png?generation=1761605329379498&alt=media)\n\n*Pipeline Overview. Click [here](https://ibb.co/Lzqg59YK) for the full size the image.*\n\n## Models\n\n### ROI Skull cropping\n\nNothing fancy, the model is a simple 3D-Unet trained with some open-source data found online. The task is quite easy and the model works on all the orientations. This allows for more uniformity across the dataset.\n\n### Vessel segmentation\n\nThe task is very similar to the skull segmentation task, but way harder. This model only supports axial data, and requires some tricks to achieve good performance:\n- Position encoding to feed left / right information to the CNN\n- Mask dilation + dice to fight class imbalance\n\nThis model is only used to sample good negatives to feed the 2D models. The initial goal was also to train a vessel crop classification model - but this approach did quite poorly.\n\n### 2D classification\n\n#### Main points\n\nThis is where most of the heavy lifting comes from. We heavily optimized 2D-CNN models by training them with cleverly sampled frames. \n\nArchitectures used are `coatnet_rmlp_2_rw_384` and `maxvit_rmlp_base_rw_384`. The Relative Position Encoding MLP(rmlp) helps a lot since position information is very important to distinguish the 13 classes. On top of it, a custom pooling is used not to lose the left/right information before the logits layer.\n\nThis model works on coronal and axial data: sagittal stacks are converted to axial during inference, and ignored during training.\n\n#### More Ideas\n\nEach of the Following approximately brought around +0.01 CV\n\n- Further cropping the ROI by using fixed ratios to restrict the skull to areas that actually contain aneurysm.\n- Strong ShiftScaleRotate and color augmentations\n- Cutmix (or Mixup) to prevent overfitting\n- Horizontal flip augmentation that also flips the targets\n- Use 2 adjacent frames as 3 channels\n- Use SliceSpacing information to sample the adjacent frames further if the spacing is small.\n- Ian manually refined the localizers labels to obtain segmentation masks. This allowed for more accurate frame sampling for positives\n\nOther things that did not really help CV but were kept for robustness\n- External data from OpenNeuro.org\n- Axial -> Coronal augmentation\n\n### Sequence models\n\nUsing a simple max aggregation using predictions of the 2D models on all the frames already gave 0.87 CV. Adding a custom sequence model further improved results to 0.88. This also allowed for easy ensembling, with a 3-model ensemble reaching **0.895 CV**.\n\nTo reduce inference time, the maximum number of frames is limited to a fixed number depending on the modality. Furthermore, only half of the frames are inferred for stacks of size > 64, and a quarter for stacks > 128. The model uses 3 adjacent frames as input which means the information is not lost anyways.\n\n## Final words\n\nUsing 2 models instead of 3 in the pipeline would make the pipeline run in 7 hours, which would probably have been enough to blend-in a 3D pipeline. We also had strong results with such approaches (LB 0.79).\nHowever CV improvements were small (<0.01) and due to submission runtime instability and our CV being concerningly high compared to LB, we decided not to invest more time there.\n\n*Thanks for reading !*\n",
    "3302457": "@theoviel Nice solution! Have you considered using a 3D Spatial Transformer Network (STN) to automatically learn and focus on the regions of interest during the initial stage?",
    "3302361": "Congratulations to the entire team!\nAll images are simply minmaxed normalized? No windowing?\n\nDid you try training from scratch the 2.5D model?"
  }
}