{
  "id": 613534,
  "title": "8th Place Solution",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/613534",
  "author_name": "Konni",
  "post_date": "2025-10-27T17:06:16.541000",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<h2>Approach:</h2>\n<ul>\n<li>Two stage approach to find aneurysms in sequences of MR/CT images without using resampling, 3D approaches or segmentation. </li>\n<li>Approach the problem as outlier prediction task: we have only a small number of aneurysm spots in the training sequences, but can sample a huge number of negative spots, where no aneurysm is present.</li>\n<li><strong>Stage-1:</strong> Train a classifier on single images.</li>\n<li><strong>Stage-2:</strong> Train a Transformer to classify the whole sequence of images based on the extracted features of Stage-1.</li>\n</ul>\n<h2>Stage 1:</h2>\n<p>Based on the given <strong>train_localizers.csv</strong> sample positive (aneurysm present) and negative (no aneurysm present) cases and train a simple image classifier. Use the surrounding images as R and B channel of the image. Two different step sizes are used for stacking. Step 2 simply means not stacking three consecutive frames but instead leaving a gap and taking the next frame. As backbone a Convnext Base Dinov3 is used with a simple classification head. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2298475%2Fcb023dda34b71ba52d5a862d3893bdf7%2Fstage1.png?generation=1761583583043279&amp;alt=media\" alt=\"stage-1\"></p>\n<h3><strong>Augmentations:</strong></h3>\n<pre><code>valid_transforms = A.Compose([\n                              A.CenterCrop(448, 448),\n                              ]) \n\n\ntrain_transforms = A.Compose([\n                              A.ShiftScaleRotate(rotate_limit=(-5, 5), =0.5),\n                              A.RandomCrop(448, 448),\n                              A.RandomRotate90(=1.0),\n                              A.OneOf([\n                                        A.GridDropout(=0.4, =0.5),\n                                        A.CoarseDropout(=25,\n                                                        =int(0.2*448),\n                                                        =int(0.2*448),\n                                                        =10,\n                                                        =int(0.1*448),\n                                                        =int(0.1*448),\n                                                        =0.5),\n                                        A.GridDistortion(=1.0),\n                                        ], =0.5),\n                              ])\n</code></pre>\n<p>Images are getting resized to <strong>512x512</strong> during pre-processing. Rotation invariant training and in addition a left to right flip by simultaneously flipping the labels is applied. </p>\n<h3><strong>Sampling:</strong></h3>\n<p>To sample negative spots for training of Stage-1, image from sequences without an aneurysm but also images from sequences with aneurysms are used. In the second case, spots where an aneurysm is present are excluded with some boarders. Furthermore, based on the OOF predictions of an early trained classifier, spots where the model predicts false positives are sampled more often. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2298475%2Fab28e796d8a103ce5e7ee38f5b1fc79b%2Fsampling.png?generation=1761583609064432&amp;alt=media\" alt=\"sampling\"></p>\n<h2>Stage 2:</h2>\n<p>Based on the extracted features of Stage-1 a Transformer is trained on the complete sequence per UID. For augmentation in this stage, sequences with step-1 and step-2 are extracted and used during training of Stage-2. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2298475%2Fa21b25418847a7b1d18bc3f16c0583f4%2Fstage2.png?generation=1761586827480711&amp;alt=media\" alt=\"stage2\"></p>\n<h2>Ensemble:</h2>\n<ul>\n<li>An ensemble of 4 Convnext Base models with two different step sizes is used.</li>\n<li>On sequences &gt; 192 images, only every n image is selected to shrink sequence length (n is calculated dynamically based on original sequence length).</li>\n<li>To shrink the sequence length is only necessary because of the 12h time limit and the incredible crappy Kaggle hardware with totally random runtime based on the assigned server. </li>\n<li>Nearly half of my submissions are simply timeouts, sorry Kaggle, but it can’t be that the same code without changes hit the 12h time limit or finishes in less than 8h.</li>\n</ul>\n<h2>Links:</h2>\n<p><a href=\"https://github.com/KonradHabel/rsna\" target=\"_blank\">Training</a></p>\n<p><a href=\"https://www.kaggle.com/code/khabel/rsna-dual-gpu-ensemble-2step\" target=\"_blank\">Inference</a></p>\n<p><a href=\"https://www.kaggle.com/datasets/khabel/rsna-2025-weights-4xbase\" target=\"_blank\">Weights</a></p>",
  "messages": [
    {
      "id": 3307740,
      "postDate": "2025-10-27T17:06:16.540Z",
      "content": "<h2>Approach:</h2>\n<ul>\n<li>Two stage approach to find aneurysms in sequences of MR/CT images without using resampling, 3D approaches or segmentation. </li>\n<li>Approach the problem as outlier prediction task: we have only a small number of aneurysm spots in the training sequences, but can sample a huge number of negative spots, where no aneurysm is present.</li>\n<li><strong>Stage-1:</strong> Train a classifier on single images.</li>\n<li><strong>Stage-2:</strong> Train a Transformer to classify the whole sequence of images based on the extracted features of Stage-1.</li>\n</ul>\n<h2>Stage 1:</h2>\n<p>Based on the given <strong>train_localizers.csv</strong> sample positive (aneurysm present) and negative (no aneurysm present) cases and train a simple image classifier. Use the surrounding images as R and B channel of the image. Two different step sizes are used for stacking. Step 2 simply means not stacking three consecutive frames but instead leaving a gap and taking the next frame. As backbone a Convnext Base Dinov3 is used with a simple classification head. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2298475%2Fcb023dda34b71ba52d5a862d3893bdf7%2Fstage1.png?generation=1761583583043279&amp;alt=media\" alt=\"stage-1\"></p>\n<h3><strong>Augmentations:</strong></h3>\n<pre><code>valid_transforms = A.Compose([\n                              A.CenterCrop(448, 448),\n                              ]) \n\n\ntrain_transforms = A.Compose([\n                              A.ShiftScaleRotate(rotate_limit=(-5, 5), =0.5),\n                              A.RandomCrop(448, 448),\n                              A.RandomRotate90(=1.0),\n                              A.OneOf([\n                                        A.GridDropout(=0.4, =0.5),\n                                        A.CoarseDropout(=25,\n                                                        =int(0.2*448),\n                                                        =int(0.2*448),\n                                                        =10,\n                                                        =int(0.1*448),\n                                                        =int(0.1*448),\n                                                        =0.5),\n                                        A.GridDistortion(=1.0),\n                                        ], =0.5),\n                              ])\n</code></pre>\n<p>Images are getting resized to <strong>512x512</strong> during pre-processing. Rotation invariant training and in addition a left to right flip by simultaneously flipping the labels is applied. </p>\n<h3><strong>Sampling:</strong></h3>\n<p>To sample negative spots for training of Stage-1, image from sequences without an aneurysm but also images from sequences with aneurysms are used. In the second case, spots where an aneurysm is present are excluded with some boarders. Furthermore, based on the OOF predictions of an early trained classifier, spots where the model predicts false positives are sampled more often. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2298475%2Fab28e796d8a103ce5e7ee38f5b1fc79b%2Fsampling.png?generation=1761583609064432&amp;alt=media\" alt=\"sampling\"></p>\n<h2>Stage 2:</h2>\n<p>Based on the extracted features of Stage-1 a Transformer is trained on the complete sequence per UID. For augmentation in this stage, sequences with step-1 and step-2 are extracted and used during training of Stage-2. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2298475%2Fa21b25418847a7b1d18bc3f16c0583f4%2Fstage2.png?generation=1761586827480711&amp;alt=media\" alt=\"stage2\"></p>\n<h2>Ensemble:</h2>\n<ul>\n<li>An ensemble of 4 Convnext Base models with two different step sizes is used.</li>\n<li>On sequences &gt; 192 images, only every n image is selected to shrink sequence length (n is calculated dynamically based on original sequence length).</li>\n<li>To shrink the sequence length is only necessary because of the 12h time limit and the incredible crappy Kaggle hardware with totally random runtime based on the assigned server. </li>\n<li>Nearly half of my submissions are simply timeouts, sorry Kaggle, but it can’t be that the same code without changes hit the 12h time limit or finishes in less than 8h.</li>\n</ul>\n<h2>Links:</h2>\n<p><a href=\"https://github.com/KonradHabel/rsna\" target=\"_blank\">Training</a></p>\n<p><a href=\"https://www.kaggle.com/code/khabel/rsna-dual-gpu-ensemble-2step\" target=\"_blank\">Inference</a></p>\n<p><a href=\"https://www.kaggle.com/datasets/khabel/rsna-2025-weights-4xbase\" target=\"_blank\">Weights</a></p>",
      "rawMarkdown": "## Approach:\n- Two stage approach to find aneurysms in sequences of MR/CT images without using resampling, 3D approaches or segmentation. \n- Approach the problem as outlier prediction task: we have only a small number of aneurysm spots in the training sequences, but can sample a huge number of negative spots, where no aneurysm is present.\n- **Stage-1:** Train a classifier on single images.\n- **Stage-2:** Train a Transformer to classify the whole sequence of images based on the extracted features of Stage-1.\n\n## Stage 1:\nBased on the given **train_localizers.csv** sample positive (aneurysm present) and negative (no aneurysm present) cases and train a simple image classifier. Use the surrounding images as R and B channel of the image. Two different step sizes are used for stacking. Step 2 simply means not stacking three consecutive frames but instead leaving a gap and taking the next frame. As backbone a Convnext Base Dinov3 is used with a simple classification head. \n\n![stage-1](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2298475%2Fcb023dda34b71ba52d5a862d3893bdf7%2Fstage1.png?generation=1761583583043279&alt=media)\n\n### **Augmentations:**\n```\nvalid_transforms = A.Compose([\n                              A.CenterCrop(448, 448),\n                              ]) \n\n\ntrain_transforms = A.Compose([\n                              A.ShiftScaleRotate(rotate_limit=(-5, 5), p=0.5),\n                              A.RandomCrop(448, 448),\n                              A.RandomRotate90(p=1.0),\n                              A.OneOf([\n                                        A.GridDropout(ratio=0.4, p=0.5),\n                                        A.CoarseDropout(max_holes=25,\n                                                        max_height=int(0.2*448),\n                                                        max_width=int(0.2*448),\n                                                        min_holes=10,\n                                                        min_height=int(0.1*448),\n                                                        min_width=int(0.1*448),\n                                                        p=0.5),\n                                        A.GridDistortion(p=1.0),\n                                        ], p=0.5),\n                              ])\n\n```\n\nImages are getting resized to **512x512** during pre-processing. Rotation invariant training and in addition a left to right flip by simultaneously flipping the labels is applied. \n\n### **Sampling:**\nTo sample negative spots for training of Stage-1, image from sequences without an aneurysm but also images from sequences with aneurysms are used. In the second case, spots where an aneurysm is present are excluded with some boarders. Furthermore, based on the OOF predictions of an early trained classifier, spots where the model predicts false positives are sampled more often. \n\n![sampling](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2298475%2Fab28e796d8a103ce5e7ee38f5b1fc79b%2Fsampling.png?generation=1761583609064432&alt=media)\n\n## Stage 2:\nBased on the extracted features of Stage-1 a Transformer is trained on the complete sequence per UID. For augmentation in this stage, sequences with step-1 and step-2 are extracted and used during training of Stage-2. \n\n![stage2](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2298475%2Fa21b25418847a7b1d18bc3f16c0583f4%2Fstage2.png?generation=1761586827480711&alt=media)\n\n## Ensemble:\n- An ensemble of 4 Convnext Base models with two different step sizes is used.\n- On sequences > 192 images, only every n image is selected to shrink sequence length (n is calculated dynamically based on original sequence length).\n- To shrink the sequence length is only necessary because of the 12h time limit and the incredible crappy Kaggle hardware with totally random runtime based on the assigned server. \n- Nearly half of my submissions are simply timeouts, sorry Kaggle, but it can’t be that the same code without changes hit the 12h time limit or finishes in less than 8h.\n\n## Links:\n[Training](https://github.com/KonradHabel/rsna)\n\n[Inference](https://www.kaggle.com/code/khabel/rsna-dual-gpu-ensemble-2step)\n\n[Weights](https://www.kaggle.com/datasets/khabel/rsna-2025-weights-4xbase)\n",
      "votes": 7
    },
    {
      "id": 3452378,
      "postDate": "2026-05-03T07:58:52.387Z",
      "content": "<p>this is the kind of top solution someday i want to make. Very concise and creative. Thx for sharing ur code. thx!</p>",
      "rawMarkdown": "this is the kind of top solution someday i want to make. Very concise and creative. Thx for sharing ur code. thx!"
    }
  ],
  "comments": [
    {
      "id": 3452378,
      "author_name": "Simon Beck",
      "author_url": "",
      "post_date": "2026-05-03T07:58:52.387000",
      "content": "<p>this is the kind of top solution someday i want to make. Very concise and creative. Thx for sharing ur code. thx!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3307740": "## Approach:\n- Two stage approach to find aneurysms in sequences of MR/CT images without using resampling, 3D approaches or segmentation. \n- Approach the problem as outlier prediction task: we have only a small number of aneurysm spots in the training sequences, but can sample a huge number of negative spots, where no aneurysm is present.\n- **Stage-1:** Train a classifier on single images.\n- **Stage-2:** Train a Transformer to classify the whole sequence of images based on the extracted features of Stage-1.\n\n## Stage 1:\nBased on the given **train_localizers.csv** sample positive (aneurysm present) and negative (no aneurysm present) cases and train a simple image classifier. Use the surrounding images as R and B channel of the image. Two different step sizes are used for stacking. Step 2 simply means not stacking three consecutive frames but instead leaving a gap and taking the next frame. As backbone a Convnext Base Dinov3 is used with a simple classification head. \n\n![stage-1](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2298475%2Fcb023dda34b71ba52d5a862d3893bdf7%2Fstage1.png?generation=1761583583043279&alt=media)\n\n### **Augmentations:**\n```\nvalid_transforms = A.Compose([\n                              A.CenterCrop(448, 448),\n                              ]) \n\n\ntrain_transforms = A.Compose([\n                              A.ShiftScaleRotate(rotate_limit=(-5, 5), p=0.5),\n                              A.RandomCrop(448, 448),\n                              A.RandomRotate90(p=1.0),\n                              A.OneOf([\n                                        A.GridDropout(ratio=0.4, p=0.5),\n                                        A.CoarseDropout(max_holes=25,\n                                                        max_height=int(0.2*448),\n                                                        max_width=int(0.2*448),\n                                                        min_holes=10,\n                                                        min_height=int(0.1*448),\n                                                        min_width=int(0.1*448),\n                                                        p=0.5),\n                                        A.GridDistortion(p=1.0),\n                                        ], p=0.5),\n                              ])\n\n```\n\nImages are getting resized to **512x512** during pre-processing. Rotation invariant training and in addition a left to right flip by simultaneously flipping the labels is applied. \n\n### **Sampling:**\nTo sample negative spots for training of Stage-1, image from sequences without an aneurysm but also images from sequences with aneurysms are used. In the second case, spots where an aneurysm is present are excluded with some boarders. Furthermore, based on the OOF predictions of an early trained classifier, spots where the model predicts false positives are sampled more often. \n\n![sampling](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2298475%2Fab28e796d8a103ce5e7ee38f5b1fc79b%2Fsampling.png?generation=1761583609064432&alt=media)\n\n## Stage 2:\nBased on the extracted features of Stage-1 a Transformer is trained on the complete sequence per UID. For augmentation in this stage, sequences with step-1 and step-2 are extracted and used during training of Stage-2. \n\n![stage2](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2298475%2Fa21b25418847a7b1d18bc3f16c0583f4%2Fstage2.png?generation=1761586827480711&alt=media)\n\n## Ensemble:\n- An ensemble of 4 Convnext Base models with two different step sizes is used.\n- On sequences > 192 images, only every n image is selected to shrink sequence length (n is calculated dynamically based on original sequence length).\n- To shrink the sequence length is only necessary because of the 12h time limit and the incredible crappy Kaggle hardware with totally random runtime based on the assigned server. \n- Nearly half of my submissions are simply timeouts, sorry Kaggle, but it can’t be that the same code without changes hit the 12h time limit or finishes in less than 8h.\n\n## Links:\n[Training](https://github.com/KonradHabel/rsna)\n\n[Inference](https://www.kaggle.com/code/khabel/rsna-dual-gpu-ensemble-2step)\n\n[Weights](https://www.kaggle.com/datasets/khabel/rsna-2025-weights-4xbase)\n",
    "3452378": "this is the kind of top solution someday i want to make. Very concise and creative. Thx for sharing ur code. thx!"
  }
}