{
  "id": 611849,
  "title": "5th place solution with code",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/611849",
  "author_name": "HoangHuyen",
  "post_date": "2025-10-15T05:09:21.560000",
  "votes": 54,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Thanks to Kaggle and RSNA for hosting this exciting competition. <br>\nIt was a great learning experience and it was very interesting to see how much of my computer vision experience could also be applied to medical imaging. I’m quite disappointed with the result, but I see it as an opportunity to learn from other teams.</p>\n<p><strong>Github code</strong>: <a href=\"https://github.com/hoanghuyen797/RSNA-Intracranial-Aneurysm-Detection\" target=\"_blank\">https://github.com/hoanghuyen797/RSNA-Intracranial-Aneurysm-Detection</a><br>\n<strong>Inference notebook</strong>: <a href=\"https://www.kaggle.com/code/longb173/rsna-iad-final-nb?scriptVersionId=266774552\" target=\"_blank\">https://www.kaggle.com/code/longb173/rsna-iad-final-nb?scriptVersionId=266774552</a><br>\n<strong>Demo notebook</strong>: <a href=\"https://github.com/hoanghuyen797/RSNA-Intracranial-Aneurysm-Detection/blob/main/src/demo-test/test.ipynb\" target=\"_blank\">https://github.com/hoanghuyen797/RSNA-Intracranial-Aneurysm-Detection/blob/main/src/demo-test/test.ipynb</a></p>\n<h1>Overall  Pipeline</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F28100758%2F14753e43854508d3eb8e5bb82af6c77e%2Frsna_pipeline.drawio.png?generation=1760504882020544&amp;alt=media\" alt=\"\"></p>\n<h2>1. 2.5D image</h2>\n<p>For each slice t, I combine slices t-1 and t+1 to create a 3-channel image corresponding to [t-1, t, t+1]</p>\n<h2>2. Exp0: Aneurysm detection</h2>\n<p>In my experience, using only the classification labels (train.csv) is not as accurate as combining the classification and localization labels (train_localizers.csv). Therefore, based on the labels provided in train_localizers.csv by the host, for each aneurysm centroid, I searched within ±10 neighboring slices and manually annotate bounding boxes for the aneurysm using <a href=\"https://github.com/HumanSignal/labelImg\" target=\"_blank\">LabelImg</a>. <br>\nThis process does not require specialized medical knowledge since the aneurysm centroids are already provided.<br>\n2 classes: aneurysm (modality CTA, MRA, MRI T1post) and aneurysm_mri_t2 (modality MRI T2)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F28100758%2F43722e92d294e8b2c3d7681c7dfa961f%2Faneurysm_detection.png?generation=1760504909286014&amp;alt=media\" alt=\"\"><br>\nThen I train 5 models (5 folds) using YOLOv11x 1280</p>\n<table>\n<thead>\n<tr>\n<th>metric</th>\n<th>Fold 0</th>\n<th>Fold 1</th>\n<th>Fold 2</th>\n<th>Fold 3</th>\n<th>Fold 4</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>mAP50</td>\n<td>0.705</td>\n<td>0.647</td>\n<td>0.766</td>\n<td>0.702</td>\n<td>0.691</td>\n</tr>\n<tr>\n<td>mAP50-95</td>\n<td>0.460</td>\n<td>0.429</td>\n<td>0.504</td>\n<td>0.482</td>\n<td>0.449</td>\n</tr>\n</tbody>\n</table>\n<h2>3. Exp1: Brain detection</h2>\n<p>For each SeriesInstanceUID in the training set, I generate a single image by averaging all slices. <br>\nThen, I manually annotate the brain bounding box as the following 2 classes: brain (brain in axial view) and abnormal (brain in other views). Each slice in the series will be cropped according to the bounding box predicted by this model. This reduces background noise (especially for slices containing lung regions …), which improves the model’s accuracy by about <strong>0.03-0.05</strong>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F28100758%2F11c4dde4be41f43e33e293cc2c20d214%2Fbrain_det.png?generation=1760504934002847&amp;alt=media\" alt=\"\"><br>\nThen I train 1 model using yolov5n 640: mAP50-95 = 0.948</p>\n<h2>4. Image augmentation</h2>\n<p>`train_transform = albu.Compose([<br>\n    albu.RandomResizedCrop(size=(self.image_size, self.image_size), scale=(0.5, 1.0), ratio=(0.75, 1.3333), p=1),<br>\n    albu.ShiftScaleRotate(rotate_limit=15, border_mode=0, p=0.5),<br>\n    albu.OneOf([<br>\n        albu.MotionBlur(blur_limit=5),<br>\n        albu.MedianBlur(blur_limit=5),<br>\n        albu.GaussianBlur(blur_limit=5),<br>\n        albu.GaussNoise(var_limit=(5.0, 30.0)),<br>\n    ], p=0.5),<br>\n    albu.CLAHE(clip_limit=4.0, p=0.5),<br>\n    albu.HueSaturationValue(p=0.5),<br>\n    albu.RandomBrightnessContrast(p=0.5),</p>\n<p>val_transform = albu.Compose([<br>\n    albu.Resize(self.image_size, self.image_size),<br>\n])`</p>\n<p>Horizontal Flip: It may sound unreasonable, but I applied horizontal flipping to the images and adjusted the labels as following:<br>\nLeft Infraclinoid Internal Carotid Artery &lt;-&gt; Right Infraclinoid Internal Carotid Artery<br>\nLeft Supraclinoid Internal Carotid Artery &lt;-&gt;Right Supraclinoid Internal Carotid Artery<br>\nLeft Middle Cerebral Artery &lt;-&gt;Right Middle Cerebral Artery<br>\nLeft Anterior Cerebral Artery &lt;-&gt;Right Anterior Cerebral Artery<br>\nLeft Posterior Communicating Artery &lt;-&gt;Right Posterior Communicating Artery<br>\nThis worked, and my model’s accuracy improved by about 0.01</p>\n<h2>5. Exp2: 2 classification models</h2>\n<p>2 multi-label classification models (14 classes), image size=384, trained on this competition dataset. For negative series I use all slices, while for positive series I only use slices that contain aneurysm boxes (as reviewed in section 2). The label of each slice is the same as the series label in the train.csv<br>\nvit large 384: OOF AUC = 0.8503<br>\neva large 384: OOF AUC = 0.8551<br>\nDue to time constraints, I couldn't use 5 models (5 folds) for prediction. I could only use a single model trained on almost the full dataset, using only 50 series for evaluation, and selecting the best epoch based on the public leaderboard.</p>\n<h2>6. Exp3: Multi-task classification + segmentation</h2>\n<p>1 model (mit b4 fpn), image size = 384, trained on the RSNA dataset. Similar to exp2, for negative series I use all slices, while for positive series I only use slices that contain aneurysm boxes (as reviewed in section 2). The label of each slice is the same as the series label in the train.csv. For the submission, I only used the prediction from the classification task. To create masks for the segmentation task, I use the aneurysm bounding box (2). <br>\nmit-b4 FPN : OOF AUC = 0.8549</p>\n<h2>7. External dataset</h2>\n<p>I use the competition dataset and 2 external datasets:<br>\nLausanne_TOFMRA: <a href=\"https://openneuro.org/datasets/ds003949/versions/1.0.1\" target=\"_blank\">https://openneuro.org/datasets/ds003949/versions/1.0.1</a> <br>\nRoyal_Brisbane_TOFMRA: <a href=\"https://openneuro.org/datasets/ds005096/versions/1.0.3\" target=\"_blank\">https://openneuro.org/datasets/ds005096/versions/1.0.3</a><br>\nFor the external data, I generate the series label using prediction of 2 models (vit large exp2 + mit b4 exp3), and generate the localization label using prediction of model exp0-aneurysm detection  </p>\n<h2>8. Clean trainset</h2>\n<p>For negative series in the trainset with 'Aneurysm Present' prediction score of 2 models (vit large exp2 + mit b4 exp3) &gt; 0.9, I will change it to positive and use model exp0-aneurysm detection to create localization label</p>\n<h2>9. Exp4: 2 classification models</h2>\n<p>2 classification models trained on the cleaned RSNA dataset (8) and external dataset (Lausanne_TOFMRA + Royal_Brisbane_TOFMRA) pseudo labeling (7)<br>\nvit large 384: OOF AUC = 0.8558<br>\neva large 384: OOF AUC = 0.8579</p>\n<h2>10: Exp5:  Multi-task classification + segmentation</h2>\n<p>1 model (mit b4 fpn), image size = 384, trained on the cleaned RSNA dataset (8) and external dataset (Lausanne_TOFMRA + Royal_Brisbane_TOFMRA) pseudo labeling (7)<br>\nmit-b4 FPN 384: OOF AUC = 0.8629</p>\n<h2>11. Final submission</h2>\n<p>The final submission is the ensemble of 6 models:</p>\n<ul>\n<li>Final1: 0.25 exp3_mit_b4 + 0.25 exp5_mit_b4 + 0.125 exp2_vit_large + 0.125 exp4_vit_large + 0.125 exp2_eva_large + 0.125 exp2_eva_large</li>\n<li>Final2: 0.5 exp3_mit_b4 + 0.125 exp2_vit_large + 0.125 exp4_vit_large + 0.125 exp2_eva_large + 0.125 exp2_eva_large</li>\n</ul>\n<p>Final1: OOF AUC = 0.8823, public LB = 0.89<br>\nFinal2: OOF AUC = 0.8767, public LB = 0.89</p>",
  "messages": [
    {
      "id": 3302118,
      "postDate": "2025-10-15T05:09:21.560Z",
      "content": "<p>Thanks to Kaggle and RSNA for hosting this exciting competition. <br>\nIt was a great learning experience and it was very interesting to see how much of my computer vision experience could also be applied to medical imaging. I’m quite disappointed with the result, but I see it as an opportunity to learn from other teams.</p>\n<p><strong>Github code</strong>: <a href=\"https://github.com/hoanghuyen797/RSNA-Intracranial-Aneurysm-Detection\" target=\"_blank\">https://github.com/hoanghuyen797/RSNA-Intracranial-Aneurysm-Detection</a><br>\n<strong>Inference notebook</strong>: <a href=\"https://www.kaggle.com/code/longb173/rsna-iad-final-nb?scriptVersionId=266774552\" target=\"_blank\">https://www.kaggle.com/code/longb173/rsna-iad-final-nb?scriptVersionId=266774552</a><br>\n<strong>Demo notebook</strong>: <a href=\"https://github.com/hoanghuyen797/RSNA-Intracranial-Aneurysm-Detection/blob/main/src/demo-test/test.ipynb\" target=\"_blank\">https://github.com/hoanghuyen797/RSNA-Intracranial-Aneurysm-Detection/blob/main/src/demo-test/test.ipynb</a></p>\n<h1>Overall  Pipeline</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F28100758%2F14753e43854508d3eb8e5bb82af6c77e%2Frsna_pipeline.drawio.png?generation=1760504882020544&amp;alt=media\" alt=\"\"></p>\n<h2>1. 2.5D image</h2>\n<p>For each slice t, I combine slices t-1 and t+1 to create a 3-channel image corresponding to [t-1, t, t+1]</p>\n<h2>2. Exp0: Aneurysm detection</h2>\n<p>In my experience, using only the classification labels (train.csv) is not as accurate as combining the classification and localization labels (train_localizers.csv). Therefore, based on the labels provided in train_localizers.csv by the host, for each aneurysm centroid, I searched within ±10 neighboring slices and manually annotate bounding boxes for the aneurysm using <a href=\"https://github.com/HumanSignal/labelImg\" target=\"_blank\">LabelImg</a>. <br>\nThis process does not require specialized medical knowledge since the aneurysm centroids are already provided.<br>\n2 classes: aneurysm (modality CTA, MRA, MRI T1post) and aneurysm_mri_t2 (modality MRI T2)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F28100758%2F43722e92d294e8b2c3d7681c7dfa961f%2Faneurysm_detection.png?generation=1760504909286014&amp;alt=media\" alt=\"\"><br>\nThen I train 5 models (5 folds) using YOLOv11x 1280</p>\n<table>\n<thead>\n<tr>\n<th>metric</th>\n<th>Fold 0</th>\n<th>Fold 1</th>\n<th>Fold 2</th>\n<th>Fold 3</th>\n<th>Fold 4</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>mAP50</td>\n<td>0.705</td>\n<td>0.647</td>\n<td>0.766</td>\n<td>0.702</td>\n<td>0.691</td>\n</tr>\n<tr>\n<td>mAP50-95</td>\n<td>0.460</td>\n<td>0.429</td>\n<td>0.504</td>\n<td>0.482</td>\n<td>0.449</td>\n</tr>\n</tbody>\n</table>\n<h2>3. Exp1: Brain detection</h2>\n<p>For each SeriesInstanceUID in the training set, I generate a single image by averaging all slices. <br>\nThen, I manually annotate the brain bounding box as the following 2 classes: brain (brain in axial view) and abnormal (brain in other views). Each slice in the series will be cropped according to the bounding box predicted by this model. This reduces background noise (especially for slices containing lung regions …), which improves the model’s accuracy by about <strong>0.03-0.05</strong>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F28100758%2F11c4dde4be41f43e33e293cc2c20d214%2Fbrain_det.png?generation=1760504934002847&amp;alt=media\" alt=\"\"><br>\nThen I train 1 model using yolov5n 640: mAP50-95 = 0.948</p>\n<h2>4. Image augmentation</h2>\n<p>`train_transform = albu.Compose([<br>\n    albu.RandomResizedCrop(size=(self.image_size, self.image_size), scale=(0.5, 1.0), ratio=(0.75, 1.3333), p=1),<br>\n    albu.ShiftScaleRotate(rotate_limit=15, border_mode=0, p=0.5),<br>\n    albu.OneOf([<br>\n        albu.MotionBlur(blur_limit=5),<br>\n        albu.MedianBlur(blur_limit=5),<br>\n        albu.GaussianBlur(blur_limit=5),<br>\n        albu.GaussNoise(var_limit=(5.0, 30.0)),<br>\n    ], p=0.5),<br>\n    albu.CLAHE(clip_limit=4.0, p=0.5),<br>\n    albu.HueSaturationValue(p=0.5),<br>\n    albu.RandomBrightnessContrast(p=0.5),</p>\n<p>val_transform = albu.Compose([<br>\n    albu.Resize(self.image_size, self.image_size),<br>\n])`</p>\n<p>Horizontal Flip: It may sound unreasonable, but I applied horizontal flipping to the images and adjusted the labels as following:<br>\nLeft Infraclinoid Internal Carotid Artery &lt;-&gt; Right Infraclinoid Internal Carotid Artery<br>\nLeft Supraclinoid Internal Carotid Artery &lt;-&gt;Right Supraclinoid Internal Carotid Artery<br>\nLeft Middle Cerebral Artery &lt;-&gt;Right Middle Cerebral Artery<br>\nLeft Anterior Cerebral Artery &lt;-&gt;Right Anterior Cerebral Artery<br>\nLeft Posterior Communicating Artery &lt;-&gt;Right Posterior Communicating Artery<br>\nThis worked, and my model’s accuracy improved by about 0.01</p>\n<h2>5. Exp2: 2 classification models</h2>\n<p>2 multi-label classification models (14 classes), image size=384, trained on this competition dataset. For negative series I use all slices, while for positive series I only use slices that contain aneurysm boxes (as reviewed in section 2). The label of each slice is the same as the series label in the train.csv<br>\nvit large 384: OOF AUC = 0.8503<br>\neva large 384: OOF AUC = 0.8551<br>\nDue to time constraints, I couldn't use 5 models (5 folds) for prediction. I could only use a single model trained on almost the full dataset, using only 50 series for evaluation, and selecting the best epoch based on the public leaderboard.</p>\n<h2>6. Exp3: Multi-task classification + segmentation</h2>\n<p>1 model (mit b4 fpn), image size = 384, trained on the RSNA dataset. Similar to exp2, for negative series I use all slices, while for positive series I only use slices that contain aneurysm boxes (as reviewed in section 2). The label of each slice is the same as the series label in the train.csv. For the submission, I only used the prediction from the classification task. To create masks for the segmentation task, I use the aneurysm bounding box (2). <br>\nmit-b4 FPN : OOF AUC = 0.8549</p>\n<h2>7. External dataset</h2>\n<p>I use the competition dataset and 2 external datasets:<br>\nLausanne_TOFMRA: <a href=\"https://openneuro.org/datasets/ds003949/versions/1.0.1\" target=\"_blank\">https://openneuro.org/datasets/ds003949/versions/1.0.1</a> <br>\nRoyal_Brisbane_TOFMRA: <a href=\"https://openneuro.org/datasets/ds005096/versions/1.0.3\" target=\"_blank\">https://openneuro.org/datasets/ds005096/versions/1.0.3</a><br>\nFor the external data, I generate the series label using prediction of 2 models (vit large exp2 + mit b4 exp3), and generate the localization label using prediction of model exp0-aneurysm detection  </p>\n<h2>8. Clean trainset</h2>\n<p>For negative series in the trainset with 'Aneurysm Present' prediction score of 2 models (vit large exp2 + mit b4 exp3) &gt; 0.9, I will change it to positive and use model exp0-aneurysm detection to create localization label</p>\n<h2>9. Exp4: 2 classification models</h2>\n<p>2 classification models trained on the cleaned RSNA dataset (8) and external dataset (Lausanne_TOFMRA + Royal_Brisbane_TOFMRA) pseudo labeling (7)<br>\nvit large 384: OOF AUC = 0.8558<br>\neva large 384: OOF AUC = 0.8579</p>\n<h2>10: Exp5:  Multi-task classification + segmentation</h2>\n<p>1 model (mit b4 fpn), image size = 384, trained on the cleaned RSNA dataset (8) and external dataset (Lausanne_TOFMRA + Royal_Brisbane_TOFMRA) pseudo labeling (7)<br>\nmit-b4 FPN 384: OOF AUC = 0.8629</p>\n<h2>11. Final submission</h2>\n<p>The final submission is the ensemble of 6 models:</p>\n<ul>\n<li>Final1: 0.25 exp3_mit_b4 + 0.25 exp5_mit_b4 + 0.125 exp2_vit_large + 0.125 exp4_vit_large + 0.125 exp2_eva_large + 0.125 exp2_eva_large</li>\n<li>Final2: 0.5 exp3_mit_b4 + 0.125 exp2_vit_large + 0.125 exp4_vit_large + 0.125 exp2_eva_large + 0.125 exp2_eva_large</li>\n</ul>\n<p>Final1: OOF AUC = 0.8823, public LB = 0.89<br>\nFinal2: OOF AUC = 0.8767, public LB = 0.89</p>",
      "rawMarkdown": "Thanks to Kaggle and RSNA for hosting this exciting competition. \nIt was a great learning experience and it was very interesting to see how much of my computer vision experience could also be applied to medical imaging. I’m quite disappointed with the result, but I see it as an opportunity to learn from other teams.\n\n**Github code**: https://github.com/hoanghuyen797/RSNA-Intracranial-Aneurysm-Detection\n**Inference notebook**: https://www.kaggle.com/code/longb173/rsna-iad-final-nb?scriptVersionId=266774552\n**Demo notebook**: https://github.com/hoanghuyen797/RSNA-Intracranial-Aneurysm-Detection/blob/main/src/demo-test/test.ipynb\n\n#Overall  Pipeline\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F28100758%2F14753e43854508d3eb8e5bb82af6c77e%2Frsna_pipeline.drawio.png?generation=1760504882020544&alt=media)\n## 1. 2.5D image \nFor each slice t, I combine slices t-1 and t+1 to create a 3-channel image corresponding to [t-1, t, t+1]\n\n## 2. Exp0: Aneurysm detection\nIn my experience, using only the classification labels (train.csv) is not as accurate as combining the classification and localization labels (train_localizers.csv). Therefore, based on the labels provided in train_localizers.csv by the host, for each aneurysm centroid, I searched within ±10 neighboring slices and manually annotate bounding boxes for the aneurysm using [LabelImg](https://github.com/HumanSignal/labelImg). \nThis process does not require specialized medical knowledge since the aneurysm centroids are already provided.\n2 classes: aneurysm (modality CTA, MRA, MRI T1post) and aneurysm_mri_t2 (modality MRI T2)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F28100758%2F43722e92d294e8b2c3d7681c7dfa961f%2Faneurysm_detection.png?generation=1760504909286014&alt=media)\nThen I train 5 models (5 folds) using YOLOv11x 1280\nmetric   | Fold 0 | Fold 1 | Fold 2 | Fold 3 | Fold 4\n:-------:|:------:|:------:|:------:|:------:|:------:\nmAP50    | 0.705  | 0.647  | 0.766  | 0.702  | 0.691\nmAP50-95 | 0.460  | 0.429  | 0.504  | 0.482  | 0.449\n\n## 3. Exp1: Brain detection\nFor each SeriesInstanceUID in the training set, I generate a single image by averaging all slices. \nThen, I manually annotate the brain bounding box as the following 2 classes: brain (brain in axial view) and abnormal (brain in other views). Each slice in the series will be cropped according to the bounding box predicted by this model. This reduces background noise (especially for slices containing lung regions ...), which improves the model’s accuracy by about **0.03-0.05**.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F28100758%2F11c4dde4be41f43e33e293cc2c20d214%2Fbrain_det.png?generation=1760504934002847&alt=media)\nThen I train 1 model using yolov5n 640: mAP50-95 = 0.948\n\n## 4. Image augmentation\n`train_transform = albu.Compose([\n    albu.RandomResizedCrop(size=(self.image_size, self.image_size), scale=(0.5, 1.0), ratio=(0.75, 1.3333), p=1),\n    albu.ShiftScaleRotate(rotate_limit=15, border_mode=0, p=0.5),\n    albu.OneOf([\n        albu.MotionBlur(blur_limit=5),\n        albu.MedianBlur(blur_limit=5),\n        albu.GaussianBlur(blur_limit=5),\n        albu.GaussNoise(var_limit=(5.0, 30.0)),\n    ], p=0.5),\n    albu.CLAHE(clip_limit=4.0, p=0.5),\n    albu.HueSaturationValue(p=0.5),\n    albu.RandomBrightnessContrast(p=0.5),\n\nval_transform = albu.Compose([\n    albu.Resize(self.image_size, self.image_size),\n])`\n\nHorizontal Flip: It may sound unreasonable, but I applied horizontal flipping to the images and adjusted the labels as following:\nLeft Infraclinoid Internal Carotid Artery <-> Right Infraclinoid Internal Carotid Artery\nLeft Supraclinoid Internal Carotid Artery <->Right Supraclinoid Internal Carotid Artery\nLeft Middle Cerebral Artery <->Right Middle Cerebral Artery\nLeft Anterior Cerebral Artery <->Right Anterior Cerebral Artery\nLeft Posterior Communicating Artery <->Right Posterior Communicating Artery\nThis worked, and my model’s accuracy improved by about 0.01\n\n## 5. Exp2: 2 classification models\n2 multi-label classification models (14 classes), image size=384, trained on this competition dataset. For negative series I use all slices, while for positive series I only use slices that contain aneurysm boxes (as reviewed in section 2). The label of each slice is the same as the series label in the train.csv\nvit large 384: OOF AUC = 0.8503\neva large 384: OOF AUC = 0.8551\nDue to time constraints, I couldn't use 5 models (5 folds) for prediction. I could only use a single model trained on almost the full dataset, using only 50 series for evaluation, and selecting the best epoch based on the public leaderboard.\n\n## 6. Exp3: Multi-task classification + segmentation\n1 model (mit b4 fpn), image size = 384, trained on the RSNA dataset. Similar to exp2, for negative series I use all slices, while for positive series I only use slices that contain aneurysm boxes (as reviewed in section 2). The label of each slice is the same as the series label in the train.csv. For the submission, I only used the prediction from the classification task. To create masks for the segmentation task, I use the aneurysm bounding box (2). \nmit-b4 FPN : OOF AUC = 0.8549\n\n## 7. External dataset \nI use the competition dataset and 2 external datasets:\nLausanne_TOFMRA: https://openneuro.org/datasets/ds003949/versions/1.0.1 \nRoyal_Brisbane_TOFMRA: https://openneuro.org/datasets/ds005096/versions/1.0.3\nFor the external data, I generate the series label using prediction of 2 models (vit large exp2 + mit b4 exp3), and generate the localization label using prediction of model exp0-aneurysm detection  \n\n## 8. Clean trainset\nFor negative series in the trainset with 'Aneurysm Present' prediction score of 2 models (vit large exp2 + mit b4 exp3) > 0.9, I will change it to positive and use model exp0-aneurysm detection to create localization label\n\n## 9. Exp4: 2 classification models\n2 classification models trained on the cleaned RSNA dataset (8) and external dataset (Lausanne_TOFMRA + Royal_Brisbane_TOFMRA) pseudo labeling (7)\nvit large 384: OOF AUC = 0.8558\neva large 384: OOF AUC = 0.8579\n\n## 10: Exp5:  Multi-task classification + segmentation\n1 model (mit b4 fpn), image size = 384, trained on the cleaned RSNA dataset (8) and external dataset (Lausanne_TOFMRA + Royal_Brisbane_TOFMRA) pseudo labeling (7)\nmit-b4 FPN 384: OOF AUC = 0.8629\n\n## 11. Final submission\nThe final submission is the ensemble of 6 models:\n- Final1: 0.25 exp3_mit_b4 + 0.25 exp5_mit_b4 + 0.125 exp2_vit_large + 0.125 exp4_vit_large + 0.125 exp2_eva_large + 0.125 exp2_eva_large\n- Final2: 0.5 exp3_mit_b4 + 0.125 exp2_vit_large + 0.125 exp4_vit_large + 0.125 exp2_eva_large + 0.125 exp2_eva_large\n\nFinal1: OOF AUC = 0.8823, public LB = 0.89\nFinal2: OOF AUC = 0.8767, public LB = 0.89",
      "votes": 54
    },
    {
      "id": 3325808,
      "postDate": "2025-11-15T07:27:47.070Z",
      "content": "<p>Hello, could you specify whether the data splitting employs k-fold cross-validation, or if a simple random split is used? </p>",
      "rawMarkdown": "Hello, could you specify whether the data splitting employs k-fold cross-validation, or if a simple random split is used? "
    },
    {
      "id": 3302701,
      "postDate": "2025-10-16T09:51:54.557Z",
      "content": "<p>First part is extremely effort, consider only you in the team. Good work.</p>",
      "rawMarkdown": "First part is extremely effort, consider only you in the team. Good work."
    },
    {
      "id": 3302161,
      "postDate": "2025-10-15T08:05:23.100Z",
      "content": "<p>I would like to ask how the final model ensemble weights were calculated?</p>",
      "rawMarkdown": "I would like to ask how the final model ensemble weights were calculated?",
      "replies": [
        {
          "id": 3302166,
          "postDate": "2025-10-15T08:22:06Z",
          "content": "<p>try a few weight combinations to achieve the highest OOF AUC and the highest public LB</p>",
          "rawMarkdown": "try a few weight combinations to achieve the highest OOF AUC and the highest public LB"
        }
      ]
    },
    {
      "id": 3302136,
      "postDate": "2025-10-15T06:31:02.227Z",
      "content": "<p>Congratulations! This is a great solution , 2.5D seems to work wonders in this competition. I have a few questions (some may sound stupid, sorry) - </p>\n<ol>\n<li>Is there a reason why you manually annotated bboxes, and also separated MRI T2 aneurysms as a separate class? I trained a similar yolo12L to detect aneurysms , searching in 7 neighboring slices using 1D images (each slice as image) and it could get mAP50 0.8+. </li>\n<li>~0.85 AUC using only classification models is really impressive (exp2). Are they trained using 2.5D slices only? </li>\n<li>How do you do inference with the slice-wise classification models? Do you infer on all slices , and then aggregate predictions? </li>\n</ol>\n<p>Btw, the codebase is really concise, well maintained and easy to read. A lot to learn from you.</p>\n<p>Thank you, and congratulations again!</p>",
      "rawMarkdown": "Congratulations! This is a great solution , 2.5D seems to work wonders in this competition. I have a few questions (some may sound stupid, sorry) - \n1. Is there a reason why you manually annotated bboxes, and also separated MRI T2 aneurysms as a separate class? I trained a similar yolo12L to detect aneurysms , searching in 7 neighboring slices using 1D images (each slice as image) and it could get mAP50 0.8+. \n2.  ~0.85 AUC using only classification models is really impressive (exp2). Are they trained using 2.5D slices only? \n3. How do you do inference with the slice-wise classification models? Do you infer on all slices , and then aggregate predictions? \n\nBtw, the codebase is really concise, well maintained and easy to read. A lot to learn from you.\n\nThank you, and congratulations again!",
      "replies": [
        {
          "id": 3302137,
          "postDate": "2025-10-15T06:31:54.157Z",
          "content": "<p>Also , you mentioned in another discussion that you are renting GPUs. I was doing the same, just curious what GPU you rented, and how much did it cost you overall ( if you are ok sharing the information) :D</p>",
          "rawMarkdown": "Also , you mentioned in another discussion that you are renting GPUs. I was doing the same, just curious what GPU you rented, and how much did it cost you overall ( if you are ok sharing the information) :D",
          "replies": [
            {
              "id": 3302146,
              "postDate": "2025-10-15T06:56:00.787Z",
              "content": "<p>~$700, if we pretend sleep doesn’t matter 😂</p>",
              "rawMarkdown": "~$700, if we pretend sleep doesn’t matter 😂\n",
              "votes": 1
            },
            {
              "id": 3302148,
              "postDate": "2025-10-15T06:59:11.503Z",
              "content": "<p>omg hahaha, that's quite alot! May I know what platform you used, and which GPU? I used vast.ai , and mostly 2x5090 / 1xRTX 6000 Pro for all my tasks. No more than $150 overall :3. I really was stingy about renting large amounts of disk space since it made things much more expensive, looking back I do have some regrets now</p>",
              "rawMarkdown": "omg hahaha, that's quite alot! May I know what platform you used, and which GPU? I used vast.ai , and mostly 2x5090 / 1xRTX 6000 Pro for all my tasks. No more than $150 overall :3. I really was stingy about renting large amounts of disk space since it made things much more expensive, looking back I do have some regrets now"
            }
          ]
        },
        {
          "id": 3302144,
          "postDate": "2025-10-15T06:53:36.160Z",
          "content": "<ol>\n<li>In MRI T2, the aneurysm seems dark, unlike in other modalities where it appears bright (based on the observed centroids).</li>\n<li>Yes 2.5D, you can check it on my GitHub repo</li>\n<li>Yes, and the series prediction is the maximum in z-axis (slice): num_slicesx14 -&gt; 14 probs</li>\n</ol>",
          "rawMarkdown": "1. In MRI T2, the aneurysm seems dark, unlike in other modalities where it appears bright (based on the observed centroids).\n2. Yes 2.5D, you can check it on my GitHub repo\n3. Yes, and the series prediction is the maximum in z-axis (slice): num_slicesx14 -> 14 probs",
          "votes": 1
        }
      ]
    },
    {
      "id": 3302130,
      "postDate": "2025-10-15T06:16:07.807Z",
      "content": "<p>Congratulations on your great results, nice work! Regarding using the model to clean the dataset, how many cases did you clean in total?</p>",
      "rawMarkdown": "Congratulations on your great results, nice work! Regarding using the model to clean the dataset, how many cases did you clean in total?",
      "replies": [
        {
          "id": 3302135,
          "postDate": "2025-10-15T06:29:03.127Z",
          "content": "<p>33 negative cases</p>",
          "rawMarkdown": "33 negative cases"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3325808,
      "author_name": "Seeing Times",
      "author_url": "",
      "post_date": "2025-11-15T07:27:47.070000",
      "content": "<p>Hello, could you specify whether the data splitting employs k-fold cross-validation, or if a simple random split is used? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3302701,
      "author_name": "Tom",
      "author_url": "",
      "post_date": "2025-10-16T09:51:54.557000",
      "content": "<p>First part is extremely effort, consider only you in the team. Good work.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3302161,
      "author_name": "Hide on bush",
      "author_url": "",
      "post_date": "2025-10-15T08:05:23.100000",
      "content": "<p>I would like to ask how the final model ensemble weights were calculated?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3302166,
          "author_name": "HoangHuyen",
          "author_url": "",
          "post_date": "2025-10-15T08:22:06",
          "content": "<p>try a few weight combinations to achieve the highest OOF AUC and the highest public LB</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3302136,
      "author_name": "Satwik",
      "author_url": "",
      "post_date": "2025-10-15T06:31:02.227000",
      "content": "<p>Congratulations! This is a great solution , 2.5D seems to work wonders in this competition. I have a few questions (some may sound stupid, sorry) - </p>\n<ol>\n<li>Is there a reason why you manually annotated bboxes, and also separated MRI T2 aneurysms as a separate class? I trained a similar yolo12L to detect aneurysms , searching in 7 neighboring slices using 1D images (each slice as image) and it could get mAP50 0.8+. </li>\n<li>~0.85 AUC using only classification models is really impressive (exp2). Are they trained using 2.5D slices only? </li>\n<li>How do you do inference with the slice-wise classification models? Do you infer on all slices , and then aggregate predictions? </li>\n</ol>\n<p>Btw, the codebase is really concise, well maintained and easy to read. A lot to learn from you.</p>\n<p>Thank you, and congratulations again!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3302137,
          "author_name": "Satwik",
          "author_url": "",
          "post_date": "2025-10-15T06:31:54.157000",
          "content": "<p>Also , you mentioned in another discussion that you are renting GPUs. I was doing the same, just curious what GPU you rented, and how much did it cost you overall ( if you are ok sharing the information) :D</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3302146,
              "author_name": "HoangHuyen",
              "author_url": "",
              "post_date": "2025-10-15T06:56:00.787000",
              "content": "<p>~$700, if we pretend sleep doesn’t matter 😂</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3302148,
              "author_name": "Satwik",
              "author_url": "",
              "post_date": "2025-10-15T06:59:11.503000",
              "content": "<p>omg hahaha, that's quite alot! May I know what platform you used, and which GPU? I used vast.ai , and mostly 2x5090 / 1xRTX 6000 Pro for all my tasks. No more than $150 overall :3. I really was stingy about renting large amounts of disk space since it made things much more expensive, looking back I do have some regrets now</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3302144,
          "author_name": "HoangHuyen",
          "author_url": "",
          "post_date": "2025-10-15T06:53:36.160000",
          "content": "<ol>\n<li>In MRI T2, the aneurysm seems dark, unlike in other modalities where it appears bright (based on the observed centroids).</li>\n<li>Yes 2.5D, you can check it on my GitHub repo</li>\n<li>Yes, and the series prediction is the maximum in z-axis (slice): num_slicesx14 -&gt; 14 probs</li>\n</ol>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3302130,
      "author_name": "Shuolin Liu",
      "author_url": "",
      "post_date": "2025-10-15T06:16:07.807000",
      "content": "<p>Congratulations on your great results, nice work! Regarding using the model to clean the dataset, how many cases did you clean in total?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3302135,
          "author_name": "HoangHuyen",
          "author_url": "",
          "post_date": "2025-10-15T06:29:03.127000",
          "content": "<p>33 negative cases</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3302118": "Thanks to Kaggle and RSNA for hosting this exciting competition. \nIt was a great learning experience and it was very interesting to see how much of my computer vision experience could also be applied to medical imaging. I’m quite disappointed with the result, but I see it as an opportunity to learn from other teams.\n\n**Github code**: https://github.com/hoanghuyen797/RSNA-Intracranial-Aneurysm-Detection\n**Inference notebook**: https://www.kaggle.com/code/longb173/rsna-iad-final-nb?scriptVersionId=266774552\n**Demo notebook**: https://github.com/hoanghuyen797/RSNA-Intracranial-Aneurysm-Detection/blob/main/src/demo-test/test.ipynb\n\n#Overall  Pipeline\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F28100758%2F14753e43854508d3eb8e5bb82af6c77e%2Frsna_pipeline.drawio.png?generation=1760504882020544&alt=media)\n## 1. 2.5D image \nFor each slice t, I combine slices t-1 and t+1 to create a 3-channel image corresponding to [t-1, t, t+1]\n\n## 2. Exp0: Aneurysm detection\nIn my experience, using only the classification labels (train.csv) is not as accurate as combining the classification and localization labels (train_localizers.csv). Therefore, based on the labels provided in train_localizers.csv by the host, for each aneurysm centroid, I searched within ±10 neighboring slices and manually annotate bounding boxes for the aneurysm using [LabelImg](https://github.com/HumanSignal/labelImg). \nThis process does not require specialized medical knowledge since the aneurysm centroids are already provided.\n2 classes: aneurysm (modality CTA, MRA, MRI T1post) and aneurysm_mri_t2 (modality MRI T2)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F28100758%2F43722e92d294e8b2c3d7681c7dfa961f%2Faneurysm_detection.png?generation=1760504909286014&alt=media)\nThen I train 5 models (5 folds) using YOLOv11x 1280\nmetric   | Fold 0 | Fold 1 | Fold 2 | Fold 3 | Fold 4\n:-------:|:------:|:------:|:------:|:------:|:------:\nmAP50    | 0.705  | 0.647  | 0.766  | 0.702  | 0.691\nmAP50-95 | 0.460  | 0.429  | 0.504  | 0.482  | 0.449\n\n## 3. Exp1: Brain detection\nFor each SeriesInstanceUID in the training set, I generate a single image by averaging all slices. \nThen, I manually annotate the brain bounding box as the following 2 classes: brain (brain in axial view) and abnormal (brain in other views). Each slice in the series will be cropped according to the bounding box predicted by this model. This reduces background noise (especially for slices containing lung regions ...), which improves the model’s accuracy by about **0.03-0.05**.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F28100758%2F11c4dde4be41f43e33e293cc2c20d214%2Fbrain_det.png?generation=1760504934002847&alt=media)\nThen I train 1 model using yolov5n 640: mAP50-95 = 0.948\n\n## 4. Image augmentation\n`train_transform = albu.Compose([\n    albu.RandomResizedCrop(size=(self.image_size, self.image_size), scale=(0.5, 1.0), ratio=(0.75, 1.3333), p=1),\n    albu.ShiftScaleRotate(rotate_limit=15, border_mode=0, p=0.5),\n    albu.OneOf([\n        albu.MotionBlur(blur_limit=5),\n        albu.MedianBlur(blur_limit=5),\n        albu.GaussianBlur(blur_limit=5),\n        albu.GaussNoise(var_limit=(5.0, 30.0)),\n    ], p=0.5),\n    albu.CLAHE(clip_limit=4.0, p=0.5),\n    albu.HueSaturationValue(p=0.5),\n    albu.RandomBrightnessContrast(p=0.5),\n\nval_transform = albu.Compose([\n    albu.Resize(self.image_size, self.image_size),\n])`\n\nHorizontal Flip: It may sound unreasonable, but I applied horizontal flipping to the images and adjusted the labels as following:\nLeft Infraclinoid Internal Carotid Artery <-> Right Infraclinoid Internal Carotid Artery\nLeft Supraclinoid Internal Carotid Artery <->Right Supraclinoid Internal Carotid Artery\nLeft Middle Cerebral Artery <->Right Middle Cerebral Artery\nLeft Anterior Cerebral Artery <->Right Anterior Cerebral Artery\nLeft Posterior Communicating Artery <->Right Posterior Communicating Artery\nThis worked, and my model’s accuracy improved by about 0.01\n\n## 5. Exp2: 2 classification models\n2 multi-label classification models (14 classes), image size=384, trained on this competition dataset. For negative series I use all slices, while for positive series I only use slices that contain aneurysm boxes (as reviewed in section 2). The label of each slice is the same as the series label in the train.csv\nvit large 384: OOF AUC = 0.8503\neva large 384: OOF AUC = 0.8551\nDue to time constraints, I couldn't use 5 models (5 folds) for prediction. I could only use a single model trained on almost the full dataset, using only 50 series for evaluation, and selecting the best epoch based on the public leaderboard.\n\n## 6. Exp3: Multi-task classification + segmentation\n1 model (mit b4 fpn), image size = 384, trained on the RSNA dataset. Similar to exp2, for negative series I use all slices, while for positive series I only use slices that contain aneurysm boxes (as reviewed in section 2). The label of each slice is the same as the series label in the train.csv. For the submission, I only used the prediction from the classification task. To create masks for the segmentation task, I use the aneurysm bounding box (2). \nmit-b4 FPN : OOF AUC = 0.8549\n\n## 7. External dataset \nI use the competition dataset and 2 external datasets:\nLausanne_TOFMRA: https://openneuro.org/datasets/ds003949/versions/1.0.1 \nRoyal_Brisbane_TOFMRA: https://openneuro.org/datasets/ds005096/versions/1.0.3\nFor the external data, I generate the series label using prediction of 2 models (vit large exp2 + mit b4 exp3), and generate the localization label using prediction of model exp0-aneurysm detection  \n\n## 8. Clean trainset\nFor negative series in the trainset with 'Aneurysm Present' prediction score of 2 models (vit large exp2 + mit b4 exp3) > 0.9, I will change it to positive and use model exp0-aneurysm detection to create localization label\n\n## 9. Exp4: 2 classification models\n2 classification models trained on the cleaned RSNA dataset (8) and external dataset (Lausanne_TOFMRA + Royal_Brisbane_TOFMRA) pseudo labeling (7)\nvit large 384: OOF AUC = 0.8558\neva large 384: OOF AUC = 0.8579\n\n## 10: Exp5:  Multi-task classification + segmentation\n1 model (mit b4 fpn), image size = 384, trained on the cleaned RSNA dataset (8) and external dataset (Lausanne_TOFMRA + Royal_Brisbane_TOFMRA) pseudo labeling (7)\nmit-b4 FPN 384: OOF AUC = 0.8629\n\n## 11. Final submission\nThe final submission is the ensemble of 6 models:\n- Final1: 0.25 exp3_mit_b4 + 0.25 exp5_mit_b4 + 0.125 exp2_vit_large + 0.125 exp4_vit_large + 0.125 exp2_eva_large + 0.125 exp2_eva_large\n- Final2: 0.5 exp3_mit_b4 + 0.125 exp2_vit_large + 0.125 exp4_vit_large + 0.125 exp2_eva_large + 0.125 exp2_eva_large\n\nFinal1: OOF AUC = 0.8823, public LB = 0.89\nFinal2: OOF AUC = 0.8767, public LB = 0.89",
    "3325808": "Hello, could you specify whether the data splitting employs k-fold cross-validation, or if a simple random split is used? ",
    "3302701": "First part is extremely effort, consider only you in the team. Good work.",
    "3302161": "I would like to ask how the final model ensemble weights were calculated?",
    "3302136": "Congratulations! This is a great solution , 2.5D seems to work wonders in this competition. I have a few questions (some may sound stupid, sorry) - \n1. Is there a reason why you manually annotated bboxes, and also separated MRI T2 aneurysms as a separate class? I trained a similar yolo12L to detect aneurysms , searching in 7 neighboring slices using 1D images (each slice as image) and it could get mAP50 0.8+. \n2.  ~0.85 AUC using only classification models is really impressive (exp2). Are they trained using 2.5D slices only? \n3. How do you do inference with the slice-wise classification models? Do you infer on all slices , and then aggregate predictions? \n\nBtw, the codebase is really concise, well maintained and easy to read. A lot to learn from you.\n\nThank you, and congratulations again!",
    "3302130": "Congratulations on your great results, nice work! Regarding using the model to clean the dataset, how many cases did you clean in total?"
  }
}