{
  "id": 612006,
  "title": "18th Place Solution",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/612006",
  "author_name": "koooeo",
  "post_date": "2025-10-16T05:52:21.062000",
  "votes": 12,
  "comment_count": 7,
  "views": 0,
  "content": "<h1>First of all</h1>\n<p>I would like to express my gratitude to the hosts for providing such a rewarding competition, to the Kaggle staff, and to all the participants who shared their insights in the discussions.</p>\n<h1>Strategy</h1>\n<p>From the early stage of the competition, based on the provided dataset, I devised a strategy to generate ROIs using a segmentation model and then classify them.</p>\n<h1>Pipeline overview</h1>\n<ol>\n<li>Stage1: Vessel Mask Prediction and ROI at the vessel-class level Extraction by a 2.5D Segmentation Model</li>\n<li>Stage2: ROI Classification</li>\n<li>Stage3: ROI Scoring and Series-level aggregation</li>\n</ol>\n<h1>Learning Process</h1>\n<h2>Stage1 (Segment Model)</h2>\n<ul>\n<li>We performed slice-level mask prediction using smp.Unet with an EfficientNet-B0 backbone.</li>\n<li>For preprocessing, windowing was applied to CT images and percentile normalization to MRI images. Basic data augmentation techniques were also used.</li>\n<li>The input size was (C, H, W) = (3, 512, 512).</li>\n<li>To address class imbalance among negative, positive, and vessel-specific positive samples, we used a random sampler with slice-level weighting.</li>\n<li>Using 3-channel inputs achieved better accuracy than 1-channel inputs.</li>\n<li>The final Dice score was approximately 0.66.</li>\n</ul>\n<h2>Stage1 (Vessel-class-level ROI Extraction）</h2>\n<ul>\n<li>We extracted ROIs using a segmentation model.<br>\nEach ROI was generated for every vessel class as the minimum bounding rectangle of the predicted mask with an added margin.<br>\nWhen multiple separated regions of the same class were detected within a slice, individual ROIs were created for each object (e.g., class_i_object_j).</li>\n<li>For aneurysm-positive slices, ROIs were extracted according to the above definition, and those containing the aneurysm coordinates were labeled as positive ROIs, achieving approximately 94% hit rate.</li>\n<li>Negative ROIs were sampled from both negative series and regions located more than 40 mm (in xyz space) away from aneurysm coordinates in positive series.</li>\n</ul>\n<h2>Stage2</h2>\n<ul>\n<li>For ROI classification, a 2.5D input was created by applying the same bounding box to the ROI slice and its adjacent slices.<br>\nThe image was resized while maintaining its aspect ratio, by scaling it so that the longer side matched the target size.<br>\nThe final input size was (C, H, W) = (3, 224, 224), and padding was applied as necessary.</li>\n<li>Multiple backbone networks were evaluated, including EfficientNet-B0 and EfficientNet-V2-S.<br>\nThe model achieved an ROI-level AUC of approximately 0.94.</li>\n</ul>\n<h2>Stage3</h2>\n<ul>\n<li>For each ROI, the aneurysm probability (prob) predicted by the ROI classification model was multiplied by the vessel class probabilities (class_scores) obtained from the segmentation model to compute the 13 class-wise scores.<br>\nThe binary aneurysm score was taken directly from the ROI probability (prob) without any modification.</li>\n<li>Series-level aggregation was performed using the top-k mean of ROI-level scores.</li>\n<li>The local validation AUC reached approximately 0.81.</li>\n</ul>\n<h1>Submission Process</h1>\n<ul>\n<li>The DICOM slices were aligned in the LPS coordinate system, and only the last 150 mm along the z-axis (corresponding to the head region) were used for inference.<br>\nTo limit the number of slices per series, we set max_slice = 200.<br>\nThe best public leaderboard score achieved was 0.81.</li>\n</ul>\n<h1>What Worked Well</h1>\n<ul>\n<li>We successfully implemented ROI extraction and ROI-level classification.</li>\n<li>The derivation of class-wise scores also worked as intended.</li>\n</ul>\n<h1>What Didn’t Work Well</h1>\n<ul>\n<li>The overall score decreased after series-level aggregation.<br>\nSeveral post-processing methods were explored to filter out false-positive ROIs, but none of them were effective.<br>\nWe should have considered classification methods that incorporate spatial context, such as 2.5D LSTM or 3D CNN.</li>\n<li>The performance on MRI T2 images remained consistently low, and even when using a dedicated model for this modality, no improvement was observed.</li>\n</ul>\n<blockquote>\n  <p>Thank You for Reading!</p>\n</blockquote>",
  "messages": [
    {
      "id": 3302592,
      "postDate": "2025-10-16T05:52:21.063Z",
      "content": "<h1>First of all</h1>\n<p>I would like to express my gratitude to the hosts for providing such a rewarding competition, to the Kaggle staff, and to all the participants who shared their insights in the discussions.</p>\n<h1>Strategy</h1>\n<p>From the early stage of the competition, based on the provided dataset, I devised a strategy to generate ROIs using a segmentation model and then classify them.</p>\n<h1>Pipeline overview</h1>\n<ol>\n<li>Stage1: Vessel Mask Prediction and ROI at the vessel-class level Extraction by a 2.5D Segmentation Model</li>\n<li>Stage2: ROI Classification</li>\n<li>Stage3: ROI Scoring and Series-level aggregation</li>\n</ol>\n<h1>Learning Process</h1>\n<h2>Stage1 (Segment Model)</h2>\n<ul>\n<li>We performed slice-level mask prediction using smp.Unet with an EfficientNet-B0 backbone.</li>\n<li>For preprocessing, windowing was applied to CT images and percentile normalization to MRI images. Basic data augmentation techniques were also used.</li>\n<li>The input size was (C, H, W) = (3, 512, 512).</li>\n<li>To address class imbalance among negative, positive, and vessel-specific positive samples, we used a random sampler with slice-level weighting.</li>\n<li>Using 3-channel inputs achieved better accuracy than 1-channel inputs.</li>\n<li>The final Dice score was approximately 0.66.</li>\n</ul>\n<h2>Stage1 (Vessel-class-level ROI Extraction）</h2>\n<ul>\n<li>We extracted ROIs using a segmentation model.<br>\nEach ROI was generated for every vessel class as the minimum bounding rectangle of the predicted mask with an added margin.<br>\nWhen multiple separated regions of the same class were detected within a slice, individual ROIs were created for each object (e.g., class_i_object_j).</li>\n<li>For aneurysm-positive slices, ROIs were extracted according to the above definition, and those containing the aneurysm coordinates were labeled as positive ROIs, achieving approximately 94% hit rate.</li>\n<li>Negative ROIs were sampled from both negative series and regions located more than 40 mm (in xyz space) away from aneurysm coordinates in positive series.</li>\n</ul>\n<h2>Stage2</h2>\n<ul>\n<li>For ROI classification, a 2.5D input was created by applying the same bounding box to the ROI slice and its adjacent slices.<br>\nThe image was resized while maintaining its aspect ratio, by scaling it so that the longer side matched the target size.<br>\nThe final input size was (C, H, W) = (3, 224, 224), and padding was applied as necessary.</li>\n<li>Multiple backbone networks were evaluated, including EfficientNet-B0 and EfficientNet-V2-S.<br>\nThe model achieved an ROI-level AUC of approximately 0.94.</li>\n</ul>\n<h2>Stage3</h2>\n<ul>\n<li>For each ROI, the aneurysm probability (prob) predicted by the ROI classification model was multiplied by the vessel class probabilities (class_scores) obtained from the segmentation model to compute the 13 class-wise scores.<br>\nThe binary aneurysm score was taken directly from the ROI probability (prob) without any modification.</li>\n<li>Series-level aggregation was performed using the top-k mean of ROI-level scores.</li>\n<li>The local validation AUC reached approximately 0.81.</li>\n</ul>\n<h1>Submission Process</h1>\n<ul>\n<li>The DICOM slices were aligned in the LPS coordinate system, and only the last 150 mm along the z-axis (corresponding to the head region) were used for inference.<br>\nTo limit the number of slices per series, we set max_slice = 200.<br>\nThe best public leaderboard score achieved was 0.81.</li>\n</ul>\n<h1>What Worked Well</h1>\n<ul>\n<li>We successfully implemented ROI extraction and ROI-level classification.</li>\n<li>The derivation of class-wise scores also worked as intended.</li>\n</ul>\n<h1>What Didn’t Work Well</h1>\n<ul>\n<li>The overall score decreased after series-level aggregation.<br>\nSeveral post-processing methods were explored to filter out false-positive ROIs, but none of them were effective.<br>\nWe should have considered classification methods that incorporate spatial context, such as 2.5D LSTM or 3D CNN.</li>\n<li>The performance on MRI T2 images remained consistently low, and even when using a dedicated model for this modality, no improvement was observed.</li>\n</ul>\n<blockquote>\n  <p>Thank You for Reading!</p>\n</blockquote>",
      "rawMarkdown": "# First of all\nI would like to express my gratitude to the hosts for providing such a rewarding competition, to the Kaggle staff, and to all the participants who shared their insights in the discussions.\n\n# Strategy\nFrom the early stage of the competition, based on the provided dataset, I devised a strategy to generate ROIs using a segmentation model and then classify them.\n\n# Pipeline overview\n1. Stage1: Vessel Mask Prediction and ROI at the vessel-class level Extraction by a 2.5D Segmentation Model\n2. Stage2: ROI Classification\n3. Stage3: ROI Scoring and Series-level aggregation\n\n# Learning Process\n##  Stage1 (Segment Model)\n- We performed slice-level mask prediction using smp.Unet with an EfficientNet-B0 backbone.\n- For preprocessing, windowing was applied to CT images and percentile normalization to MRI images. Basic data augmentation techniques were also used.\n- The input size was (C, H, W) = (3, 512, 512).\n- To address class imbalance among negative, positive, and vessel-specific positive samples, we used a random sampler with slice-level weighting.\n- Using 3-channel inputs achieved better accuracy than 1-channel inputs.\n- The final Dice score was approximately 0.66.\n\n## Stage1 (Vessel-class-level ROI Extraction）\n- We extracted ROIs using a segmentation model.\nEach ROI was generated for every vessel class as the minimum bounding rectangle of the predicted mask with an added margin.\nWhen multiple separated regions of the same class were detected within a slice, individual ROIs were created for each object (e.g., class_i_object_j).\n- For aneurysm-positive slices, ROIs were extracted according to the above definition, and those containing the aneurysm coordinates were labeled as positive ROIs, achieving approximately 94% hit rate.\n- Negative ROIs were sampled from both negative series and regions located more than 40 mm (in xyz space) away from aneurysm coordinates in positive series.\n\n## Stage2\n- For ROI classification, a 2.5D input was created by applying the same bounding box to the ROI slice and its adjacent slices.\nThe image was resized while maintaining its aspect ratio, by scaling it so that the longer side matched the target size.\nThe final input size was (C, H, W) = (3, 224, 224), and padding was applied as necessary.\n- Multiple backbone networks were evaluated, including EfficientNet-B0 and EfficientNet-V2-S.\nThe model achieved an ROI-level AUC of approximately 0.94.\n\n## Stage3\n- For each ROI, the aneurysm probability (prob) predicted by the ROI classification model was multiplied by the vessel class probabilities (class_scores) obtained from the segmentation model to compute the 13 class-wise scores.\nThe binary aneurysm score was taken directly from the ROI probability (prob) without any modification.\n- Series-level aggregation was performed using the top-k mean of ROI-level scores.\n- The local validation AUC reached approximately 0.81.\n\n# Submission Process\n- The DICOM slices were aligned in the LPS coordinate system, and only the last 150 mm along the z-axis (corresponding to the head region) were used for inference.\nTo limit the number of slices per series, we set max_slice = 200.\nThe best public leaderboard score achieved was 0.81.\n\n# What Worked Well\n- We successfully implemented ROI extraction and ROI-level classification.\n- The derivation of class-wise scores also worked as intended.\n# What Didn’t Work Well\n- The overall score decreased after series-level aggregation.\nSeveral post-processing methods were explored to filter out false-positive ROIs, but none of them were effective.\nWe should have considered classification methods that incorporate spatial context, such as 2.5D LSTM or 3D CNN.\n- The performance on MRI T2 images remained consistently low, and even when using a dedicated model for this modality, no improvement was observed.\n\n> Thank You for Reading!",
      "votes": 12
    },
    {
      "id": 3302652,
      "postDate": "2025-10-16T08:37:01.213Z",
      "content": "<p><a href=\"https://www.kaggle.com/koooeo\" target=\"_blank\">@koooeo</a> can I ask the best private score your team achieved? I think ROI classification is winning approach in this competition.</p>",
      "rawMarkdown": "@koooeo can I ask the best private score your team achieved? I think ROI classification is winning approach in this competition.",
      "votes": 2,
      "replies": [
        {
          "id": 3302677,
          "postDate": "2025-10-16T09:14:15.207Z",
          "content": "<p>Thanks! Our best private score was around 0.80. I also think ROI classification turned out to be a key approach in this competition.</p>",
          "rawMarkdown": "Thanks! Our best private score was around 0.80. I also think ROI classification turned out to be a key approach in this competition.",
          "replies": [
            {
              "id": 3302690,
              "postDate": "2025-10-16T09:32:02.257Z",
              "content": "<p>I also tried filtering out false positives, but this approach might remove valid high-probability predictions that help improve the AUROC ranking.</p>",
              "rawMarkdown": "I also tried filtering out false positives, but this approach might remove valid high-probability predictions that help improve the AUROC ranking.",
              "votes": 1
            },
            {
              "id": 3302700,
              "postDate": "2025-10-16T09:51:31.747Z",
              "content": "<p>Exactly. We struggled a lot with false positives.<br>\nMaybe adding location information of common aneurysm sites or stacking 3-directional MIPs as additional channels could have helped.</p>",
              "rawMarkdown": "Exactly. We struggled a lot with false positives.\nMaybe adding location information of common aneurysm sites or stacking 3-directional MIPs as additional channels could have helped.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3426061,
      "postDate": "2026-03-21T20:05:25.220Z",
      "content": "<p><a href=\"https://www.kaggle.com/koooeo\" target=\"_blank\">@koooeo</a> Congratulations on the win! Can you share the related code for training? I want to study about the details of your code designs</p>",
      "rawMarkdown": "@koooeo Congratulations on the win! Can you share the related code for training? I want to study about the details of your code designs"
    },
    {
      "id": 3303917,
      "postDate": "2025-10-19T09:19:12.323Z",
      "content": "<p>Congratulations on the win!<br>\nCan you explain how did you do Vessel Mask Prediction and ROI at the vessel-class level Extraction? What specific model and dataset did you use to train it? Is it a pretrained model? </p>",
      "rawMarkdown": "Congratulations on the win!\nCan you explain how did you do Vessel Mask Prediction and ROI at the vessel-class level Extraction? What specific model and dataset did you use to train it? Is it a pretrained model? \n",
      "replies": [
        {
          "id": 3303944,
          "postDate": "2025-10-19T10:45:46.530Z",
          "content": "<p>Thanks for your question!</p>\n<p>Vascular mask prediction is performed on a slice-by-slice basis. During the inference process, predictions are made for all slices within the specified volume.</p>\n<p>A minimum bounding box (BBOX) is calculated for each predicted vascular class mask, and the same BBOX is used to create a 3ch image from the slices before and after. This forms one ROI unit.</p>\n<p>The model used was smp.Unet, with a backbone of timm-efficientnet-b0.</p>\n<p>Since slice-by-slice training was required, a dataset was created by expanding the provided nii file to fit each slice. To improve training speed, the nii file was saved in npz format for each slice. Enabling pre-training weights speeds convergence, but this did not significantly affect the final results.</p>\n<p>Note that the nii file is rotated 90° clockwise relative to the corresponding DICOM file, so it must be rotated 90° counterclockwise. Additionally, for some MRI data, the images were flipped left and right, so it was necessary to rotate them 90° counterclockwise before flipping them left and right.</p>",
          "rawMarkdown": "Thanks for your question!\n\nVascular mask prediction is performed on a slice-by-slice basis. During the inference process, predictions are made for all slices within the specified volume.\n\nA minimum bounding box (BBOX) is calculated for each predicted vascular class mask, and the same BBOX is used to create a 3ch image from the slices before and after. This forms one ROI unit.\n\nThe model used was smp.Unet, with a backbone of timm-efficientnet-b0.\n\nSince slice-by-slice training was required, a dataset was created by expanding the provided nii file to fit each slice. To improve training speed, the nii file was saved in npz format for each slice. Enabling pre-training weights speeds convergence, but this did not significantly affect the final results.\n\nNote that the nii file is rotated 90° clockwise relative to the corresponding DICOM file, so it must be rotated 90° counterclockwise. Additionally, for some MRI data, the images were flipped left and right, so it was necessary to rotate them 90° counterclockwise before flipping them left and right."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3302652,
      "author_name": "Tom",
      "author_url": "",
      "post_date": "2025-10-16T08:37:01.213000",
      "content": "<p><a href=\"https://www.kaggle.com/koooeo\" target=\"_blank\">@koooeo</a> can I ask the best private score your team achieved? I think ROI classification is winning approach in this competition.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3302677,
          "author_name": "Shun Kuraishi",
          "author_url": "",
          "post_date": "2025-10-16T09:14:15.207000",
          "content": "<p>Thanks! Our best private score was around 0.80. I also think ROI classification turned out to be a key approach in this competition.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3302690,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2025-10-16T09:32:02.257000",
              "content": "<p>I also tried filtering out false positives, but this approach might remove valid high-probability predictions that help improve the AUROC ranking.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3302700,
              "author_name": "Shun Kuraishi",
              "author_url": "",
              "post_date": "2025-10-16T09:51:31.747000",
              "content": "<p>Exactly. We struggled a lot with false positives.<br>\nMaybe adding location information of common aneurysm sites or stacking 3-directional MIPs as additional channels could have helped.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3426061,
      "author_name": "Brench",
      "author_url": "",
      "post_date": "2026-03-21T20:05:25.220000",
      "content": "<p><a href=\"https://www.kaggle.com/koooeo\" target=\"_blank\">@koooeo</a> Congratulations on the win! Can you share the related code for training? I want to study about the details of your code designs</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3303917,
      "author_name": "Jeslin Thomaskutty",
      "author_url": "",
      "post_date": "2025-10-19T09:19:12.323000",
      "content": "<p>Congratulations on the win!<br>\nCan you explain how did you do Vessel Mask Prediction and ROI at the vessel-class level Extraction? What specific model and dataset did you use to train it? Is it a pretrained model? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3303944,
          "author_name": "koooeo",
          "author_url": "",
          "post_date": "2025-10-19T10:45:46.530000",
          "content": "<p>Thanks for your question!</p>\n<p>Vascular mask prediction is performed on a slice-by-slice basis. During the inference process, predictions are made for all slices within the specified volume.</p>\n<p>A minimum bounding box (BBOX) is calculated for each predicted vascular class mask, and the same BBOX is used to create a 3ch image from the slices before and after. This forms one ROI unit.</p>\n<p>The model used was smp.Unet, with a backbone of timm-efficientnet-b0.</p>\n<p>Since slice-by-slice training was required, a dataset was created by expanding the provided nii file to fit each slice. To improve training speed, the nii file was saved in npz format for each slice. Enabling pre-training weights speeds convergence, but this did not significantly affect the final results.</p>\n<p>Note that the nii file is rotated 90° clockwise relative to the corresponding DICOM file, so it must be rotated 90° counterclockwise. Additionally, for some MRI data, the images were flipped left and right, so it was necessary to rotate them 90° counterclockwise before flipping them left and right.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3302592": "# First of all\nI would like to express my gratitude to the hosts for providing such a rewarding competition, to the Kaggle staff, and to all the participants who shared their insights in the discussions.\n\n# Strategy\nFrom the early stage of the competition, based on the provided dataset, I devised a strategy to generate ROIs using a segmentation model and then classify them.\n\n# Pipeline overview\n1. Stage1: Vessel Mask Prediction and ROI at the vessel-class level Extraction by a 2.5D Segmentation Model\n2. Stage2: ROI Classification\n3. Stage3: ROI Scoring and Series-level aggregation\n\n# Learning Process\n##  Stage1 (Segment Model)\n- We performed slice-level mask prediction using smp.Unet with an EfficientNet-B0 backbone.\n- For preprocessing, windowing was applied to CT images and percentile normalization to MRI images. Basic data augmentation techniques were also used.\n- The input size was (C, H, W) = (3, 512, 512).\n- To address class imbalance among negative, positive, and vessel-specific positive samples, we used a random sampler with slice-level weighting.\n- Using 3-channel inputs achieved better accuracy than 1-channel inputs.\n- The final Dice score was approximately 0.66.\n\n## Stage1 (Vessel-class-level ROI Extraction）\n- We extracted ROIs using a segmentation model.\nEach ROI was generated for every vessel class as the minimum bounding rectangle of the predicted mask with an added margin.\nWhen multiple separated regions of the same class were detected within a slice, individual ROIs were created for each object (e.g., class_i_object_j).\n- For aneurysm-positive slices, ROIs were extracted according to the above definition, and those containing the aneurysm coordinates were labeled as positive ROIs, achieving approximately 94% hit rate.\n- Negative ROIs were sampled from both negative series and regions located more than 40 mm (in xyz space) away from aneurysm coordinates in positive series.\n\n## Stage2\n- For ROI classification, a 2.5D input was created by applying the same bounding box to the ROI slice and its adjacent slices.\nThe image was resized while maintaining its aspect ratio, by scaling it so that the longer side matched the target size.\nThe final input size was (C, H, W) = (3, 224, 224), and padding was applied as necessary.\n- Multiple backbone networks were evaluated, including EfficientNet-B0 and EfficientNet-V2-S.\nThe model achieved an ROI-level AUC of approximately 0.94.\n\n## Stage3\n- For each ROI, the aneurysm probability (prob) predicted by the ROI classification model was multiplied by the vessel class probabilities (class_scores) obtained from the segmentation model to compute the 13 class-wise scores.\nThe binary aneurysm score was taken directly from the ROI probability (prob) without any modification.\n- Series-level aggregation was performed using the top-k mean of ROI-level scores.\n- The local validation AUC reached approximately 0.81.\n\n# Submission Process\n- The DICOM slices were aligned in the LPS coordinate system, and only the last 150 mm along the z-axis (corresponding to the head region) were used for inference.\nTo limit the number of slices per series, we set max_slice = 200.\nThe best public leaderboard score achieved was 0.81.\n\n# What Worked Well\n- We successfully implemented ROI extraction and ROI-level classification.\n- The derivation of class-wise scores also worked as intended.\n# What Didn’t Work Well\n- The overall score decreased after series-level aggregation.\nSeveral post-processing methods were explored to filter out false-positive ROIs, but none of them were effective.\nWe should have considered classification methods that incorporate spatial context, such as 2.5D LSTM or 3D CNN.\n- The performance on MRI T2 images remained consistently low, and even when using a dedicated model for this modality, no improvement was observed.\n\n> Thank You for Reading!",
    "3302652": "@koooeo can I ask the best private score your team achieved? I think ROI classification is winning approach in this competition.",
    "3426061": "@koooeo Congratulations on the win! Can you share the related code for training? I want to study about the details of your code designs",
    "3303917": "Congratulations on the win!\nCan you explain how did you do Vessel Mask Prediction and ROI at the vessel-class level Extraction? What specific model and dataset did you use to train it? Is it a pretrained model? \n"
  }
}