{
  "id": 612186,
  "title": "11th Place Solution",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/612186",
  "author_name": "RihanPiggy",
  "post_date": "2025-10-17T11:38:28.324000",
  "votes": 23,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Congratulations to all the winners! Thanks to the organizers for hosting such an interesting competition. It is really a hard one and I enjoyed it a lot. Here let me share my solution.</p>\n<h1>TL;DR</h1>\n<p>My solution consists of 4 stages:</p>\n<ol>\n<li>Generate brain mask with TotalSegmentator, train a 3D Unet to crop the ROI region which contains CoW Area.</li>\n<li>Train 5 fold 2.5D CoW Seg model (cls loss + seg loss) and extract features.</li>\n<li>Train 5 fold 2.5D Aneurysm model (cls + keypoint loss) and extract features.</li>\n<li>Train RNNs on concatenated features from CoW Seg model and Aneurysm model across depth direction.</li>\n</ol>\n<h1>Preprocessing</h1>\n<p>I used <a href=\"https://github.com/icometrix/dicom2nifti/tree/main\" target=\"_blank\">dicom2nifti</a> to convert dicom files to nifti files because TotalSegmentator uses this library.<br>\nI wrote a custom version to transform the coordinates and frame indexes from train_localizers.csv to the new ones which align with the nifti files.</p>\n<h1>Stage 1: Brain Mask</h1>\n<h2>Brain Mask Generation</h2>\n<p>I used TotalSegmentator to generate brain masks for all scans.<br>\nStrangely, some of the CT scans have better brain masks generated by TotalSegmentator MRI than TotalSegmentator CT. Thus I generated 2 brain masks for each scan and selected the one with more pixels.<br>\nI also filtered out small masks with less than 1 million pixels. Finally, I have 4386 brain masks.</p>\n<h2>3D Unet for Brain Segmentation</h2>\n<p>I trained a 3D Unet to segment the brain. The train code is reimplementation from <a href=\"https://www.kaggle.com/code/haqishen/rsna-2022-1st-place-solution-train-stage1\" target=\"_blank\">Qishen Ha's RSNA2022 solution</a>.<br>\nI first pretrained the model with <a href=\"https://zenodo.org/records/10047292\" target=\"_blank\">CT</a> and <a href=\"https://zenodo.org/records/14710732\" target=\"_blank\">MRI</a> scans from Zenodo, which is a subset of the training data used to train TotalSegmentator. Then, I trained the model with 4386 brain masks. The model is trained with 5 folds. The input size is (128, 128, 128).<br>\nI used the trained 3D Unet to generate brain masks and crop the ROI region for all scans. The parameters for cropping are adjusted according to the CoWSeg and train_localizers.csv to ensure the cropped region contains the CoW Area.</p>\n<h1>Stage 2: CoW Segmentation Model</h1>\n<p>Unlike other competitors, I did not managed to train a good 3D CoW Segmentation model. Instead, I trained a 2.5D CoW Segmentation model with classification head. The purpose of this model is to extract positional features to feed into the final RNN model. According to my experience in RSNA2022, positional features boosted the CV. So although segmentation loss is even worse with 2.5D model (dice=0.56), the extracted features are still useful.<br>\nThe training code is reimplementation from <a href=\"https://www.kaggle.com/code/awsaf49/uwmgi-2-5d-train-pytorch\" target=\"_blank\">this amazing notebook</a> from UWMGI. The backbone I used is tf_efficientnetv2_b0.in1k (img size=224).<br>\nAlso, using the OOF classification logits from the CoW Seg model like in stage 1, I filter out slices which do not contain CoW area.</p>\n<h1>Stage 3: Aneurysm Model</h1>\n<p>Similar to CoW Seg model, I trained a 2.5D Aneurysm model with classification head and keypoint head to extract features.<br>\nThe key for the successful training of this model is as follows:</p>\n<ul>\n<li>3 rounds of negative sampling. In the first round, I used classification logits from CoW Seg model to sample slices which contains vessels. In the next 2 rounds, I used OOF classification logits from the Aneurysm model itself to sample hard negative slices.</li>\n<li>Heuristic cropping to make the model focus on the vessel area. Crop ratio is adjusted according to the coordinates. (ensuring all the coordinates are in the cropped region)</li>\n<li>Focal loss with alpha=0.75</li>\n<li>Self distillation with OOF logits to deal with label noise.</li>\n<li>EMA</li>\n</ul>\n<p>The backbones I used in this stage are:</p>\n<ul>\n<li>coatnet_rmlp_2_rw_384: 5 fold, img_size=384</li>\n<li>coatnet_rmlp_2_rw_384: all data train. img_size=384</li>\n<li>tf_efficientnetv2_s.in21k_ft_in1k: another 5 fold, img_size=384</li>\n</ul>\n<p>All data train used a different revision of dataset with less hard negative images. I first trained a 5 fold model and recorded the epoch with best CV to calculate the total training steps for all data train.</p>\n<h1>Stage 4: RNN Model</h1>\n<p>I trained 4 types of RNN models for each aneurysm model. The model starts to overfit quickly, EMA and heavy mixup(mixup ratio=0.5) are used to regularize the training.</p>\n<p>For 5 fold Coatnet:</p>\n<ul>\n<li>Residual LSTM: aneurysm feature + CoW seg feature</li>\n<li>Residual GRU: aneurysm feature only</li>\n<li>BiLSTM: aneurysm feature only</li>\n<li>Bert: aneurysm feature only</li>\n</ul>\n<p>For 5 fold Efficientnet:</p>\n<ul>\n<li>Residual LSTM: aneurysm feature + CoW seg feature</li>\n<li>Residual GRU: aneurysm feature + CoW seg feature</li>\n<li>LSTM: aneurysm feature only</li>\n<li>Bert: aneurysm feature only</li>\n</ul>\n<p>For all data train Coatnet:</p>\n<ul>\n<li>Residual LSTM: aneurysm feature + CoW seg feature</li>\n<li>Residual GRU: aneurysm feature + CoW seg feature</li>\n<li>BiLSTM: aneurysm feature + CoW seg feature</li>\n<li>Bert: aneurysm feature + CoW seg feature</li>\n</ul>\n<p>The ensemble of these RNNs reached auc=0.8909 and by self-distillation with OOF logits, another submission reached auc=0.9035</p>\n<ul>\n<li>submission 1: CV: 0.8909, Public: 0.87689, Private: 0.82201</li>\n<li>submission 2: CV: 0.9035, Public: 0.86703, Private: 0.81321<br>\nI am still very confused about the huge gap between Public and Private scores.</li>\n</ul>\n<h1>Maximizing GPU Utilization</h1>\n<p>To maximize the inference speed, I used Monai to preprocess the images on GPU. I used T4 x 2 and split the patient volume into 2 parts to parallelize the inference of 2.5D models. For RNN models, I simply split the models into 2 groups and run them in parallel on 2 GPUs. That's how I managed to finish the inference of 11 aneurysm models and 45 RNN models within 12 hours. <br>\nI suffered from timeout issue in the final days. Hoping that Kaggle will fix the new evaluation framework.</p>",
  "messages": [
    {
      "id": 3303213,
      "postDate": "2025-10-17T11:38:28.323Z",
      "content": "<p>Congratulations to all the winners! Thanks to the organizers for hosting such an interesting competition. It is really a hard one and I enjoyed it a lot. Here let me share my solution.</p>\n<h1>TL;DR</h1>\n<p>My solution consists of 4 stages:</p>\n<ol>\n<li>Generate brain mask with TotalSegmentator, train a 3D Unet to crop the ROI region which contains CoW Area.</li>\n<li>Train 5 fold 2.5D CoW Seg model (cls loss + seg loss) and extract features.</li>\n<li>Train 5 fold 2.5D Aneurysm model (cls + keypoint loss) and extract features.</li>\n<li>Train RNNs on concatenated features from CoW Seg model and Aneurysm model across depth direction.</li>\n</ol>\n<h1>Preprocessing</h1>\n<p>I used <a href=\"https://github.com/icometrix/dicom2nifti/tree/main\" target=\"_blank\">dicom2nifti</a> to convert dicom files to nifti files because TotalSegmentator uses this library.<br>\nI wrote a custom version to transform the coordinates and frame indexes from train_localizers.csv to the new ones which align with the nifti files.</p>\n<h1>Stage 1: Brain Mask</h1>\n<h2>Brain Mask Generation</h2>\n<p>I used TotalSegmentator to generate brain masks for all scans.<br>\nStrangely, some of the CT scans have better brain masks generated by TotalSegmentator MRI than TotalSegmentator CT. Thus I generated 2 brain masks for each scan and selected the one with more pixels.<br>\nI also filtered out small masks with less than 1 million pixels. Finally, I have 4386 brain masks.</p>\n<h2>3D Unet for Brain Segmentation</h2>\n<p>I trained a 3D Unet to segment the brain. The train code is reimplementation from <a href=\"https://www.kaggle.com/code/haqishen/rsna-2022-1st-place-solution-train-stage1\" target=\"_blank\">Qishen Ha's RSNA2022 solution</a>.<br>\nI first pretrained the model with <a href=\"https://zenodo.org/records/10047292\" target=\"_blank\">CT</a> and <a href=\"https://zenodo.org/records/14710732\" target=\"_blank\">MRI</a> scans from Zenodo, which is a subset of the training data used to train TotalSegmentator. Then, I trained the model with 4386 brain masks. The model is trained with 5 folds. The input size is (128, 128, 128).<br>\nI used the trained 3D Unet to generate brain masks and crop the ROI region for all scans. The parameters for cropping are adjusted according to the CoWSeg and train_localizers.csv to ensure the cropped region contains the CoW Area.</p>\n<h1>Stage 2: CoW Segmentation Model</h1>\n<p>Unlike other competitors, I did not managed to train a good 3D CoW Segmentation model. Instead, I trained a 2.5D CoW Segmentation model with classification head. The purpose of this model is to extract positional features to feed into the final RNN model. According to my experience in RSNA2022, positional features boosted the CV. So although segmentation loss is even worse with 2.5D model (dice=0.56), the extracted features are still useful.<br>\nThe training code is reimplementation from <a href=\"https://www.kaggle.com/code/awsaf49/uwmgi-2-5d-train-pytorch\" target=\"_blank\">this amazing notebook</a> from UWMGI. The backbone I used is tf_efficientnetv2_b0.in1k (img size=224).<br>\nAlso, using the OOF classification logits from the CoW Seg model like in stage 1, I filter out slices which do not contain CoW area.</p>\n<h1>Stage 3: Aneurysm Model</h1>\n<p>Similar to CoW Seg model, I trained a 2.5D Aneurysm model with classification head and keypoint head to extract features.<br>\nThe key for the successful training of this model is as follows:</p>\n<ul>\n<li>3 rounds of negative sampling. In the first round, I used classification logits from CoW Seg model to sample slices which contains vessels. In the next 2 rounds, I used OOF classification logits from the Aneurysm model itself to sample hard negative slices.</li>\n<li>Heuristic cropping to make the model focus on the vessel area. Crop ratio is adjusted according to the coordinates. (ensuring all the coordinates are in the cropped region)</li>\n<li>Focal loss with alpha=0.75</li>\n<li>Self distillation with OOF logits to deal with label noise.</li>\n<li>EMA</li>\n</ul>\n<p>The backbones I used in this stage are:</p>\n<ul>\n<li>coatnet_rmlp_2_rw_384: 5 fold, img_size=384</li>\n<li>coatnet_rmlp_2_rw_384: all data train. img_size=384</li>\n<li>tf_efficientnetv2_s.in21k_ft_in1k: another 5 fold, img_size=384</li>\n</ul>\n<p>All data train used a different revision of dataset with less hard negative images. I first trained a 5 fold model and recorded the epoch with best CV to calculate the total training steps for all data train.</p>\n<h1>Stage 4: RNN Model</h1>\n<p>I trained 4 types of RNN models for each aneurysm model. The model starts to overfit quickly, EMA and heavy mixup(mixup ratio=0.5) are used to regularize the training.</p>\n<p>For 5 fold Coatnet:</p>\n<ul>\n<li>Residual LSTM: aneurysm feature + CoW seg feature</li>\n<li>Residual GRU: aneurysm feature only</li>\n<li>BiLSTM: aneurysm feature only</li>\n<li>Bert: aneurysm feature only</li>\n</ul>\n<p>For 5 fold Efficientnet:</p>\n<ul>\n<li>Residual LSTM: aneurysm feature + CoW seg feature</li>\n<li>Residual GRU: aneurysm feature + CoW seg feature</li>\n<li>LSTM: aneurysm feature only</li>\n<li>Bert: aneurysm feature only</li>\n</ul>\n<p>For all data train Coatnet:</p>\n<ul>\n<li>Residual LSTM: aneurysm feature + CoW seg feature</li>\n<li>Residual GRU: aneurysm feature + CoW seg feature</li>\n<li>BiLSTM: aneurysm feature + CoW seg feature</li>\n<li>Bert: aneurysm feature + CoW seg feature</li>\n</ul>\n<p>The ensemble of these RNNs reached auc=0.8909 and by self-distillation with OOF logits, another submission reached auc=0.9035</p>\n<ul>\n<li>submission 1: CV: 0.8909, Public: 0.87689, Private: 0.82201</li>\n<li>submission 2: CV: 0.9035, Public: 0.86703, Private: 0.81321<br>\nI am still very confused about the huge gap between Public and Private scores.</li>\n</ul>\n<h1>Maximizing GPU Utilization</h1>\n<p>To maximize the inference speed, I used Monai to preprocess the images on GPU. I used T4 x 2 and split the patient volume into 2 parts to parallelize the inference of 2.5D models. For RNN models, I simply split the models into 2 groups and run them in parallel on 2 GPUs. That's how I managed to finish the inference of 11 aneurysm models and 45 RNN models within 12 hours. <br>\nI suffered from timeout issue in the final days. Hoping that Kaggle will fix the new evaluation framework.</p>",
      "rawMarkdown": "Congratulations to all the winners! Thanks to the organizers for hosting such an interesting competition. It is really a hard one and I enjoyed it a lot. Here let me share my solution.\n\n# TL;DR\nMy solution consists of 4 stages:\n1. Generate brain mask with TotalSegmentator, train a 3D Unet to crop the ROI region which contains CoW Area.\n2. Train 5 fold 2.5D CoW Seg model (cls loss + seg loss) and extract features.\n3. Train 5 fold 2.5D Aneurysm model (cls + keypoint loss) and extract features.\n4. Train RNNs on concatenated features from CoW Seg model and Aneurysm model across depth direction.\n\n# Preprocessing\nI used [dicom2nifti](https://github.com/icometrix/dicom2nifti/tree/main) to convert dicom files to nifti files because TotalSegmentator uses this library.\nI wrote a custom version to transform the coordinates and frame indexes from train_localizers.csv to the new ones which align with the nifti files.\n\n# Stage 1: Brain Mask\n\n## Brain Mask Generation\nI used TotalSegmentator to generate brain masks for all scans.\nStrangely, some of the CT scans have better brain masks generated by TotalSegmentator MRI than TotalSegmentator CT. Thus I generated 2 brain masks for each scan and selected the one with more pixels.\nI also filtered out small masks with less than 1 million pixels. Finally, I have 4386 brain masks.\n\n## 3D Unet for Brain Segmentation\nI trained a 3D Unet to segment the brain. The train code is reimplementation from [Qishen Ha's RSNA2022 solution](https://www.kaggle.com/code/haqishen/rsna-2022-1st-place-solution-train-stage1).\nI first pretrained the model with [CT](https://zenodo.org/records/10047292) and [MRI](https://zenodo.org/records/14710732) scans from Zenodo, which is a subset of the training data used to train TotalSegmentator. Then, I trained the model with 4386 brain masks. The model is trained with 5 folds. The input size is (128, 128, 128).\nI used the trained 3D Unet to generate brain masks and crop the ROI region for all scans. The parameters for cropping are adjusted according to the CoWSeg and train_localizers.csv to ensure the cropped region contains the CoW Area.\n\n# Stage 2: CoW Segmentation Model\nUnlike other competitors, I did not managed to train a good 3D CoW Segmentation model. Instead, I trained a 2.5D CoW Segmentation model with classification head. The purpose of this model is to extract positional features to feed into the final RNN model. According to my experience in RSNA2022, positional features boosted the CV. So although segmentation loss is even worse with 2.5D model (dice=0.56), the extracted features are still useful.\nThe training code is reimplementation from [this amazing notebook](https://www.kaggle.com/code/awsaf49/uwmgi-2-5d-train-pytorch) from UWMGI. The backbone I used is tf_efficientnetv2_b0.in1k (img size=224).\nAlso, using the OOF classification logits from the CoW Seg model like in stage 1, I filter out slices which do not contain CoW area.\n\n# Stage 3: Aneurysm Model\nSimilar to CoW Seg model, I trained a 2.5D Aneurysm model with classification head and keypoint head to extract features.\nThe key for the successful training of this model is as follows:\n- 3 rounds of negative sampling. In the first round, I used classification logits from CoW Seg model to sample slices which contains vessels. In the next 2 rounds, I used OOF classification logits from the Aneurysm model itself to sample hard negative slices.\n- Heuristic cropping to make the model focus on the vessel area. Crop ratio is adjusted according to the coordinates. (ensuring all the coordinates are in the cropped region)\n- Focal loss with alpha=0.75\n- Self distillation with OOF logits to deal with label noise.\n- EMA\n\nThe backbones I used in this stage are:\n- coatnet_rmlp_2_rw_384: 5 fold, img_size=384\n- coatnet_rmlp_2_rw_384: all data train. img_size=384\n- tf_efficientnetv2_s.in21k_ft_in1k: another 5 fold, img_size=384\n\nAll data train used a different revision of dataset with less hard negative images. I first trained a 5 fold model and recorded the epoch with best CV to calculate the total training steps for all data train.\n\n# Stage 4: RNN Model\nI trained 4 types of RNN models for each aneurysm model. The model starts to overfit quickly, EMA and heavy mixup(mixup ratio=0.5) are used to regularize the training.\n\nFor 5 fold Coatnet:\n- Residual LSTM: aneurysm feature + CoW seg feature\n- Residual GRU: aneurysm feature only\n- BiLSTM: aneurysm feature only\n- Bert: aneurysm feature only\n\nFor 5 fold Efficientnet:\n- Residual LSTM: aneurysm feature + CoW seg feature\n- Residual GRU: aneurysm feature + CoW seg feature\n- LSTM: aneurysm feature only\n- Bert: aneurysm feature only\n\nFor all data train Coatnet:\n- Residual LSTM: aneurysm feature + CoW seg feature\n- Residual GRU: aneurysm feature + CoW seg feature\n- BiLSTM: aneurysm feature + CoW seg feature\n- Bert: aneurysm feature + CoW seg feature\n\nThe ensemble of these RNNs reached auc=0.8909 and by self-distillation with OOF logits, another submission reached auc=0.9035\n- submission 1: CV: 0.8909, Public: 0.87689, Private: 0.82201\n- submission 2: CV: 0.9035, Public: 0.86703, Private: 0.81321\nI am still very confused about the huge gap between Public and Private scores.\n\n# Maximizing GPU Utilization\nTo maximize the inference speed, I used Monai to preprocess the images on GPU. I used T4 x 2 and split the patient volume into 2 parts to parallelize the inference of 2.5D models. For RNN models, I simply split the models into 2 groups and run them in parallel on 2 GPUs. That's how I managed to finish the inference of 11 aneurysm models and 45 RNN models within 12 hours. \nI suffered from timeout issue in the final days. Hoping that Kaggle will fix the new evaluation framework.",
      "votes": 23
    },
    {
      "id": 3303479,
      "postDate": "2025-10-18T06:19:43.110Z",
      "content": "<p>Insightful solution! I observed a similar discrepancy between public and private LB. After implementing strict exception handling by removing a fallback to 0.5 predictions, I noticed a consistent 3-4% gap. I suspect this is due to a difference in the distribution of abnormal data between the public and private test sets.</p>",
      "rawMarkdown": "Insightful solution! I observed a similar discrepancy between public and private LB. After implementing strict exception handling by removing a fallback to 0.5 predictions, I noticed a consistent 3-4% gap. I suspect this is due to a difference in the distribution of abnormal data between the public and private test sets.",
      "votes": 2,
      "replies": [
        {
          "id": 3304483,
          "postDate": "2025-10-20T16:02:51.677Z",
          "content": "<p>Oh, glad to know that I am not alone.</p>\n<blockquote>\n  <p>After implementing strict exception handling by removing a fallback to 0.5 predictions, I noticed a consistent 3-4% gap</p>\n</blockquote>\n<p>Do you mean that the gap occured after applying strict exception handling!? Meaning that without strict handling, we can get a better private LB?<br>\nWould you mind to share what kind of exception occurs in your inference pipeline? For me, because I used dicom2nifti library which refers dicom tags, so for multiframe dicom, I catch the exception and set the default value for tag like Orientation and Position.</p>",
          "rawMarkdown": "Oh, glad to know that I am not alone.\n>  After implementing strict exception handling by removing a fallback to 0.5 predictions, I noticed a consistent 3-4% gap\n\nDo you mean that the gap occured after applying strict exception handling!? Meaning that without strict handling, we can get a better private LB?\nWould you mind to share what kind of exception occurs in your inference pipeline? For me, because I used dicom2nifti library which refers dicom tags, so for multiframe dicom, I catch the exception and set the default value for tag like Orientation and Position.",
          "votes": 1,
          "replies": [
            {
              "id": 3304495,
              "postDate": "2025-10-20T16:23:11.227Z",
              "content": "<p>I also used the dicom2nifti library. Regarding abnormal data, primarily multiframe data—specifically, the discrepancy appeared after I implemented processing for multiframe data using pydicom and mapping spacing based on shape:</p>\n<pre><code> (dicom_files) == :\n    ds = pydicom.dcmread(dicom_files[], force=)\n\n    pixel_array = ds.pixel_array.astype(np.int16)\n\n    input_img_np = pixel_array[]\n    original_spacing = get_spacing_by_shape(pixel_array.shape)[::-]\n\n    probs = predict_aneurysm(input_img_np, original_spacing, DEVICE)\n</code></pre>\n<p>After implementing this, the gap consistently remained above 2%. After adding orientation correction for T2 data, the gap increased to over 3%. These are what I refer to as exception handling—there's no need to use try-except and return 0.5. </p>\n<p>I suspect the main reason is that the proportion of multiframe data in the private test set is lower than in the public test set. For detailed inference code, please see: <a href=\"https://www.kaggle.com/code/pengchengshi/bravecowcow-2nd-place-inference\" target=\"_blank\">https://www.kaggle.com/code/pengchengshi/bravecowcow-2nd-place-inference</a></p>",
              "rawMarkdown": "I also used the dicom2nifti library. Regarding abnormal data, primarily multiframe data—specifically, the discrepancy appeared after I implemented processing for multiframe data using pydicom and mapping spacing based on shape:\n\n```python\nelif len(dicom_files) == 1:\n    ds = pydicom.dcmread(dicom_files[0], force=True)\n\n    pixel_array = ds.pixel_array.astype(np.int16)\n\n    input_img_np = pixel_array[None]\n    original_spacing = get_spacing_by_shape(pixel_array.shape)[::-1]\n\n    probs = predict_aneurysm(input_img_np, original_spacing, DEVICE)\n```\n\nAfter implementing this, the gap consistently remained above 2%. After adding orientation correction for T2 data, the gap increased to over 3%. These are what I refer to as exception handling—there's no need to use try-except and return 0.5. \n\nI suspect the main reason is that the proportion of multiframe data in the private test set is lower than in the public test set. For detailed inference code, please see: https://www.kaggle.com/code/pengchengshi/bravecowcow-2nd-place-inference",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3303225,
      "postDate": "2025-10-17T12:42:27.913Z",
      "content": "<p><a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> that's a good solution, especially utilizing 2.5d segmentation, I would like to use it in my next rsna competition. But seems one of your stage causes overfit, I think improving its generalization can make this solution to 1st.</p>",
      "rawMarkdown": "@honglihang that's a good solution, especially utilizing 2.5d segmentation, I would like to use it in my next rsna competition. But seems one of your stage causes overfit, I think improving its generalization can make this solution to 1st.",
      "replies": [
        {
          "id": 3303232,
          "postDate": "2025-10-17T13:25:15.393Z",
          "content": "<p>Thank you. I kept the kfold same in all the stages, so I believe there is no leak in training process, and public LB is quite corelated to CV, which is very confusing for me. <br>\nConsidering that test scans are collected from different sites, maybe there is a domain shift in test data which cause the brain segmentator not working well? Another possibility is that there is an issue in dicom tag because I used tags to perform orienting and scaling.</p>",
          "rawMarkdown": "Thank you. I kept the kfold same in all the stages, so I believe there is no leak in training process, and public LB is quite corelated to CV, which is very confusing for me. \nConsidering that test scans are collected from different sites, maybe there is a domain shift in test data which cause the brain segmentator not working well? Another possibility is that there is an issue in dicom tag because I used tags to perform orienting and scaling.",
          "votes": 3,
          "replies": [
            {
              "id": 3303246,
              "postDate": "2025-10-17T13:56:02.717Z",
              "content": "<p>Likely due to domain shift: the segmentor failed to generate valid masks for cases absent from the training set and the public hidden test set but present in the private hidden test set, causing the pipeline to fail.</p>",
              "rawMarkdown": "Likely due to domain shift: the segmentor failed to generate valid masks for cases absent from the training set and the public hidden test set but present in the private hidden test set, causing the pipeline to fail."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3303479,
      "author_name": "Pengcheng Shi",
      "author_url": "",
      "post_date": "2025-10-18T06:19:43.110000",
      "content": "<p>Insightful solution! I observed a similar discrepancy between public and private LB. After implementing strict exception handling by removing a fallback to 0.5 predictions, I noticed a consistent 3-4% gap. I suspect this is due to a difference in the distribution of abnormal data between the public and private test sets.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3304483,
          "author_name": "RihanPiggy",
          "author_url": "",
          "post_date": "2025-10-20T16:02:51.677000",
          "content": "<p>Oh, glad to know that I am not alone.</p>\n<blockquote>\n  <p>After implementing strict exception handling by removing a fallback to 0.5 predictions, I noticed a consistent 3-4% gap</p>\n</blockquote>\n<p>Do you mean that the gap occured after applying strict exception handling!? Meaning that without strict handling, we can get a better private LB?<br>\nWould you mind to share what kind of exception occurs in your inference pipeline? For me, because I used dicom2nifti library which refers dicom tags, so for multiframe dicom, I catch the exception and set the default value for tag like Orientation and Position.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3304495,
              "author_name": "Pengcheng Shi",
              "author_url": "",
              "post_date": "2025-10-20T16:23:11.227000",
              "content": "<p>I also used the dicom2nifti library. Regarding abnormal data, primarily multiframe data—specifically, the discrepancy appeared after I implemented processing for multiframe data using pydicom and mapping spacing based on shape:</p>\n<pre><code> (dicom_files) == :\n    ds = pydicom.dcmread(dicom_files[], force=)\n\n    pixel_array = ds.pixel_array.astype(np.int16)\n\n    input_img_np = pixel_array[]\n    original_spacing = get_spacing_by_shape(pixel_array.shape)[::-]\n\n    probs = predict_aneurysm(input_img_np, original_spacing, DEVICE)\n</code></pre>\n<p>After implementing this, the gap consistently remained above 2%. After adding orientation correction for T2 data, the gap increased to over 3%. These are what I refer to as exception handling—there's no need to use try-except and return 0.5. </p>\n<p>I suspect the main reason is that the proportion of multiframe data in the private test set is lower than in the public test set. For detailed inference code, please see: <a href=\"https://www.kaggle.com/code/pengchengshi/bravecowcow-2nd-place-inference\" target=\"_blank\">https://www.kaggle.com/code/pengchengshi/bravecowcow-2nd-place-inference</a></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3303225,
      "author_name": "Tom",
      "author_url": "",
      "post_date": "2025-10-17T12:42:27.913000",
      "content": "<p><a href=\"https://www.kaggle.com/honglihang\" target=\"_blank\">@honglihang</a> that's a good solution, especially utilizing 2.5d segmentation, I would like to use it in my next rsna competition. But seems one of your stage causes overfit, I think improving its generalization can make this solution to 1st.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3303232,
          "author_name": "RihanPiggy",
          "author_url": "",
          "post_date": "2025-10-17T13:25:15.393000",
          "content": "<p>Thank you. I kept the kfold same in all the stages, so I believe there is no leak in training process, and public LB is quite corelated to CV, which is very confusing for me. <br>\nConsidering that test scans are collected from different sites, maybe there is a domain shift in test data which cause the brain segmentator not working well? Another possibility is that there is an issue in dicom tag because I used tags to perform orienting and scaling.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3303246,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2025-10-17T13:56:02.717000",
              "content": "<p>Likely due to domain shift: the segmentor failed to generate valid masks for cases absent from the training set and the public hidden test set but present in the private hidden test set, causing the pipeline to fail.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3303213": "Congratulations to all the winners! Thanks to the organizers for hosting such an interesting competition. It is really a hard one and I enjoyed it a lot. Here let me share my solution.\n\n# TL;DR\nMy solution consists of 4 stages:\n1. Generate brain mask with TotalSegmentator, train a 3D Unet to crop the ROI region which contains CoW Area.\n2. Train 5 fold 2.5D CoW Seg model (cls loss + seg loss) and extract features.\n3. Train 5 fold 2.5D Aneurysm model (cls + keypoint loss) and extract features.\n4. Train RNNs on concatenated features from CoW Seg model and Aneurysm model across depth direction.\n\n# Preprocessing\nI used [dicom2nifti](https://github.com/icometrix/dicom2nifti/tree/main) to convert dicom files to nifti files because TotalSegmentator uses this library.\nI wrote a custom version to transform the coordinates and frame indexes from train_localizers.csv to the new ones which align with the nifti files.\n\n# Stage 1: Brain Mask\n\n## Brain Mask Generation\nI used TotalSegmentator to generate brain masks for all scans.\nStrangely, some of the CT scans have better brain masks generated by TotalSegmentator MRI than TotalSegmentator CT. Thus I generated 2 brain masks for each scan and selected the one with more pixels.\nI also filtered out small masks with less than 1 million pixels. Finally, I have 4386 brain masks.\n\n## 3D Unet for Brain Segmentation\nI trained a 3D Unet to segment the brain. The train code is reimplementation from [Qishen Ha's RSNA2022 solution](https://www.kaggle.com/code/haqishen/rsna-2022-1st-place-solution-train-stage1).\nI first pretrained the model with [CT](https://zenodo.org/records/10047292) and [MRI](https://zenodo.org/records/14710732) scans from Zenodo, which is a subset of the training data used to train TotalSegmentator. Then, I trained the model with 4386 brain masks. The model is trained with 5 folds. The input size is (128, 128, 128).\nI used the trained 3D Unet to generate brain masks and crop the ROI region for all scans. The parameters for cropping are adjusted according to the CoWSeg and train_localizers.csv to ensure the cropped region contains the CoW Area.\n\n# Stage 2: CoW Segmentation Model\nUnlike other competitors, I did not managed to train a good 3D CoW Segmentation model. Instead, I trained a 2.5D CoW Segmentation model with classification head. The purpose of this model is to extract positional features to feed into the final RNN model. According to my experience in RSNA2022, positional features boosted the CV. So although segmentation loss is even worse with 2.5D model (dice=0.56), the extracted features are still useful.\nThe training code is reimplementation from [this amazing notebook](https://www.kaggle.com/code/awsaf49/uwmgi-2-5d-train-pytorch) from UWMGI. The backbone I used is tf_efficientnetv2_b0.in1k (img size=224).\nAlso, using the OOF classification logits from the CoW Seg model like in stage 1, I filter out slices which do not contain CoW area.\n\n# Stage 3: Aneurysm Model\nSimilar to CoW Seg model, I trained a 2.5D Aneurysm model with classification head and keypoint head to extract features.\nThe key for the successful training of this model is as follows:\n- 3 rounds of negative sampling. In the first round, I used classification logits from CoW Seg model to sample slices which contains vessels. In the next 2 rounds, I used OOF classification logits from the Aneurysm model itself to sample hard negative slices.\n- Heuristic cropping to make the model focus on the vessel area. Crop ratio is adjusted according to the coordinates. (ensuring all the coordinates are in the cropped region)\n- Focal loss with alpha=0.75\n- Self distillation with OOF logits to deal with label noise.\n- EMA\n\nThe backbones I used in this stage are:\n- coatnet_rmlp_2_rw_384: 5 fold, img_size=384\n- coatnet_rmlp_2_rw_384: all data train. img_size=384\n- tf_efficientnetv2_s.in21k_ft_in1k: another 5 fold, img_size=384\n\nAll data train used a different revision of dataset with less hard negative images. I first trained a 5 fold model and recorded the epoch with best CV to calculate the total training steps for all data train.\n\n# Stage 4: RNN Model\nI trained 4 types of RNN models for each aneurysm model. The model starts to overfit quickly, EMA and heavy mixup(mixup ratio=0.5) are used to regularize the training.\n\nFor 5 fold Coatnet:\n- Residual LSTM: aneurysm feature + CoW seg feature\n- Residual GRU: aneurysm feature only\n- BiLSTM: aneurysm feature only\n- Bert: aneurysm feature only\n\nFor 5 fold Efficientnet:\n- Residual LSTM: aneurysm feature + CoW seg feature\n- Residual GRU: aneurysm feature + CoW seg feature\n- LSTM: aneurysm feature only\n- Bert: aneurysm feature only\n\nFor all data train Coatnet:\n- Residual LSTM: aneurysm feature + CoW seg feature\n- Residual GRU: aneurysm feature + CoW seg feature\n- BiLSTM: aneurysm feature + CoW seg feature\n- Bert: aneurysm feature + CoW seg feature\n\nThe ensemble of these RNNs reached auc=0.8909 and by self-distillation with OOF logits, another submission reached auc=0.9035\n- submission 1: CV: 0.8909, Public: 0.87689, Private: 0.82201\n- submission 2: CV: 0.9035, Public: 0.86703, Private: 0.81321\nI am still very confused about the huge gap between Public and Private scores.\n\n# Maximizing GPU Utilization\nTo maximize the inference speed, I used Monai to preprocess the images on GPU. I used T4 x 2 and split the patient volume into 2 parts to parallelize the inference of 2.5D models. For RNN models, I simply split the models into 2 groups and run them in parallel on 2 GPUs. That's how I managed to finish the inference of 11 aneurysm models and 45 RNN models within 12 hours. \nI suffered from timeout issue in the final days. Hoping that Kaggle will fix the new evaluation framework.",
    "3303479": "Insightful solution! I observed a similar discrepancy between public and private LB. After implementing strict exception handling by removing a fallback to 0.5 predictions, I noticed a consistent 3-4% gap. I suspect this is due to a difference in the distribution of abnormal data between the public and private test sets.",
    "3303225": "@honglihang that's a good solution, especially utilizing 2.5d segmentation, I would like to use it in my next rsna competition. But seems one of your stage causes overfit, I think improving its generalization can make this solution to 1st."
  }
}