{
  "id": 612774,
  "title": "15th Place Solution",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/612774",
  "author_name": "TakafumiOuchi",
  "post_date": "2025-10-22T02:48:33.913000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks to RSNA and the Kaggle team for hosting such an interesting competition.<br>\nThrough this competition, I was able to learn a great deal about image processing and artificial intelligence.<br>\nHere, let me share my solution overview.</p>\n<h2>Overview</h2>\n<ul>\n<li>Ensemble of two Transformers (weighted average of 14 outputs: 13 locations + aneurysm presence)</li>\n<li>Each Transformer aggregates RoIs extracted by the first-stage Faster R-CNN</li>\n<li>Model 1: 2.5D Faster R-CNN (3-channel input, ResNet-50 backbone) + DeiT-Small Transformer</li>\n<li>Model 2: 2.5D Faster R-CNN (5-channel input, ConvNeXtV2-Tiny backbone) + DeiT-Small Transformer</li>\n</ul>\n<h2>Strategy</h2>\n<p>As a radiologist working at a general hospital, I wanted to develop a model leveraging domain knowledge.<br>\nThrough exploratory data analysis, I found that the dataset contained axial, coronal, and sagittal images.<br>\nI also discovered errors in the direction cosine matrix of some volume data.<br>\nSince the orientation of all images could not be determined solely based on DICOM tags,<br>\nI decided to develop a model that can make inferences regardless of image orientation.</p>\n<h2>Data Pipeline</h2>\n<ul>\n<li>All volumes were resampled to the same voxel spacing (2.0 mm slice thickness, 0.5 mm in-plane pixel spacing)</li>\n<li>Cropped and padded to depth × 384 px × 384 px by a simple rule-based method</li>\n<li>Training: all anatomical planes (axial, coronal, sagittal) extracted from each volume</li>\n<li>Inference: only the finest anatomical plane (lesser interpolation) extracted</li>\n</ul>\n<h2>Stage 1: 2.5D Faster R-CNN</h2>\n<h3>Common Training Strategy</h3>\n<ul>\n<li>Loss: Default of torchvision Faster R-CNN</li>\n<li>Optimizer: AdamW</li>\n<li>Scheduler: OneCycleLR</li>\n<li>Batch size: 32</li>\n<li>Sampling: All positive slices + 3–10 negative slices (gradually increased during training)</li>\n<li>Augmentation:<ul>\n<li>Affine, ElasticTransform, GridDistortion, MotionBlur, RandomBrightnessContrast, GaussNoise using albumentations</li>\n<li>Horizontal flip with label swapping (left ↔ right)</li></ul></li>\n</ul>\n<h3>Model 1</h3>\n<ul>\n<li>Backbone: torchvision/fasterrcnn_resnet50_fpn_v2</li>\n<li>Input: 3-channel (center slice + two adjacent slices, 384 × 384 px)</li>\n<li>Ground truth: 96 px (48 mm) bounding box around aneurysm center</li>\n<li>Anchors: 5 sizes (64, 80, 96, 112, 128 px) × 3 aspect ratios (0.5, 1.0, 2.0)</li>\n<li>Learning rate: 1e-4 → max 1e-3</li>\n</ul>\n<h3>Model 2</h3>\n<ul>\n<li>Backbone: timm/convnextv2_tiny.fcmae_ft_in22k_in1k</li>\n<li>Input: 5-channel (center slice + four adjacent slices, 384 × 384 px)</li>\n<li>Ground truth: 64 px (32 mm) bounding box around aneurysm center</li>\n<li>Anchors: 3 sizes (56, 64, 72 px) × 3 aspect ratios (0.8, 1.0, 1.2)</li>\n<li>Learning rate: 2e-5 → max 2e-4</li>\n</ul>\n<h2>Stage 2: DeiT-Small Transformer</h2>\n<ul>\n<li>Pre-trained weights: timm/deit_small_patch16_224</li>\n<li>Flow: 1024-dim Faster R-CNN RoI features → 384-dim ViT embeddings → 3D sinusoidal positional encoding (based on Faster R-CNN outputs)</li>\n<li>Location head: 13-way multi-label classification (BCEWithLogitsLoss with label smoothing)</li>\n<li>Presence head: Binary classification (BCEWithLogitsLoss with label smoothing)</li>\n<li>Loss weighting: 1.0 × location loss + 0.1 × presence loss</li>\n<li>Two-stage training:<ol>\n<li>Freeze all ViT backbone layers</li>\n<li>Unfreeze all weights</li></ol></li>\n<li>Optimizer: AdamW</li>\n<li>Scheduler: OneCycleLR (stage 2)</li>\n<li>Layer-wise learning rates (stage 2):<ul>\n<li>Classification heads: 4e-4 (max learning rate)</li>\n<li>ViT top 4 blocks: 2e-4 (max learning rate)</li>\n<li>ViT bottom blocks: 1e-4 (max learning rate)</li>\n<li>RoI projector: 1e-4 (max learning rate)</li></ul></li>\n<li>Batch size: 32</li>\n<li>Models were trained using 20× augmented cached RoI features</li>\n</ul>\n<h2>Scores</h2>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model 1 (Faster R-CNN only) *</td>\n<td>0.78602</td>\n<td>0.77517</td>\n</tr>\n<tr>\n<td>Model 2 (Faster R-CNN only) *</td>\n<td>0.78177</td>\n<td>0.75816</td>\n</tr>\n<tr>\n<td>Model 1 (Faster R-CNN + Transformer)</td>\n<td>0.80415</td>\n<td>0.78740</td>\n</tr>\n<tr>\n<td>Model 2 (Faster R-CNN + Transformer)</td>\n<td>0.80140</td>\n<td>0.77955</td>\n</tr>\n<tr>\n<td>Final Ensemble (weighted average)</td>\n<td>0.82494</td>\n<td>0.80592</td>\n</tr>\n<tr>\n<td>*: To calculate 14 probabilities, I extracted max probability per location and used max as aneurysm presence.</td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 3305097,
      "postDate": "2025-10-22T02:48:33.913Z",
      "content": "<p>Thanks to RSNA and the Kaggle team for hosting such an interesting competition.<br>\nThrough this competition, I was able to learn a great deal about image processing and artificial intelligence.<br>\nHere, let me share my solution overview.</p>\n<h2>Overview</h2>\n<ul>\n<li>Ensemble of two Transformers (weighted average of 14 outputs: 13 locations + aneurysm presence)</li>\n<li>Each Transformer aggregates RoIs extracted by the first-stage Faster R-CNN</li>\n<li>Model 1: 2.5D Faster R-CNN (3-channel input, ResNet-50 backbone) + DeiT-Small Transformer</li>\n<li>Model 2: 2.5D Faster R-CNN (5-channel input, ConvNeXtV2-Tiny backbone) + DeiT-Small Transformer</li>\n</ul>\n<h2>Strategy</h2>\n<p>As a radiologist working at a general hospital, I wanted to develop a model leveraging domain knowledge.<br>\nThrough exploratory data analysis, I found that the dataset contained axial, coronal, and sagittal images.<br>\nI also discovered errors in the direction cosine matrix of some volume data.<br>\nSince the orientation of all images could not be determined solely based on DICOM tags,<br>\nI decided to develop a model that can make inferences regardless of image orientation.</p>\n<h2>Data Pipeline</h2>\n<ul>\n<li>All volumes were resampled to the same voxel spacing (2.0 mm slice thickness, 0.5 mm in-plane pixel spacing)</li>\n<li>Cropped and padded to depth × 384 px × 384 px by a simple rule-based method</li>\n<li>Training: all anatomical planes (axial, coronal, sagittal) extracted from each volume</li>\n<li>Inference: only the finest anatomical plane (lesser interpolation) extracted</li>\n</ul>\n<h2>Stage 1: 2.5D Faster R-CNN</h2>\n<h3>Common Training Strategy</h3>\n<ul>\n<li>Loss: Default of torchvision Faster R-CNN</li>\n<li>Optimizer: AdamW</li>\n<li>Scheduler: OneCycleLR</li>\n<li>Batch size: 32</li>\n<li>Sampling: All positive slices + 3–10 negative slices (gradually increased during training)</li>\n<li>Augmentation:<ul>\n<li>Affine, ElasticTransform, GridDistortion, MotionBlur, RandomBrightnessContrast, GaussNoise using albumentations</li>\n<li>Horizontal flip with label swapping (left ↔ right)</li></ul></li>\n</ul>\n<h3>Model 1</h3>\n<ul>\n<li>Backbone: torchvision/fasterrcnn_resnet50_fpn_v2</li>\n<li>Input: 3-channel (center slice + two adjacent slices, 384 × 384 px)</li>\n<li>Ground truth: 96 px (48 mm) bounding box around aneurysm center</li>\n<li>Anchors: 5 sizes (64, 80, 96, 112, 128 px) × 3 aspect ratios (0.5, 1.0, 2.0)</li>\n<li>Learning rate: 1e-4 → max 1e-3</li>\n</ul>\n<h3>Model 2</h3>\n<ul>\n<li>Backbone: timm/convnextv2_tiny.fcmae_ft_in22k_in1k</li>\n<li>Input: 5-channel (center slice + four adjacent slices, 384 × 384 px)</li>\n<li>Ground truth: 64 px (32 mm) bounding box around aneurysm center</li>\n<li>Anchors: 3 sizes (56, 64, 72 px) × 3 aspect ratios (0.8, 1.0, 1.2)</li>\n<li>Learning rate: 2e-5 → max 2e-4</li>\n</ul>\n<h2>Stage 2: DeiT-Small Transformer</h2>\n<ul>\n<li>Pre-trained weights: timm/deit_small_patch16_224</li>\n<li>Flow: 1024-dim Faster R-CNN RoI features → 384-dim ViT embeddings → 3D sinusoidal positional encoding (based on Faster R-CNN outputs)</li>\n<li>Location head: 13-way multi-label classification (BCEWithLogitsLoss with label smoothing)</li>\n<li>Presence head: Binary classification (BCEWithLogitsLoss with label smoothing)</li>\n<li>Loss weighting: 1.0 × location loss + 0.1 × presence loss</li>\n<li>Two-stage training:<ol>\n<li>Freeze all ViT backbone layers</li>\n<li>Unfreeze all weights</li></ol></li>\n<li>Optimizer: AdamW</li>\n<li>Scheduler: OneCycleLR (stage 2)</li>\n<li>Layer-wise learning rates (stage 2):<ul>\n<li>Classification heads: 4e-4 (max learning rate)</li>\n<li>ViT top 4 blocks: 2e-4 (max learning rate)</li>\n<li>ViT bottom blocks: 1e-4 (max learning rate)</li>\n<li>RoI projector: 1e-4 (max learning rate)</li></ul></li>\n<li>Batch size: 32</li>\n<li>Models were trained using 20× augmented cached RoI features</li>\n</ul>\n<h2>Scores</h2>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model 1 (Faster R-CNN only) *</td>\n<td>0.78602</td>\n<td>0.77517</td>\n</tr>\n<tr>\n<td>Model 2 (Faster R-CNN only) *</td>\n<td>0.78177</td>\n<td>0.75816</td>\n</tr>\n<tr>\n<td>Model 1 (Faster R-CNN + Transformer)</td>\n<td>0.80415</td>\n<td>0.78740</td>\n</tr>\n<tr>\n<td>Model 2 (Faster R-CNN + Transformer)</td>\n<td>0.80140</td>\n<td>0.77955</td>\n</tr>\n<tr>\n<td>Final Ensemble (weighted average)</td>\n<td>0.82494</td>\n<td>0.80592</td>\n</tr>\n<tr>\n<td>*: To calculate 14 probabilities, I extracted max probability per location and used max as aneurysm presence.</td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Thanks to RSNA and the Kaggle team for hosting such an interesting competition.\nThrough this competition, I was able to learn a great deal about image processing and artificial intelligence.\nHere, let me share my solution overview.\n\n## Overview\n- Ensemble of two Transformers (weighted average of 14 outputs: 13 locations + aneurysm presence)\n- Each Transformer aggregates RoIs extracted by the first-stage Faster R-CNN\n- Model 1: 2.5D Faster R-CNN (3-channel input, ResNet-50 backbone) + DeiT-Small Transformer\n- Model 2: 2.5D Faster R-CNN (5-channel input, ConvNeXtV2-Tiny backbone) + DeiT-Small Transformer\n\n## Strategy\nAs a radiologist working at a general hospital, I wanted to develop a model leveraging domain knowledge.\nThrough exploratory data analysis, I found that the dataset contained axial, coronal, and sagittal images.\nI also discovered errors in the direction cosine matrix of some volume data.\nSince the orientation of all images could not be determined solely based on DICOM tags,\nI decided to develop a model that can make inferences regardless of image orientation.\n\n## Data Pipeline\n- All volumes were resampled to the same voxel spacing (2.0 mm slice thickness, 0.5 mm in-plane pixel spacing)\n- Cropped and padded to depth × 384 px × 384 px by a simple rule-based method\n- Training: all anatomical planes (axial, coronal, sagittal) extracted from each volume\n- Inference: only the finest anatomical plane (lesser interpolation) extracted\n\n## Stage 1: 2.5D Faster R-CNN\n### Common Training Strategy\n- Loss: Default of torchvision Faster R-CNN\n- Optimizer: AdamW\n- Scheduler: OneCycleLR\n- Batch size: 32\n- Sampling: All positive slices + 3–10 negative slices (gradually increased during training)\n- Augmentation:\n    - Affine, ElasticTransform, GridDistortion, MotionBlur, RandomBrightnessContrast, GaussNoise using albumentations\n    - Horizontal flip with label swapping (left ↔ right)\n\n### Model 1\n- Backbone: torchvision/fasterrcnn_resnet50_fpn_v2\n- Input: 3-channel (center slice + two adjacent slices, 384 × 384 px)\n- Ground truth: 96 px (48 mm) bounding box around aneurysm center\n- Anchors: 5 sizes (64, 80, 96, 112, 128 px) × 3 aspect ratios (0.5, 1.0, 2.0)\n- Learning rate: 1e-4 → max 1e-3\n\n### Model 2\n- Backbone: timm/convnextv2_tiny.fcmae_ft_in22k_in1k\n- Input: 5-channel (center slice + four adjacent slices, 384 × 384 px)\n- Ground truth: 64 px (32 mm) bounding box around aneurysm center\n- Anchors: 3 sizes (56, 64, 72 px) × 3 aspect ratios (0.8, 1.0, 1.2)\n- Learning rate: 2e-5 → max 2e-4\n\n## Stage 2: DeiT-Small Transformer\n- Pre-trained weights: timm/deit_small_patch16_224\n- Flow: 1024-dim Faster R-CNN RoI features → 384-dim ViT embeddings → 3D sinusoidal positional encoding (based on Faster R-CNN outputs)\n- Location head: 13-way multi-label classification (BCEWithLogitsLoss with label smoothing)\n- Presence head: Binary classification (BCEWithLogitsLoss with label smoothing)\n- Loss weighting: 1.0 × location loss + 0.1 × presence loss\n- Two-stage training:\n    1. Freeze all ViT backbone layers\n    2. Unfreeze all weights\n- Optimizer: AdamW\n- Scheduler: OneCycleLR (stage 2)\n- Layer-wise learning rates (stage 2):\n    - Classification heads: 4e-4 (max learning rate)\n    - ViT top 4 blocks: 2e-4 (max learning rate)\n    - ViT bottom blocks: 1e-4 (max learning rate)\n    - RoI projector: 1e-4 (max learning rate)\n- Batch size: 32\n- Models were trained using 20× augmented cached RoI features\n\n## Scores\n| Model | Public | Private |\n|---|---|---|\n| Model 1 (Faster R-CNN only) \\* | 0.78602 | 0.77517 |\n| Model 2 (Faster R-CNN only) \\* | 0.78177 | 0.75816 |\n| Model 1 (Faster R-CNN + Transformer) | 0.80415 | 0.78740 |\n| Model 2 (Faster R-CNN + Transformer) | 0.80140 | 0.77955 |\n| Final Ensemble (weighted average) | 0.82494 | 0.80592 |\n\\*: To calculate 14 probabilities, I extracted max probability per location and used max as aneurysm presence."
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3305097": "Thanks to RSNA and the Kaggle team for hosting such an interesting competition.\nThrough this competition, I was able to learn a great deal about image processing and artificial intelligence.\nHere, let me share my solution overview.\n\n## Overview\n- Ensemble of two Transformers (weighted average of 14 outputs: 13 locations + aneurysm presence)\n- Each Transformer aggregates RoIs extracted by the first-stage Faster R-CNN\n- Model 1: 2.5D Faster R-CNN (3-channel input, ResNet-50 backbone) + DeiT-Small Transformer\n- Model 2: 2.5D Faster R-CNN (5-channel input, ConvNeXtV2-Tiny backbone) + DeiT-Small Transformer\n\n## Strategy\nAs a radiologist working at a general hospital, I wanted to develop a model leveraging domain knowledge.\nThrough exploratory data analysis, I found that the dataset contained axial, coronal, and sagittal images.\nI also discovered errors in the direction cosine matrix of some volume data.\nSince the orientation of all images could not be determined solely based on DICOM tags,\nI decided to develop a model that can make inferences regardless of image orientation.\n\n## Data Pipeline\n- All volumes were resampled to the same voxel spacing (2.0 mm slice thickness, 0.5 mm in-plane pixel spacing)\n- Cropped and padded to depth × 384 px × 384 px by a simple rule-based method\n- Training: all anatomical planes (axial, coronal, sagittal) extracted from each volume\n- Inference: only the finest anatomical plane (lesser interpolation) extracted\n\n## Stage 1: 2.5D Faster R-CNN\n### Common Training Strategy\n- Loss: Default of torchvision Faster R-CNN\n- Optimizer: AdamW\n- Scheduler: OneCycleLR\n- Batch size: 32\n- Sampling: All positive slices + 3–10 negative slices (gradually increased during training)\n- Augmentation:\n    - Affine, ElasticTransform, GridDistortion, MotionBlur, RandomBrightnessContrast, GaussNoise using albumentations\n    - Horizontal flip with label swapping (left ↔ right)\n\n### Model 1\n- Backbone: torchvision/fasterrcnn_resnet50_fpn_v2\n- Input: 3-channel (center slice + two adjacent slices, 384 × 384 px)\n- Ground truth: 96 px (48 mm) bounding box around aneurysm center\n- Anchors: 5 sizes (64, 80, 96, 112, 128 px) × 3 aspect ratios (0.5, 1.0, 2.0)\n- Learning rate: 1e-4 → max 1e-3\n\n### Model 2\n- Backbone: timm/convnextv2_tiny.fcmae_ft_in22k_in1k\n- Input: 5-channel (center slice + four adjacent slices, 384 × 384 px)\n- Ground truth: 64 px (32 mm) bounding box around aneurysm center\n- Anchors: 3 sizes (56, 64, 72 px) × 3 aspect ratios (0.8, 1.0, 1.2)\n- Learning rate: 2e-5 → max 2e-4\n\n## Stage 2: DeiT-Small Transformer\n- Pre-trained weights: timm/deit_small_patch16_224\n- Flow: 1024-dim Faster R-CNN RoI features → 384-dim ViT embeddings → 3D sinusoidal positional encoding (based on Faster R-CNN outputs)\n- Location head: 13-way multi-label classification (BCEWithLogitsLoss with label smoothing)\n- Presence head: Binary classification (BCEWithLogitsLoss with label smoothing)\n- Loss weighting: 1.0 × location loss + 0.1 × presence loss\n- Two-stage training:\n    1. Freeze all ViT backbone layers\n    2. Unfreeze all weights\n- Optimizer: AdamW\n- Scheduler: OneCycleLR (stage 2)\n- Layer-wise learning rates (stage 2):\n    - Classification heads: 4e-4 (max learning rate)\n    - ViT top 4 blocks: 2e-4 (max learning rate)\n    - ViT bottom blocks: 1e-4 (max learning rate)\n    - RoI projector: 1e-4 (max learning rate)\n- Batch size: 32\n- Models were trained using 20× augmented cached RoI features\n\n## Scores\n| Model | Public | Private |\n|---|---|---|\n| Model 1 (Faster R-CNN only) \\* | 0.78602 | 0.77517 |\n| Model 2 (Faster R-CNN only) \\* | 0.78177 | 0.75816 |\n| Model 1 (Faster R-CNN + Transformer) | 0.80415 | 0.78740 |\n| Model 2 (Faster R-CNN + Transformer) | 0.80140 | 0.77955 |\n| Final Ensemble (weighted average) | 0.82494 | 0.80592 |\n\\*: To calculate 14 probabilities, I extracted max probability per location and used max as aneurysm presence."
  }
}