{
  "id": 447553,
  "title": "14th place solution",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/447553",
  "author_name": "Gunes Evitan",
  "post_date": "2023-10-16T11:49:57.922000",
  "votes": 28,
  "comment_count": 3,
  "views": 0,
  "content": "<h2>Overview</h2>\n<p>I used an efficient preprocessing pipeline and small multi-task models in a single stage framework. I didn't use image level labels and segmentation masks because I forgot they were given 🤦‍♂️.</p>\n<p>Kaggle Notebook: <a href=\"https://www.kaggle.com/code/gunesevitan/rsna-2023-abdominal-trauma-detection-inference\" target=\"_blank\">https://www.kaggle.com/code/gunesevitan/rsna-2023-abdominal-trauma-detection-inference</a><br>\nKaggle Dataset: <a href=\"https://www.kaggle.com/datasets/gunesevitan/rsna-2023-abdominal-trauma-detection-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/gunesevitan/rsna-2023-abdominal-trauma-detection-dataset</a><br>\nGitHub Repository: <a href=\"https://github.com/gunesevitan/rsna-2023-abdominal-trauma-detection\" target=\"_blank\">https://github.com/gunesevitan/rsna-2023-abdominal-trauma-detection</a></p>\n<h2>Dataset</h2>\n<h3>2D Dataset</h3>\n<ul>\n<li>Bit shift with DICOM's bits allocated and stored attributes</li>\n<li>Linear pixel value rescale with DICOM's rescale slope and intercept attributes</li>\n<li>Window with DICOM's window width and center attributes (abdominal soft tissue window; width 400, center 50)</li>\n<li>Adjust minimum pixel value to 0 and scale pixel values with the new maximum</li>\n<li>Invert pixel values if DICOM's photometric interpretation attribute is MONOCHROME1</li>\n<li>Multiply pixel values with 255 and cast image to uint8</li>\n<li>Write image in lossless png format with raw size</li>\n</ul>\n<p>My 2D and 3D dataset pipelines are separated because this part can run very fast in parallel because of non-blocking IO. I can export all of the training DICOMs as pngs in approximately 20 minutes.</p>\n<h3>3D Dataset</h3>\n<p>I saved lots of different CT scans from training set as videos and examined them. I noticed each of their start and end points were different on the z dimension. Some of them were starting from the shoulders and ending just before the legs or some of them were starting from the lungs and ending somewhere around middle femur.</p>\n<p>I studied the anatomy and decided to localize ROIs. I manually annotated bounding boxes around the largest contour on axial plane. I labeled slices before the liver as \"upper\" and slices after the femur head as \"lower\". Slices between those two location are labeled as \"abdominal\". I trained a YOLOv8 nano model and it was reaching to 0.99x mAP@50 on all those classes easily. I dropped slices that are predicted as \"upper\" and \"lower\", and I used slices that are predicted as \"abdominal\" and cropped them with the predicted bounding box.<br>\n<img src=\"https://i.ibb.co/JpNsJD1/val-batch2-pred.jpg\" alt=\"yolo\"></p>\n<p>Eventually, I ditched this approach because it was too slow and it didn't improve my overall score at all. In my latest 3D pipeline, I was using a lightweight localization by simply cropping the largest contour on the axial plane and keep all slices on the z dimension.</p>\n<ul>\n<li>Read all images that are exported as pngs in a scan and stack them on the z-axis</li>\n<li>Sort z-axis in descending order by DICOMs' image position patient z attribute</li>\n<li>Flip x-axis if DICOMs' patient position attribute is HFS (head first supine)</li>\n<li>Drop partial slices (some slices at the beginning or end of the scan were partially black)</li>\n</ul>\n<p>I dropped those slices by counting all black vertical lines and their differences on z-axis. Normal slices had 0-5 all black vertical lines. If all black vertical line count suddenly increases or decreases then that slice is partial.</p>\n<pre><code>\n scan.shape[] != :\n    scan_all_zero_vertical_line_transitions = np.diff(np.(scan == , axis=).(axis=))\n    \n    slices_with_all_zero_vertical_lines = (scan_all_zero_vertical_line_transitions &gt; ) | (scan_all_zero_vertical_line_transitions &lt; -)\n    slices_with_all_zero_vertical_lines = np.append(slices_with_all_zero_vertical_lines, slices_with_all_zero_vertical_lines[-])\n    scan = scan[~slices_with_all_zero_vertical_lines]\n     scan_all_zero_vertical_line_transitions, slices_with_all_zero_vertical_lines\n</code></pre>\n<ul>\n<li>Crop the largest contour on the axial plane</li>\n</ul>\n<p>I didn't do that to each image separately because it would break the alignment of slices. I calculated bounding boxes for each slice and calculate the largest bounding box by taking minimum of starting points and maximum of ending points.</p>\n<pre><code>\nlargest_contour_bounding_boxes = np.array([dicom_utilities.get_largest_contour(image)  image  scan])\nlargest_contour_bounding_box = [\n    (largest_contour_bounding_boxes[:, ].()),\n    (largest_contour_bounding_boxes[:, ].()),\n    (largest_contour_bounding_boxes[:, ].()),\n    (largest_contour_bounding_boxes[:, ].()),\n]\nscan = scan[\n    :,\n    largest_contour_bounding_box[]:largest_contour_bounding_box[] + ,\n    largest_contour_bounding_box[]:largest_contour_bounding_box[] + ,\n]\n</code></pre>\n<ul>\n<li>Crop non-zero slices along 3 planes</li>\n</ul>\n<pre><code>\nmmin = np.array((scan &gt; ).nonzero()).(axis=)\nmmax = np.array((scan &gt; ).nonzero()).(axis=)\nscan = scan[\n    mmin[]:mmax[] + ,\n    mmin[]:mmax[] + ,\n    mmin[]:mmax[] + \n]\n</code></pre>\n<ul>\n<li>Resize 3D volume into 96x256x256 with area interpolation</li>\n<li>Write image as a numpy array file</li>\n</ul>\n<p>To conclude, those numpy arrays are used as model inputs. I wasn't able to benefit from parallel execution at this stage. </p>\n<h2>Validation</h2>\n<p>I used multi label stratified group kfold for cross-validation. Group functionality can be achieved by splitting at patient level. I converted one-hot encoded classes into ordinal encoded single columns and created another column for patient scan count. I split dataset into 5 folds and 5 ordinal encoded target columns + patient scan count column are used for stratification.</p>\n<h2>Models</h2>\n<p>I tried lots of different models, heads and necks but two simple models were the best performing ones.</p>\n<h3>MIL-like 2D multi-task classification model</h3>\n<p>This model is a very simple one that is similar to MIL approach and ironically this was my best performing model. The architecture is:</p>\n<ol>\n<li>Extract features on 2D slices</li>\n<li>Average or max pooling on z dimension</li>\n<li>Average, max, gem or attention pooling on x and y dimension</li>\n<li>Dropout</li>\n<li>5 classification heads for each target</li>\n</ol>\n<h3>RNN 2D multi-task classification model</h3>\n<p>This model is similar to what others used in previous competitions. The architecture is:</p>\n<ol>\n<li>Extract features on 2D slices</li>\n<li>Average, max or gem pooling on x and y dimension</li>\n<li>Bidirectional LSTM or GRU max while using z dimension as a sequence </li>\n<li>Dropout</li>\n<li>5 classification heads for each target</li>\n</ol>\n<h3>Backbones, necks and heads</h3>\n<ul>\n<li>I tried lots of backbones from timm and monai but my best backbones were EfficientNet b0, EfficientNet v2 tiny and DenseNet121. I think I wasn't able to make large models converge.</li>\n<li>I also tried lots of different pooling types including average, sum, logsumexp, max, gem, attention but average and attention worked best for the first model and max worked best for the second model.</li>\n<li>I only used 5 regular classification heads for 5 targets<ul>\n<li>n_features x 1 bowel head + sigmoid at inference time</li>\n<li>n_features x 1 extravasation head + sigmoid at inference time</li>\n<li>n_features x 3 kidney head + softmax at inference time</li>\n<li>n_features x 3 liver head + softmax at inference time</li>\n<li>n_features x 3 spleen head + softmax at inference time</li></ul></li>\n</ul>\n<h2>Training</h2>\n<p>I used BCEWithLogitsLoss for bowel and extravasation heads, CrossEntropyLoss for kidney, liver and spleen weights. The only modification I did was implementing exact same sample weights like this:</p>\n<pre><code> ():\n\n     ():\n\n        (SampleWeightedBCEWithLogitsLoss, self).__init__(weight=weight, reduction=reduction)\n\n        self.weight = weight\n        self.reduction = reduction\n\n     ():\n\n        loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction=, weight=self.weight)\n        loss = loss * sample_weights\n\n         self.reduction == :\n            loss = loss.mean()\n         self.reduction == :\n            loss = loss.()\n\n         loss\n</code></pre>\n<pre><code> ():\n\n     ():\n\n        (SampleWeightedCrossEntropyLoss, self).__init__(weight=weight, reduction=reduction)\n\n        self.weight = weight\n        self.reduction = reduction\n\n     ():\n\n        loss = F.cross_entropy(inputs, targets, reduction=, weight=self.weight)\n        loss = loss * sample_weights\n\n         self.reduction == :\n            loss = loss.mean()\n         self.reduction == :\n            loss = loss.()\n\n         loss\n</code></pre>\n<p>Final loss is calculated as the sum of each heads' loss and backward is called on that.</p>\n<p>Training transforms are:</p>\n<ul>\n<li>Scale by max 8 bit pixel value</li>\n<li>Random X, Y and Z flip that are independent of each other</li>\n<li>Random 90 degree rotation on axial plane</li>\n<li>Random 0-45 degree rotation on axial plane</li>\n<li>Histogram equalization or random contrast shift</li>\n<li>Random 224x224 crop on axial plane</li>\n<li>3D cutout</li>\n</ul>\n<p>Test transforms are:</p>\n<ul>\n<li>Scale by max 8 bit pixel value</li>\n<li>Center 224x224 crop on axial plane</li>\n</ul>\n<pre><code>training_transforms = T.Compose([\n    T.EnsureChannelFirst(channel_dim=),\n    T.RandFlip(spatial_axis=, prob=transform_parameters[]),\n    T.RandFlip(spatial_axis=, prob=transform_parameters[]),\n    T.RandFlip(spatial_axis=, prob=transform_parameters[]),\n    T.RandRotate90(spatial_axes=(, ), max_k=, prob=transform_parameters[]),\n    T.RandRotate(\n        range_x=transform_parameters[],\n        range_y=transform_parameters[],\n        range_z=transform_parameters[],\n        prob=transform_parameters[]\n    ),\n    T.OneOf([\n        T.RandHistogramShift(num_control_points=transform_parameters[], prob=transform_parameters[]),\n        T.RandAdjustContrast(gamma=transform_parameters[], prob=transform_parameters[])\n    ], weights=(, )),\n    T.RandSpatialCrop(roi_size=transform_parameters[], max_roi_size=, random_center=, random_size=),\n    T.RandCoarseDropout(\n        holes=transform_parameters[],\n        spatial_size=transform_parameters[],\n        dropout_holes=,\n        fill_value=,\n        max_holes=transform_parameters[],\n        max_spatial_size=transform_parameters[],\n        prob=transform_parameters[]\n    ),\n    T.ToTensor(dtype=torch.float32, track_meta=)\n])\n\ninference_transforms = T.Compose([\n    T.EnsureChannelFirst(channel_dim=),\n    T.CenterSpatialCrop(roi_size=transform_parameters[]),\n    T.ToTensor(dtype=torch.float32, track_meta=)\n])\n</code></pre>\n<pre><code>\n   \n   \n   \n   \n   \n   \n   \n   \n   \n   \n   [, ]\n   \n   [, , ]\n   \n   [, , ]\n   \n   [, , ]\n   \n</code></pre>\n<p>Batch size of 2 or 4 is used depending on the model size. Cosine annealing learning rate schedule is utilized to explore different regions with a small base and minimum learning rate. AMP is also used for faster training and regularization.</p>\n<h2>Inference</h2>\n<p>2x MIL-like model (efficientnetb0 and densenet121) and 2x RNN model (efficientnetb0 and efficientnetv2t) are used on the final ensemble. </p>\n<p>Since the models were trained with random crop augmentation, inputs are center cropped at test time. 4x TTA (xyz, xy, xz and yz flip) are applied and predictions are averaged.</p>\n<p>Predictions of 5 folds are averaged and then activated with sigmoid or softmax functions.</p>\n<h2>Post-processing</h2>\n<p>Different weights are used for 4 models for different targets. Those weights are found by minimizing the OOF score.</p>\n<pre><code>mil_efficientnetb0_bowel_weight = \nmil_densenet121_bowel_weight = \nlstm_efficientnetb0_bowel_weight = \nlstm_efficientnetv2t_bowel_weight = \n\nmil_efficientnetb0_extravasation_weight = \nmil_densenet121_extravasation_weight = \nlstm_efficientnetb0_extravasation_weight = \nlstm_efficientnetv2t_extravasation_weight = \n\nmil_efficientnetb0_kidney_weight = \nmil_densenet121_kidney_weight = \nlstm_efficientnetb0_kidney_weight = \nlstm_efficientnetv2t_kidney_weight = \n\nmil_efficientnetb0_liver_weight = \nmil_densenet121_liver_weight = \nlstm_efficientnetb0_liver_weight = \nlstm_efficientnetv2t_liver_weight = \n\nmil_efficientnetb0_spleen_weight = \nmil_densenet121_spleen_weight = \nlstm_efficientnetb0_spleen_weight = \nlstm_efficientnetv2t_spleen_weight = \n</code></pre>\n<p>I aggregated scan level predictions on patient_id and took the maximum prediction.</p>\n<p>I also scaled injury target predictions with different multipliers and they are also set by minimizing OOF score.</p>\n<pre><code>df_predictions[] *= \ndf_predictions[] *= \ndf_predictions[] *= \ndf_predictions[] *= \ndf_predictions[] *= \ndf_predictions[] *= \ndf_predictions[] *= \ndf_predictions[] *= \n</code></pre>\n<p>My final ensemble score was <strong>0.3859</strong> and target scores are listed below. I really enjoyed how my OOF scores are almost perfectly correlated with LB scores. I selected the submission that had the best OOF, public and private LB scores thanks to stable cross-validation.</p>\n<table>\n<thead>\n<tr>\n<th>bowel</th>\n<th>extravasation</th>\n<th>kidney</th>\n<th>liver</th>\n<th>spleen</th>\n<th>any</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.1282</td>\n<td>0.5070</td>\n<td>0.2831</td>\n<td>0.4186</td>\n<td>0.4736</td>\n<td>0.5050</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 2484316,
      "postDate": "2023-10-16T11:49:57.923Z",
      "content": "<h2>Overview</h2>\n<p>I used an efficient preprocessing pipeline and small multi-task models in a single stage framework. I didn't use image level labels and segmentation masks because I forgot they were given 🤦‍♂️.</p>\n<p>Kaggle Notebook: <a href=\"https://www.kaggle.com/code/gunesevitan/rsna-2023-abdominal-trauma-detection-inference\" target=\"_blank\">https://www.kaggle.com/code/gunesevitan/rsna-2023-abdominal-trauma-detection-inference</a><br>\nKaggle Dataset: <a href=\"https://www.kaggle.com/datasets/gunesevitan/rsna-2023-abdominal-trauma-detection-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/gunesevitan/rsna-2023-abdominal-trauma-detection-dataset</a><br>\nGitHub Repository: <a href=\"https://github.com/gunesevitan/rsna-2023-abdominal-trauma-detection\" target=\"_blank\">https://github.com/gunesevitan/rsna-2023-abdominal-trauma-detection</a></p>\n<h2>Dataset</h2>\n<h3>2D Dataset</h3>\n<ul>\n<li>Bit shift with DICOM's bits allocated and stored attributes</li>\n<li>Linear pixel value rescale with DICOM's rescale slope and intercept attributes</li>\n<li>Window with DICOM's window width and center attributes (abdominal soft tissue window; width 400, center 50)</li>\n<li>Adjust minimum pixel value to 0 and scale pixel values with the new maximum</li>\n<li>Invert pixel values if DICOM's photometric interpretation attribute is MONOCHROME1</li>\n<li>Multiply pixel values with 255 and cast image to uint8</li>\n<li>Write image in lossless png format with raw size</li>\n</ul>\n<p>My 2D and 3D dataset pipelines are separated because this part can run very fast in parallel because of non-blocking IO. I can export all of the training DICOMs as pngs in approximately 20 minutes.</p>\n<h3>3D Dataset</h3>\n<p>I saved lots of different CT scans from training set as videos and examined them. I noticed each of their start and end points were different on the z dimension. Some of them were starting from the shoulders and ending just before the legs or some of them were starting from the lungs and ending somewhere around middle femur.</p>\n<p>I studied the anatomy and decided to localize ROIs. I manually annotated bounding boxes around the largest contour on axial plane. I labeled slices before the liver as \"upper\" and slices after the femur head as \"lower\". Slices between those two location are labeled as \"abdominal\". I trained a YOLOv8 nano model and it was reaching to 0.99x mAP@50 on all those classes easily. I dropped slices that are predicted as \"upper\" and \"lower\", and I used slices that are predicted as \"abdominal\" and cropped them with the predicted bounding box.<br>\n<img src=\"https://i.ibb.co/JpNsJD1/val-batch2-pred.jpg\" alt=\"yolo\"></p>\n<p>Eventually, I ditched this approach because it was too slow and it didn't improve my overall score at all. In my latest 3D pipeline, I was using a lightweight localization by simply cropping the largest contour on the axial plane and keep all slices on the z dimension.</p>\n<ul>\n<li>Read all images that are exported as pngs in a scan and stack them on the z-axis</li>\n<li>Sort z-axis in descending order by DICOMs' image position patient z attribute</li>\n<li>Flip x-axis if DICOMs' patient position attribute is HFS (head first supine)</li>\n<li>Drop partial slices (some slices at the beginning or end of the scan were partially black)</li>\n</ul>\n<p>I dropped those slices by counting all black vertical lines and their differences on z-axis. Normal slices had 0-5 all black vertical lines. If all black vertical line count suddenly increases or decreases then that slice is partial.</p>\n<pre><code>\n scan.shape[] != :\n    scan_all_zero_vertical_line_transitions = np.diff(np.(scan == , axis=).(axis=))\n    \n    slices_with_all_zero_vertical_lines = (scan_all_zero_vertical_line_transitions &gt; ) | (scan_all_zero_vertical_line_transitions &lt; -)\n    slices_with_all_zero_vertical_lines = np.append(slices_with_all_zero_vertical_lines, slices_with_all_zero_vertical_lines[-])\n    scan = scan[~slices_with_all_zero_vertical_lines]\n     scan_all_zero_vertical_line_transitions, slices_with_all_zero_vertical_lines\n</code></pre>\n<ul>\n<li>Crop the largest contour on the axial plane</li>\n</ul>\n<p>I didn't do that to each image separately because it would break the alignment of slices. I calculated bounding boxes for each slice and calculate the largest bounding box by taking minimum of starting points and maximum of ending points.</p>\n<pre><code>\nlargest_contour_bounding_boxes = np.array([dicom_utilities.get_largest_contour(image)  image  scan])\nlargest_contour_bounding_box = [\n    (largest_contour_bounding_boxes[:, ].()),\n    (largest_contour_bounding_boxes[:, ].()),\n    (largest_contour_bounding_boxes[:, ].()),\n    (largest_contour_bounding_boxes[:, ].()),\n]\nscan = scan[\n    :,\n    largest_contour_bounding_box[]:largest_contour_bounding_box[] + ,\n    largest_contour_bounding_box[]:largest_contour_bounding_box[] + ,\n]\n</code></pre>\n<ul>\n<li>Crop non-zero slices along 3 planes</li>\n</ul>\n<pre><code>\nmmin = np.array((scan &gt; ).nonzero()).(axis=)\nmmax = np.array((scan &gt; ).nonzero()).(axis=)\nscan = scan[\n    mmin[]:mmax[] + ,\n    mmin[]:mmax[] + ,\n    mmin[]:mmax[] + \n]\n</code></pre>\n<ul>\n<li>Resize 3D volume into 96x256x256 with area interpolation</li>\n<li>Write image as a numpy array file</li>\n</ul>\n<p>To conclude, those numpy arrays are used as model inputs. I wasn't able to benefit from parallel execution at this stage. </p>\n<h2>Validation</h2>\n<p>I used multi label stratified group kfold for cross-validation. Group functionality can be achieved by splitting at patient level. I converted one-hot encoded classes into ordinal encoded single columns and created another column for patient scan count. I split dataset into 5 folds and 5 ordinal encoded target columns + patient scan count column are used for stratification.</p>\n<h2>Models</h2>\n<p>I tried lots of different models, heads and necks but two simple models were the best performing ones.</p>\n<h3>MIL-like 2D multi-task classification model</h3>\n<p>This model is a very simple one that is similar to MIL approach and ironically this was my best performing model. The architecture is:</p>\n<ol>\n<li>Extract features on 2D slices</li>\n<li>Average or max pooling on z dimension</li>\n<li>Average, max, gem or attention pooling on x and y dimension</li>\n<li>Dropout</li>\n<li>5 classification heads for each target</li>\n</ol>\n<h3>RNN 2D multi-task classification model</h3>\n<p>This model is similar to what others used in previous competitions. The architecture is:</p>\n<ol>\n<li>Extract features on 2D slices</li>\n<li>Average, max or gem pooling on x and y dimension</li>\n<li>Bidirectional LSTM or GRU max while using z dimension as a sequence </li>\n<li>Dropout</li>\n<li>5 classification heads for each target</li>\n</ol>\n<h3>Backbones, necks and heads</h3>\n<ul>\n<li>I tried lots of backbones from timm and monai but my best backbones were EfficientNet b0, EfficientNet v2 tiny and DenseNet121. I think I wasn't able to make large models converge.</li>\n<li>I also tried lots of different pooling types including average, sum, logsumexp, max, gem, attention but average and attention worked best for the first model and max worked best for the second model.</li>\n<li>I only used 5 regular classification heads for 5 targets<ul>\n<li>n_features x 1 bowel head + sigmoid at inference time</li>\n<li>n_features x 1 extravasation head + sigmoid at inference time</li>\n<li>n_features x 3 kidney head + softmax at inference time</li>\n<li>n_features x 3 liver head + softmax at inference time</li>\n<li>n_features x 3 spleen head + softmax at inference time</li></ul></li>\n</ul>\n<h2>Training</h2>\n<p>I used BCEWithLogitsLoss for bowel and extravasation heads, CrossEntropyLoss for kidney, liver and spleen weights. The only modification I did was implementing exact same sample weights like this:</p>\n<pre><code> ():\n\n     ():\n\n        (SampleWeightedBCEWithLogitsLoss, self).__init__(weight=weight, reduction=reduction)\n\n        self.weight = weight\n        self.reduction = reduction\n\n     ():\n\n        loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction=, weight=self.weight)\n        loss = loss * sample_weights\n\n         self.reduction == :\n            loss = loss.mean()\n         self.reduction == :\n            loss = loss.()\n\n         loss\n</code></pre>\n<pre><code> ():\n\n     ():\n\n        (SampleWeightedCrossEntropyLoss, self).__init__(weight=weight, reduction=reduction)\n\n        self.weight = weight\n        self.reduction = reduction\n\n     ():\n\n        loss = F.cross_entropy(inputs, targets, reduction=, weight=self.weight)\n        loss = loss * sample_weights\n\n         self.reduction == :\n            loss = loss.mean()\n         self.reduction == :\n            loss = loss.()\n\n         loss\n</code></pre>\n<p>Final loss is calculated as the sum of each heads' loss and backward is called on that.</p>\n<p>Training transforms are:</p>\n<ul>\n<li>Scale by max 8 bit pixel value</li>\n<li>Random X, Y and Z flip that are independent of each other</li>\n<li>Random 90 degree rotation on axial plane</li>\n<li>Random 0-45 degree rotation on axial plane</li>\n<li>Histogram equalization or random contrast shift</li>\n<li>Random 224x224 crop on axial plane</li>\n<li>3D cutout</li>\n</ul>\n<p>Test transforms are:</p>\n<ul>\n<li>Scale by max 8 bit pixel value</li>\n<li>Center 224x224 crop on axial plane</li>\n</ul>\n<pre><code>training_transforms = T.Compose([\n    T.EnsureChannelFirst(channel_dim=),\n    T.RandFlip(spatial_axis=, prob=transform_parameters[]),\n    T.RandFlip(spatial_axis=, prob=transform_parameters[]),\n    T.RandFlip(spatial_axis=, prob=transform_parameters[]),\n    T.RandRotate90(spatial_axes=(, ), max_k=, prob=transform_parameters[]),\n    T.RandRotate(\n        range_x=transform_parameters[],\n        range_y=transform_parameters[],\n        range_z=transform_parameters[],\n        prob=transform_parameters[]\n    ),\n    T.OneOf([\n        T.RandHistogramShift(num_control_points=transform_parameters[], prob=transform_parameters[]),\n        T.RandAdjustContrast(gamma=transform_parameters[], prob=transform_parameters[])\n    ], weights=(, )),\n    T.RandSpatialCrop(roi_size=transform_parameters[], max_roi_size=, random_center=, random_size=),\n    T.RandCoarseDropout(\n        holes=transform_parameters[],\n        spatial_size=transform_parameters[],\n        dropout_holes=,\n        fill_value=,\n        max_holes=transform_parameters[],\n        max_spatial_size=transform_parameters[],\n        prob=transform_parameters[]\n    ),\n    T.ToTensor(dtype=torch.float32, track_meta=)\n])\n\ninference_transforms = T.Compose([\n    T.EnsureChannelFirst(channel_dim=),\n    T.CenterSpatialCrop(roi_size=transform_parameters[]),\n    T.ToTensor(dtype=torch.float32, track_meta=)\n])\n</code></pre>\n<pre><code>\n   \n   \n   \n   \n   \n   \n   \n   \n   \n   \n   [, ]\n   \n   [, , ]\n   \n   [, , ]\n   \n   [, , ]\n   \n</code></pre>\n<p>Batch size of 2 or 4 is used depending on the model size. Cosine annealing learning rate schedule is utilized to explore different regions with a small base and minimum learning rate. AMP is also used for faster training and regularization.</p>\n<h2>Inference</h2>\n<p>2x MIL-like model (efficientnetb0 and densenet121) and 2x RNN model (efficientnetb0 and efficientnetv2t) are used on the final ensemble. </p>\n<p>Since the models were trained with random crop augmentation, inputs are center cropped at test time. 4x TTA (xyz, xy, xz and yz flip) are applied and predictions are averaged.</p>\n<p>Predictions of 5 folds are averaged and then activated with sigmoid or softmax functions.</p>\n<h2>Post-processing</h2>\n<p>Different weights are used for 4 models for different targets. Those weights are found by minimizing the OOF score.</p>\n<pre><code>mil_efficientnetb0_bowel_weight = \nmil_densenet121_bowel_weight = \nlstm_efficientnetb0_bowel_weight = \nlstm_efficientnetv2t_bowel_weight = \n\nmil_efficientnetb0_extravasation_weight = \nmil_densenet121_extravasation_weight = \nlstm_efficientnetb0_extravasation_weight = \nlstm_efficientnetv2t_extravasation_weight = \n\nmil_efficientnetb0_kidney_weight = \nmil_densenet121_kidney_weight = \nlstm_efficientnetb0_kidney_weight = \nlstm_efficientnetv2t_kidney_weight = \n\nmil_efficientnetb0_liver_weight = \nmil_densenet121_liver_weight = \nlstm_efficientnetb0_liver_weight = \nlstm_efficientnetv2t_liver_weight = \n\nmil_efficientnetb0_spleen_weight = \nmil_densenet121_spleen_weight = \nlstm_efficientnetb0_spleen_weight = \nlstm_efficientnetv2t_spleen_weight = \n</code></pre>\n<p>I aggregated scan level predictions on patient_id and took the maximum prediction.</p>\n<p>I also scaled injury target predictions with different multipliers and they are also set by minimizing OOF score.</p>\n<pre><code>df_predictions[] *= \ndf_predictions[] *= \ndf_predictions[] *= \ndf_predictions[] *= \ndf_predictions[] *= \ndf_predictions[] *= \ndf_predictions[] *= \ndf_predictions[] *= \n</code></pre>\n<p>My final ensemble score was <strong>0.3859</strong> and target scores are listed below. I really enjoyed how my OOF scores are almost perfectly correlated with LB scores. I selected the submission that had the best OOF, public and private LB scores thanks to stable cross-validation.</p>\n<table>\n<thead>\n<tr>\n<th>bowel</th>\n<th>extravasation</th>\n<th>kidney</th>\n<th>liver</th>\n<th>spleen</th>\n<th>any</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.1282</td>\n<td>0.5070</td>\n<td>0.2831</td>\n<td>0.4186</td>\n<td>0.4736</td>\n<td>0.5050</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "## Overview\n\nI used an efficient preprocessing pipeline and small multi-task models in a single stage framework. I didn't use image level labels and segmentation masks because I forgot they were given 🤦‍♂️.\n\nKaggle Notebook: https://www.kaggle.com/code/gunesevitan/rsna-2023-abdominal-trauma-detection-inference\nKaggle Dataset: https://www.kaggle.com/datasets/gunesevitan/rsna-2023-abdominal-trauma-detection-dataset\nGitHub Repository: https://github.com/gunesevitan/rsna-2023-abdominal-trauma-detection\n\n## Dataset\n\n### 2D Dataset\n\n* Bit shift with DICOM's bits allocated and stored attributes\n* Linear pixel value rescale with DICOM's rescale slope and intercept attributes\n* Window with DICOM's window width and center attributes (abdominal soft tissue window; width 400, center 50)\n* Adjust minimum pixel value to 0 and scale pixel values with the new maximum\n* Invert pixel values if DICOM's photometric interpretation attribute is MONOCHROME1\n* Multiply pixel values with 255 and cast image to uint8\n* Write image in lossless png format with raw size\n\nMy 2D and 3D dataset pipelines are separated because this part can run very fast in parallel because of non-blocking IO. I can export all of the training DICOMs as pngs in approximately 20 minutes.\n\n### 3D Dataset\n\nI saved lots of different CT scans from training set as videos and examined them. I noticed each of their start and end points were different on the z dimension. Some of them were starting from the shoulders and ending just before the legs or some of them were starting from the lungs and ending somewhere around middle femur.\n\nI studied the anatomy and decided to localize ROIs. I manually annotated bounding boxes around the largest contour on axial plane. I labeled slices before the liver as \"upper\" and slices after the femur head as \"lower\". Slices between those two location are labeled as \"abdominal\". I trained a YOLOv8 nano model and it was reaching to 0.99x mAP@50 on all those classes easily. I dropped slices that are predicted as \"upper\" and \"lower\", and I used slices that are predicted as \"abdominal\" and cropped them with the predicted bounding box.\n![yolo](https://i.ibb.co/JpNsJD1/val-batch2-pred.jpg)\n\nEventually, I ditched this approach because it was too slow and it didn't improve my overall score at all. In my latest 3D pipeline, I was using a lightweight localization by simply cropping the largest contour on the axial plane and keep all slices on the z dimension.\n\n* Read all images that are exported as pngs in a scan and stack them on the z-axis\n* Sort z-axis in descending order by DICOMs' image position patient z attribute\n* Flip x-axis if DICOMs' patient position attribute is HFS (head first supine)\n* Drop partial slices (some slices at the beginning or end of the scan were partially black)\n\nI dropped those slices by counting all black vertical lines and their differences on z-axis. Normal slices had 0-5 all black vertical lines. If all black vertical line count suddenly increases or decreases then that slice is partial.\n```python\n# Find partial slices by calculating sum of all zero vertical lines\nif scan.shape[0] != 1:\n    scan_all_zero_vertical_line_transitions = np.diff(np.all(scan == 0, axis=1).sum(axis=1))\n    # Heuristically select high and low transitions on z-axis and drop them\n    slices_with_all_zero_vertical_lines = (scan_all_zero_vertical_line_transitions > 5) | (scan_all_zero_vertical_line_transitions < -5)\n    slices_with_all_zero_vertical_lines = np.append(slices_with_all_zero_vertical_lines, slices_with_all_zero_vertical_lines[-1])\n    scan = scan[~slices_with_all_zero_vertical_lines]\n    del scan_all_zero_vertical_line_transitions, slices_with_all_zero_vertical_lines\n```\n\n* Crop the largest contour on the axial plane\n\nI didn't do that to each image separately because it would break the alignment of slices. I calculated bounding boxes for each slice and calculate the largest bounding box by taking minimum of starting points and maximum of ending points.\n\n```python\n# Crop the largest contour\nlargest_contour_bounding_boxes = np.array([dicom_utilities.get_largest_contour(image) for image in scan])\nlargest_contour_bounding_box = [\n    int(largest_contour_bounding_boxes[:, 0].min()),\n    int(largest_contour_bounding_boxes[:, 1].min()),\n    int(largest_contour_bounding_boxes[:, 2].max()),\n    int(largest_contour_bounding_boxes[:, 3].max()),\n]\nscan = scan[\n    :,\n    largest_contour_bounding_box[1]:largest_contour_bounding_box[3] + 1,\n    largest_contour_bounding_box[0]:largest_contour_bounding_box[2] + 1,\n]\n```\n\n* Crop non-zero slices along 3 planes\n```python\n# Crop non-zero slices along xz, yz and xy planes\nmmin = np.array((scan > 0).nonzero()).min(axis=1)\nmmax = np.array((scan > 0).nonzero()).max(axis=1)\nscan = scan[\n    mmin[0]:mmax[0] + 1,\n    mmin[1]:mmax[1] + 1,\n    mmin[2]:mmax[2] + 1\n]\n```\n* Resize 3D volume into 96x256x256 with area interpolation\n* Write image as a numpy array file\n\nTo conclude, those numpy arrays are used as model inputs. I wasn't able to benefit from parallel execution at this stage. \n\n## Validation\n\nI used multi label stratified group kfold for cross-validation. Group functionality can be achieved by splitting at patient level. I converted one-hot encoded classes into ordinal encoded single columns and created another column for patient scan count. I split dataset into 5 folds and 5 ordinal encoded target columns + patient scan count column are used for stratification.\n\n## Models\n\nI tried lots of different models, heads and necks but two simple models were the best performing ones.\n\n### MIL-like 2D multi-task classification model\n\nThis model is a very simple one that is similar to MIL approach and ironically this was my best performing model. The architecture is:\n1. Extract features on 2D slices\n2. Average or max pooling on z dimension\n3. Average, max, gem or attention pooling on x and y dimension\n4. Dropout\n5. 5 classification heads for each target\n\n### RNN 2D multi-task classification model\n\nThis model is similar to what others used in previous competitions. The architecture is:\n1. Extract features on 2D slices\n2. Average, max or gem pooling on x and y dimension\n3. Bidirectional LSTM or GRU max while using z dimension as a sequence \n4. Dropout\n5. 5 classification heads for each target\n\n### Backbones, necks and heads\n\n* I tried lots of backbones from timm and monai but my best backbones were EfficientNet b0, EfficientNet v2 tiny and DenseNet121. I think I wasn't able to make large models converge.\n* I also tried lots of different pooling types including average, sum, logsumexp, max, gem, attention but average and attention worked best for the first model and max worked best for the second model.\n* I only used 5 regular classification heads for 5 targets\n  * n_features x 1 bowel head + sigmoid at inference time\n  * n_features x 1 extravasation head + sigmoid at inference time\n  * n_features x 3 kidney head + softmax at inference time\n  * n_features x 3 liver head + softmax at inference time\n  * n_features x 3 spleen head + softmax at inference time\n\n## Training\n\nI used BCEWithLogitsLoss for bowel and extravasation heads, CrossEntropyLoss for kidney, liver and spleen weights. The only modification I did was implementing exact same sample weights like this:\n\n```python\nclass SampleWeightedBCEWithLogitsLoss(_WeightedLoss):\n\n    def __init__(self, weight=None, reduction='mean'):\n\n        super(SampleWeightedBCEWithLogitsLoss, self).__init__(weight=weight, reduction=reduction)\n\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets, sample_weights):\n\n        loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction='none', weight=self.weight)\n        loss = loss * sample_weights\n\n        if self.reduction == 'mean':\n            loss = loss.mean()\n        elif self.reduction == 'sum':\n            loss = loss.sum()\n\n        return loss\n```\n\n```python\nclass SampleWeightedCrossEntropyLoss(_WeightedLoss):\n\n    def __init__(self, weight=None, reduction='mean'):\n\n        super(SampleWeightedCrossEntropyLoss, self).__init__(weight=weight, reduction=reduction)\n\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets, sample_weights):\n\n        loss = F.cross_entropy(inputs, targets, reduction='none', weight=self.weight)\n        loss = loss * sample_weights\n\n        if self.reduction == 'mean':\n            loss = loss.mean()\n        elif self.reduction == 'sum':\n            loss = loss.sum()\n\n        return loss\n```\n\nFinal loss is calculated as the sum of each heads' loss and backward is called on that.\n\nTraining transforms are:\n* Scale by max 8 bit pixel value\n* Random X, Y and Z flip that are independent of each other\n* Random 90 degree rotation on axial plane\n* Random 0-45 degree rotation on axial plane\n* Histogram equalization or random contrast shift\n* Random 224x224 crop on axial plane\n* 3D cutout\n\nTest transforms are:\n* Scale by max 8 bit pixel value\n* Center 224x224 crop on axial plane\n\n```python\ntraining_transforms = T.Compose([\n    T.EnsureChannelFirst(channel_dim=0),\n    T.RandFlip(spatial_axis=0, prob=transform_parameters['random_z_flip_probability']),\n    T.RandFlip(spatial_axis=1, prob=transform_parameters['random_x_flip_probability']),\n    T.RandFlip(spatial_axis=2, prob=transform_parameters['random_y_flip_probability']),\n    T.RandRotate90(spatial_axes=(1, 2), max_k=3, prob=transform_parameters['random_axial_rotate_90_probability']),\n    T.RandRotate(\n        range_x=transform_parameters['random_rotate_range_x'],\n        range_y=transform_parameters['random_rotate_range_y'],\n        range_z=transform_parameters['random_rotate_range_z'],\n        prob=transform_parameters['random_rotate_probability']\n    ),\n    T.OneOf([\n        T.RandHistogramShift(num_control_points=transform_parameters['random_histogram_shift_num_control_points'], prob=transform_parameters['random_histogram_shift_probability']),\n        T.RandAdjustContrast(gamma=transform_parameters['random_contrast_gamma'], prob=transform_parameters['random_contrast_probability'])\n    ], weights=(0.5, 0.5)),\n    T.RandSpatialCrop(roi_size=transform_parameters['crop_roi_size'], max_roi_size=None, random_center=True, random_size=False),\n    T.RandCoarseDropout(\n        holes=transform_parameters['cutout_holes'],\n        spatial_size=transform_parameters['cutout_spatial_size'],\n        dropout_holes=True,\n        fill_value=0,\n        max_holes=transform_parameters['cutout_max_holes'],\n        max_spatial_size=transform_parameters['max_spatial_size'],\n        prob=transform_parameters['cutout_probability']\n    ),\n    T.ToTensor(dtype=torch.float32, track_meta=False)\n])\n\ninference_transforms = T.Compose([\n    T.EnsureChannelFirst(channel_dim=0),\n    T.CenterSpatialCrop(roi_size=transform_parameters['crop_roi_size']),\n    T.ToTensor(dtype=torch.float32, track_meta=False)\n])\n```\n\n```yaml\ntransforms:\n  random_z_flip_probability: 0.5\n  random_x_flip_probability: 0.5\n  random_y_flip_probability: 0.5\n  random_axial_rotate_90_probability: 0.25\n  random_rotate_range_x: 0.0\n  random_rotate_range_y: 0.0\n  random_rotate_range_z: 0.45\n  random_rotate_probability: 0.1\n  random_histogram_shift_num_control_points: 25\n  random_histogram_shift_probability: 0.05\n  random_contrast_gamma: [0.5, 2.5]\n  random_contrast_probability: 0.05\n  crop_roi_size: [-1, 224, 224]\n  cutout_holes: 8\n  cutout_spatial_size: [4, 8, 8]\n  cutout_max_holes: 16\n  max_spatial_size: [8, 16, 16]\n  cutout_probability: 0.1\n```\n\nBatch size of 2 or 4 is used depending on the model size. Cosine annealing learning rate schedule is utilized to explore different regions with a small base and minimum learning rate. AMP is also used for faster training and regularization.\n\n## Inference\n\n2x MIL-like model (efficientnetb0 and densenet121) and 2x RNN model (efficientnetb0 and efficientnetv2t) are used on the final ensemble. \n\nSince the models were trained with random crop augmentation, inputs are center cropped at test time. 4x TTA (xyz, xy, xz and yz flip) are applied and predictions are averaged.\n\nPredictions of 5 folds are averaged and then activated with sigmoid or softmax functions.\n\n## Post-processing\n\nDifferent weights are used for 4 models for different targets. Those weights are found by minimizing the OOF score.\n\n```python\nmil_efficientnetb0_bowel_weight = 0.45\nmil_densenet121_bowel_weight = 0.25\nlstm_efficientnetb0_bowel_weight = 0.15\nlstm_efficientnetv2t_bowel_weight = 0.15\n\nmil_efficientnetb0_extravasation_weight = 0.3\nmil_densenet121_extravasation_weight = 0.3\nlstm_efficientnetb0_extravasation_weight = 0.3\nlstm_efficientnetv2t_extravasation_weight = 0.1\n\nmil_efficientnetb0_kidney_weight = 0.25\nmil_densenet121_kidney_weight = 0.25\nlstm_efficientnetb0_kidney_weight = 0.25\nlstm_efficientnetv2t_kidney_weight = 0.25\n\nmil_efficientnetb0_liver_weight = 0.25\nmil_densenet121_liver_weight = 0.25\nlstm_efficientnetb0_liver_weight = 0.25\nlstm_efficientnetv2t_liver_weight = 0.25\n\nmil_efficientnetb0_spleen_weight = 0.25\nmil_densenet121_spleen_weight = 0.25\nlstm_efficientnetb0_spleen_weight = 0.25\nlstm_efficientnetv2t_spleen_weight = 0.25\n```\n\nI aggregated scan level predictions on patient_id and took the maximum prediction.\n\nI also scaled injury target predictions with different multipliers and they are also set by minimizing OOF score.\n\n```python\ndf_predictions['bowel_injury_prediction'] *= 1.\ndf_predictions['extravasation_injury_prediction'] *= 1.4\ndf_predictions['kidney_low_prediction'] *= 1.1\ndf_predictions['kidney_high_prediction'] *= 1.1\ndf_predictions['liver_low_prediction'] *= 1.3\ndf_predictions['liver_high_prediction'] *= 1.3\ndf_predictions['spleen_low_prediction'] *= 1.75\ndf_predictions['spleen_high_prediction'] *= 1.75\n```\n\nMy final ensemble score was **0.3859** and target scores are listed below. I really enjoyed how my OOF scores are almost perfectly correlated with LB scores. I selected the submission that had the best OOF, public and private LB scores thanks to stable cross-validation.\n\n| bowel  | extravasation | kidney | liver  | spleen | any    |\n|--------|---------------|--------|--------|--------|--------|\n| 0.1282 | 0.5070        | 0.2831 | 0.4186 | 0.4736 | 0.5050 |\n",
      "votes": 28
    },
    {
      "id": 2484772,
      "postDate": "2023-10-16T16:57:06.163Z",
      "content": "<p>Pretty impressive solution without any segmentation!</p>",
      "rawMarkdown": "Pretty impressive solution without any segmentation!",
      "votes": 1
    },
    {
      "id": 2484550,
      "postDate": "2023-10-16T15:04:12.403Z",
      "content": "<p>Nice work <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>. Thanks for the write up.</p>",
      "rawMarkdown": "Nice work @gunesevitan. Thanks for the write up.",
      "votes": 1
    },
    {
      "id": 2484384,
      "postDate": "2023-10-16T12:56:38.193Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> on this achievement .</p>",
      "rawMarkdown": "Congrats @gunesevitan on this achievement .",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2484772,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2023-10-16T16:57:06.163000",
      "content": "<p>Pretty impressive solution without any segmentation!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2484550,
      "author_name": "David Roberts",
      "author_url": "",
      "post_date": "2023-10-16T15:04:12.403000",
      "content": "<p>Nice work <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>. Thanks for the write up.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2484384,
      "author_name": "SOUMENDRA PRASAD MOHANTY",
      "author_url": "",
      "post_date": "2023-10-16T12:56:38.193000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> on this achievement .</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2484316": "## Overview\n\nI used an efficient preprocessing pipeline and small multi-task models in a single stage framework. I didn't use image level labels and segmentation masks because I forgot they were given 🤦‍♂️.\n\nKaggle Notebook: https://www.kaggle.com/code/gunesevitan/rsna-2023-abdominal-trauma-detection-inference\nKaggle Dataset: https://www.kaggle.com/datasets/gunesevitan/rsna-2023-abdominal-trauma-detection-dataset\nGitHub Repository: https://github.com/gunesevitan/rsna-2023-abdominal-trauma-detection\n\n## Dataset\n\n### 2D Dataset\n\n* Bit shift with DICOM's bits allocated and stored attributes\n* Linear pixel value rescale with DICOM's rescale slope and intercept attributes\n* Window with DICOM's window width and center attributes (abdominal soft tissue window; width 400, center 50)\n* Adjust minimum pixel value to 0 and scale pixel values with the new maximum\n* Invert pixel values if DICOM's photometric interpretation attribute is MONOCHROME1\n* Multiply pixel values with 255 and cast image to uint8\n* Write image in lossless png format with raw size\n\nMy 2D and 3D dataset pipelines are separated because this part can run very fast in parallel because of non-blocking IO. I can export all of the training DICOMs as pngs in approximately 20 minutes.\n\n### 3D Dataset\n\nI saved lots of different CT scans from training set as videos and examined them. I noticed each of their start and end points were different on the z dimension. Some of them were starting from the shoulders and ending just before the legs or some of them were starting from the lungs and ending somewhere around middle femur.\n\nI studied the anatomy and decided to localize ROIs. I manually annotated bounding boxes around the largest contour on axial plane. I labeled slices before the liver as \"upper\" and slices after the femur head as \"lower\". Slices between those two location are labeled as \"abdominal\". I trained a YOLOv8 nano model and it was reaching to 0.99x mAP@50 on all those classes easily. I dropped slices that are predicted as \"upper\" and \"lower\", and I used slices that are predicted as \"abdominal\" and cropped them with the predicted bounding box.\n![yolo](https://i.ibb.co/JpNsJD1/val-batch2-pred.jpg)\n\nEventually, I ditched this approach because it was too slow and it didn't improve my overall score at all. In my latest 3D pipeline, I was using a lightweight localization by simply cropping the largest contour on the axial plane and keep all slices on the z dimension.\n\n* Read all images that are exported as pngs in a scan and stack them on the z-axis\n* Sort z-axis in descending order by DICOMs' image position patient z attribute\n* Flip x-axis if DICOMs' patient position attribute is HFS (head first supine)\n* Drop partial slices (some slices at the beginning or end of the scan were partially black)\n\nI dropped those slices by counting all black vertical lines and their differences on z-axis. Normal slices had 0-5 all black vertical lines. If all black vertical line count suddenly increases or decreases then that slice is partial.\n```python\n# Find partial slices by calculating sum of all zero vertical lines\nif scan.shape[0] != 1:\n    scan_all_zero_vertical_line_transitions = np.diff(np.all(scan == 0, axis=1).sum(axis=1))\n    # Heuristically select high and low transitions on z-axis and drop them\n    slices_with_all_zero_vertical_lines = (scan_all_zero_vertical_line_transitions > 5) | (scan_all_zero_vertical_line_transitions < -5)\n    slices_with_all_zero_vertical_lines = np.append(slices_with_all_zero_vertical_lines, slices_with_all_zero_vertical_lines[-1])\n    scan = scan[~slices_with_all_zero_vertical_lines]\n    del scan_all_zero_vertical_line_transitions, slices_with_all_zero_vertical_lines\n```\n\n* Crop the largest contour on the axial plane\n\nI didn't do that to each image separately because it would break the alignment of slices. I calculated bounding boxes for each slice and calculate the largest bounding box by taking minimum of starting points and maximum of ending points.\n\n```python\n# Crop the largest contour\nlargest_contour_bounding_boxes = np.array([dicom_utilities.get_largest_contour(image) for image in scan])\nlargest_contour_bounding_box = [\n    int(largest_contour_bounding_boxes[:, 0].min()),\n    int(largest_contour_bounding_boxes[:, 1].min()),\n    int(largest_contour_bounding_boxes[:, 2].max()),\n    int(largest_contour_bounding_boxes[:, 3].max()),\n]\nscan = scan[\n    :,\n    largest_contour_bounding_box[1]:largest_contour_bounding_box[3] + 1,\n    largest_contour_bounding_box[0]:largest_contour_bounding_box[2] + 1,\n]\n```\n\n* Crop non-zero slices along 3 planes\n```python\n# Crop non-zero slices along xz, yz and xy planes\nmmin = np.array((scan > 0).nonzero()).min(axis=1)\nmmax = np.array((scan > 0).nonzero()).max(axis=1)\nscan = scan[\n    mmin[0]:mmax[0] + 1,\n    mmin[1]:mmax[1] + 1,\n    mmin[2]:mmax[2] + 1\n]\n```\n* Resize 3D volume into 96x256x256 with area interpolation\n* Write image as a numpy array file\n\nTo conclude, those numpy arrays are used as model inputs. I wasn't able to benefit from parallel execution at this stage. \n\n## Validation\n\nI used multi label stratified group kfold for cross-validation. Group functionality can be achieved by splitting at patient level. I converted one-hot encoded classes into ordinal encoded single columns and created another column for patient scan count. I split dataset into 5 folds and 5 ordinal encoded target columns + patient scan count column are used for stratification.\n\n## Models\n\nI tried lots of different models, heads and necks but two simple models were the best performing ones.\n\n### MIL-like 2D multi-task classification model\n\nThis model is a very simple one that is similar to MIL approach and ironically this was my best performing model. The architecture is:\n1. Extract features on 2D slices\n2. Average or max pooling on z dimension\n3. Average, max, gem or attention pooling on x and y dimension\n4. Dropout\n5. 5 classification heads for each target\n\n### RNN 2D multi-task classification model\n\nThis model is similar to what others used in previous competitions. The architecture is:\n1. Extract features on 2D slices\n2. Average, max or gem pooling on x and y dimension\n3. Bidirectional LSTM or GRU max while using z dimension as a sequence \n4. Dropout\n5. 5 classification heads for each target\n\n### Backbones, necks and heads\n\n* I tried lots of backbones from timm and monai but my best backbones were EfficientNet b0, EfficientNet v2 tiny and DenseNet121. I think I wasn't able to make large models converge.\n* I also tried lots of different pooling types including average, sum, logsumexp, max, gem, attention but average and attention worked best for the first model and max worked best for the second model.\n* I only used 5 regular classification heads for 5 targets\n  * n_features x 1 bowel head + sigmoid at inference time\n  * n_features x 1 extravasation head + sigmoid at inference time\n  * n_features x 3 kidney head + softmax at inference time\n  * n_features x 3 liver head + softmax at inference time\n  * n_features x 3 spleen head + softmax at inference time\n\n## Training\n\nI used BCEWithLogitsLoss for bowel and extravasation heads, CrossEntropyLoss for kidney, liver and spleen weights. The only modification I did was implementing exact same sample weights like this:\n\n```python\nclass SampleWeightedBCEWithLogitsLoss(_WeightedLoss):\n\n    def __init__(self, weight=None, reduction='mean'):\n\n        super(SampleWeightedBCEWithLogitsLoss, self).__init__(weight=weight, reduction=reduction)\n\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets, sample_weights):\n\n        loss = F.binary_cross_entropy_with_logits(inputs, targets, reduction='none', weight=self.weight)\n        loss = loss * sample_weights\n\n        if self.reduction == 'mean':\n            loss = loss.mean()\n        elif self.reduction == 'sum':\n            loss = loss.sum()\n\n        return loss\n```\n\n```python\nclass SampleWeightedCrossEntropyLoss(_WeightedLoss):\n\n    def __init__(self, weight=None, reduction='mean'):\n\n        super(SampleWeightedCrossEntropyLoss, self).__init__(weight=weight, reduction=reduction)\n\n        self.weight = weight\n        self.reduction = reduction\n\n    def forward(self, inputs, targets, sample_weights):\n\n        loss = F.cross_entropy(inputs, targets, reduction='none', weight=self.weight)\n        loss = loss * sample_weights\n\n        if self.reduction == 'mean':\n            loss = loss.mean()\n        elif self.reduction == 'sum':\n            loss = loss.sum()\n\n        return loss\n```\n\nFinal loss is calculated as the sum of each heads' loss and backward is called on that.\n\nTraining transforms are:\n* Scale by max 8 bit pixel value\n* Random X, Y and Z flip that are independent of each other\n* Random 90 degree rotation on axial plane\n* Random 0-45 degree rotation on axial plane\n* Histogram equalization or random contrast shift\n* Random 224x224 crop on axial plane\n* 3D cutout\n\nTest transforms are:\n* Scale by max 8 bit pixel value\n* Center 224x224 crop on axial plane\n\n```python\ntraining_transforms = T.Compose([\n    T.EnsureChannelFirst(channel_dim=0),\n    T.RandFlip(spatial_axis=0, prob=transform_parameters['random_z_flip_probability']),\n    T.RandFlip(spatial_axis=1, prob=transform_parameters['random_x_flip_probability']),\n    T.RandFlip(spatial_axis=2, prob=transform_parameters['random_y_flip_probability']),\n    T.RandRotate90(spatial_axes=(1, 2), max_k=3, prob=transform_parameters['random_axial_rotate_90_probability']),\n    T.RandRotate(\n        range_x=transform_parameters['random_rotate_range_x'],\n        range_y=transform_parameters['random_rotate_range_y'],\n        range_z=transform_parameters['random_rotate_range_z'],\n        prob=transform_parameters['random_rotate_probability']\n    ),\n    T.OneOf([\n        T.RandHistogramShift(num_control_points=transform_parameters['random_histogram_shift_num_control_points'], prob=transform_parameters['random_histogram_shift_probability']),\n        T.RandAdjustContrast(gamma=transform_parameters['random_contrast_gamma'], prob=transform_parameters['random_contrast_probability'])\n    ], weights=(0.5, 0.5)),\n    T.RandSpatialCrop(roi_size=transform_parameters['crop_roi_size'], max_roi_size=None, random_center=True, random_size=False),\n    T.RandCoarseDropout(\n        holes=transform_parameters['cutout_holes'],\n        spatial_size=transform_parameters['cutout_spatial_size'],\n        dropout_holes=True,\n        fill_value=0,\n        max_holes=transform_parameters['cutout_max_holes'],\n        max_spatial_size=transform_parameters['max_spatial_size'],\n        prob=transform_parameters['cutout_probability']\n    ),\n    T.ToTensor(dtype=torch.float32, track_meta=False)\n])\n\ninference_transforms = T.Compose([\n    T.EnsureChannelFirst(channel_dim=0),\n    T.CenterSpatialCrop(roi_size=transform_parameters['crop_roi_size']),\n    T.ToTensor(dtype=torch.float32, track_meta=False)\n])\n```\n\n```yaml\ntransforms:\n  random_z_flip_probability: 0.5\n  random_x_flip_probability: 0.5\n  random_y_flip_probability: 0.5\n  random_axial_rotate_90_probability: 0.25\n  random_rotate_range_x: 0.0\n  random_rotate_range_y: 0.0\n  random_rotate_range_z: 0.45\n  random_rotate_probability: 0.1\n  random_histogram_shift_num_control_points: 25\n  random_histogram_shift_probability: 0.05\n  random_contrast_gamma: [0.5, 2.5]\n  random_contrast_probability: 0.05\n  crop_roi_size: [-1, 224, 224]\n  cutout_holes: 8\n  cutout_spatial_size: [4, 8, 8]\n  cutout_max_holes: 16\n  max_spatial_size: [8, 16, 16]\n  cutout_probability: 0.1\n```\n\nBatch size of 2 or 4 is used depending on the model size. Cosine annealing learning rate schedule is utilized to explore different regions with a small base and minimum learning rate. AMP is also used for faster training and regularization.\n\n## Inference\n\n2x MIL-like model (efficientnetb0 and densenet121) and 2x RNN model (efficientnetb0 and efficientnetv2t) are used on the final ensemble. \n\nSince the models were trained with random crop augmentation, inputs are center cropped at test time. 4x TTA (xyz, xy, xz and yz flip) are applied and predictions are averaged.\n\nPredictions of 5 folds are averaged and then activated with sigmoid or softmax functions.\n\n## Post-processing\n\nDifferent weights are used for 4 models for different targets. Those weights are found by minimizing the OOF score.\n\n```python\nmil_efficientnetb0_bowel_weight = 0.45\nmil_densenet121_bowel_weight = 0.25\nlstm_efficientnetb0_bowel_weight = 0.15\nlstm_efficientnetv2t_bowel_weight = 0.15\n\nmil_efficientnetb0_extravasation_weight = 0.3\nmil_densenet121_extravasation_weight = 0.3\nlstm_efficientnetb0_extravasation_weight = 0.3\nlstm_efficientnetv2t_extravasation_weight = 0.1\n\nmil_efficientnetb0_kidney_weight = 0.25\nmil_densenet121_kidney_weight = 0.25\nlstm_efficientnetb0_kidney_weight = 0.25\nlstm_efficientnetv2t_kidney_weight = 0.25\n\nmil_efficientnetb0_liver_weight = 0.25\nmil_densenet121_liver_weight = 0.25\nlstm_efficientnetb0_liver_weight = 0.25\nlstm_efficientnetv2t_liver_weight = 0.25\n\nmil_efficientnetb0_spleen_weight = 0.25\nmil_densenet121_spleen_weight = 0.25\nlstm_efficientnetb0_spleen_weight = 0.25\nlstm_efficientnetv2t_spleen_weight = 0.25\n```\n\nI aggregated scan level predictions on patient_id and took the maximum prediction.\n\nI also scaled injury target predictions with different multipliers and they are also set by minimizing OOF score.\n\n```python\ndf_predictions['bowel_injury_prediction'] *= 1.\ndf_predictions['extravasation_injury_prediction'] *= 1.4\ndf_predictions['kidney_low_prediction'] *= 1.1\ndf_predictions['kidney_high_prediction'] *= 1.1\ndf_predictions['liver_low_prediction'] *= 1.3\ndf_predictions['liver_high_prediction'] *= 1.3\ndf_predictions['spleen_low_prediction'] *= 1.75\ndf_predictions['spleen_high_prediction'] *= 1.75\n```\n\nMy final ensemble score was **0.3859** and target scores are listed below. I really enjoyed how my OOF scores are almost perfectly correlated with LB scores. I selected the submission that had the best OOF, public and private LB scores thanks to stable cross-validation.\n\n| bowel  | extravasation | kidney | liver  | spleen | any    |\n|--------|---------------|--------|--------|--------|--------|\n| 0.1282 | 0.5070        | 0.2831 | 0.4186 | 0.4736 | 0.5050 |\n",
    "2484772": "Pretty impressive solution without any segmentation!",
    "2484550": "Nice work @gunesevitan. Thanks for the write up.",
    "2484384": "Congrats @gunesevitan on this achievement ."
  }
}