{
  "id": 363232,
  "title": "5th place solution (Team Speedrun)",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/363232",
  "author_name": "Dieter",
  "post_date": "2022-10-31T16:11:44.723000",
  "votes": 51,
  "comment_count": 7,
  "views": 0,
  "content": "<h1>5th place solution</h1>\n<p>Thanks to kaggle and the sponsors for organizing this interesting and leak-free competition which had a lot of different angles to explore. Coming directly from the DFL and joining quite late here, we were explicitly looking for the sprint aspect and to challenge ourselves to derive a good solution in only 11 days. Hence our team name: Speedrun ( <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>, <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> ) we could build on ideas shared from previous 3D-RSNA competitions as well as competitions we did before and we are proud of achieving 5th rank in only this short timespan. We also want to thank authors of various github repositories which enable us to quickly iterate and try ideas, such as <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">timm</a>, <a href=\"https://github.com/albumentations-team/albumentations\" target=\"_blank\">albumentations</a> and <a href=\"https://github.com/qubvel/segmentation_models.pytorch\" target=\"_blank\">segmentation models pytorch</a></p>\n<h2>TL;DR</h2>\n<p>Our solution is an ensemble of two quite similar approaches which follow a 3-stage paradigm:</p>\n<ol>\n<li>Train a 2D classification and segmentation model for vertebrae using the provided segmentation labels</li>\n<li>Use the resulting model to predict the vertebrae class visible in each all dicom images and multiply the result with the overall fracture label to get 2D image level labels if a certain vertebrae is fractured or not.</li>\n<li>Collect all image level labels per study and train an aggregation model to predict the given study level labels. </li>\n</ol>\n<p>The following figure gives a summary of our solution:<br>\n<img src=\"https://i.imgur.com/2Q1RU35.jpeg\" alt=\"\"></p>\n<h2>Stage 1</h2>\n<p>The first stage of our approach uses only the 87 studies for which we were given 3D-segmentation masks. We used the 3D-segmentation labels to derive two type of labels for a given 2D slice. On the one hand we created 7 dimensional binary labels for each 2D slice if a certain vertebrae is visible or not and trained an EfficientNet-B5. (model S1A). On the other hand we use a vertebrae segmentation mask for each 2D slice to train a EfficientNet-B3-UNet. (model S1B)</p>\n<p>Using model S1A we predict if the vertebrae are visible for all 2000 studys in the train set using a threshold of 0.5. We then map the target labels from study level to 2D by multiplying the prediction with the overall study labels if a vertebrae is fractured. The so created pseudo labels will be used in Stage 2. Additionally S1B is used to predict segmentation mask if any vertebrae is visible and derive a bounding box (x_min,x_max,y_min,y_max) out of it for each 2D slice. We aggregate all bounding boxes for per study using the 0.05 quantile of x_min and y_min and the 0.95 quantile for x_max and y_max to get a single box per study which will be used in stage2 for cropping the region of interest.</p>\n<h2>Stage 2</h2>\n<p>For our stage 2 models, we directly aim at predicting the probability for a fracture at each of the seven cervical vertebrae. The labels for each slice have been built by our stage 1 models.</p>\n<p>All our models in the second stage follow a 2.5D + 3D schema. The input is a 2.5D slice with 3 channels, where the center channel is the z-dimension of interest, and the two neighboring slices form the other channels. This is then run through a 2d backend, and the last layer(s) of the backbone are transformed to 3D Convolution layers and are pooled with average across the slices. For the 3D part we take a step-size of 5 and also only take two extra slices in each direction. </p>\n<p>So for a sample input frame with index 15, we would first stack the channels 14,15,16, and then add the frames at position -5, and +5, which also have 3 channels stacked, so we add the stacked channels 9,10,11 and 19,20,21. The input dimension is then (3,3,height,width).</p>\n<p>We train two types of stage 2 models:</p>\n<ul>\n<li>Full images: The input here is the full regular extracted images from the DICOM. </li>\n<li>ROI cropped images: The input here is the extracted ROI from a first stage model, see above description for details.</li>\n</ul>\n<p>It is very easy to overfit this data, so we apply heavy regularization, some of the most import ones:</p>\n<ul>\n<li>Mixup</li>\n<li>Random crop + resize</li>\n<li>Random shift, scale, rotate</li>\n<li>Random brightness</li>\n</ul>\n<p>All our models are very lightweight efficientnetv2 models. For inference, it was very helpful for us to upscale the images by at least 1.125 and sometimes also apply a center crop.</p>\n<h2>Stage 3</h2>\n<p>As our second stage model is trained on individual slices, our models cannot fully learn the actual level of interest, which is on a study level. So we train a simple 3rd stage feed-forward neural network, that directly optimizes the competition metric, and has as input only the mean, min and maximum of all individual predictions of a study. This helps specifically to boost the overall prediction score, which we cannot optimize directly also in the second stage.</p>\n<p>The final sub is a 30-70 blend of the max-aggregated output from 2nd stage and 3rd stage models and improves our scores by 2-3 LB points.</p>\n<p>edit: </p>\n<ul>\n<li>We uploaded a video explaining our solution <a href=\"https://www.youtube.com/watch?v=c_YZHwhK0Jo\" target=\"_blank\">here</a></li>\n<li>The inference kernel is shared <a href=\"https://www.kaggle.com/code/ilu000/rsna2022-5th-place-solution-inference\" target=\"_blank\">here</a></li>\n<li>Github repository with training code and instructions can be found <a href=\"https://github.com/pascal-pfeiffer/kaggle-rsna-2022-5th-place\" target=\"_blank\">here</a></li>\n</ul>",
  "messages": [
    {
      "id": 2011522,
      "postDate": "2022-10-31T16:11:44.723Z",
      "content": "<h1>5th place solution</h1>\n<p>Thanks to kaggle and the sponsors for organizing this interesting and leak-free competition which had a lot of different angles to explore. Coming directly from the DFL and joining quite late here, we were explicitly looking for the sprint aspect and to challenge ourselves to derive a good solution in only 11 days. Hence our team name: Speedrun ( <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>, <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> ) we could build on ideas shared from previous 3D-RSNA competitions as well as competitions we did before and we are proud of achieving 5th rank in only this short timespan. We also want to thank authors of various github repositories which enable us to quickly iterate and try ideas, such as <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">timm</a>, <a href=\"https://github.com/albumentations-team/albumentations\" target=\"_blank\">albumentations</a> and <a href=\"https://github.com/qubvel/segmentation_models.pytorch\" target=\"_blank\">segmentation models pytorch</a></p>\n<h2>TL;DR</h2>\n<p>Our solution is an ensemble of two quite similar approaches which follow a 3-stage paradigm:</p>\n<ol>\n<li>Train a 2D classification and segmentation model for vertebrae using the provided segmentation labels</li>\n<li>Use the resulting model to predict the vertebrae class visible in each all dicom images and multiply the result with the overall fracture label to get 2D image level labels if a certain vertebrae is fractured or not.</li>\n<li>Collect all image level labels per study and train an aggregation model to predict the given study level labels. </li>\n</ol>\n<p>The following figure gives a summary of our solution:<br>\n<img src=\"https://i.imgur.com/2Q1RU35.jpeg\" alt=\"\"></p>\n<h2>Stage 1</h2>\n<p>The first stage of our approach uses only the 87 studies for which we were given 3D-segmentation masks. We used the 3D-segmentation labels to derive two type of labels for a given 2D slice. On the one hand we created 7 dimensional binary labels for each 2D slice if a certain vertebrae is visible or not and trained an EfficientNet-B5. (model S1A). On the other hand we use a vertebrae segmentation mask for each 2D slice to train a EfficientNet-B3-UNet. (model S1B)</p>\n<p>Using model S1A we predict if the vertebrae are visible for all 2000 studys in the train set using a threshold of 0.5. We then map the target labels from study level to 2D by multiplying the prediction with the overall study labels if a vertebrae is fractured. The so created pseudo labels will be used in Stage 2. Additionally S1B is used to predict segmentation mask if any vertebrae is visible and derive a bounding box (x_min,x_max,y_min,y_max) out of it for each 2D slice. We aggregate all bounding boxes for per study using the 0.05 quantile of x_min and y_min and the 0.95 quantile for x_max and y_max to get a single box per study which will be used in stage2 for cropping the region of interest.</p>\n<h2>Stage 2</h2>\n<p>For our stage 2 models, we directly aim at predicting the probability for a fracture at each of the seven cervical vertebrae. The labels for each slice have been built by our stage 1 models.</p>\n<p>All our models in the second stage follow a 2.5D + 3D schema. The input is a 2.5D slice with 3 channels, where the center channel is the z-dimension of interest, and the two neighboring slices form the other channels. This is then run through a 2d backend, and the last layer(s) of the backbone are transformed to 3D Convolution layers and are pooled with average across the slices. For the 3D part we take a step-size of 5 and also only take two extra slices in each direction. </p>\n<p>So for a sample input frame with index 15, we would first stack the channels 14,15,16, and then add the frames at position -5, and +5, which also have 3 channels stacked, so we add the stacked channels 9,10,11 and 19,20,21. The input dimension is then (3,3,height,width).</p>\n<p>We train two types of stage 2 models:</p>\n<ul>\n<li>Full images: The input here is the full regular extracted images from the DICOM. </li>\n<li>ROI cropped images: The input here is the extracted ROI from a first stage model, see above description for details.</li>\n</ul>\n<p>It is very easy to overfit this data, so we apply heavy regularization, some of the most import ones:</p>\n<ul>\n<li>Mixup</li>\n<li>Random crop + resize</li>\n<li>Random shift, scale, rotate</li>\n<li>Random brightness</li>\n</ul>\n<p>All our models are very lightweight efficientnetv2 models. For inference, it was very helpful for us to upscale the images by at least 1.125 and sometimes also apply a center crop.</p>\n<h2>Stage 3</h2>\n<p>As our second stage model is trained on individual slices, our models cannot fully learn the actual level of interest, which is on a study level. So we train a simple 3rd stage feed-forward neural network, that directly optimizes the competition metric, and has as input only the mean, min and maximum of all individual predictions of a study. This helps specifically to boost the overall prediction score, which we cannot optimize directly also in the second stage.</p>\n<p>The final sub is a 30-70 blend of the max-aggregated output from 2nd stage and 3rd stage models and improves our scores by 2-3 LB points.</p>\n<p>edit: </p>\n<ul>\n<li>We uploaded a video explaining our solution <a href=\"https://www.youtube.com/watch?v=c_YZHwhK0Jo\" target=\"_blank\">here</a></li>\n<li>The inference kernel is shared <a href=\"https://www.kaggle.com/code/ilu000/rsna2022-5th-place-solution-inference\" target=\"_blank\">here</a></li>\n<li>Github repository with training code and instructions can be found <a href=\"https://github.com/pascal-pfeiffer/kaggle-rsna-2022-5th-place\" target=\"_blank\">here</a></li>\n</ul>",
      "rawMarkdown": "# 5th place solution\n\nThanks to kaggle and the sponsors for organizing this interesting and leak-free competition which had a lot of different angles to explore. Coming directly from the DFL and joining quite late here, we were explicitly looking for the sprint aspect and to challenge ourselves to derive a good solution in only 11 days. Hence our team name: Speedrun ( @philippsinger, @ilu000 @christofhenkel ) we could build on ideas shared from previous 3D-RSNA competitions as well as competitions we did before and we are proud of achieving 5th rank in only this short timespan. We also want to thank authors of various github repositories which enable us to quickly iterate and try ideas, such as [timm](https://github.com/rwightman/pytorch-image-models), [albumentations](https://github.com/albumentations-team/albumentations) and [segmentation models pytorch](https://github.com/qubvel/segmentation_models.pytorch)\n\n## TL;DR\n\nOur solution is an ensemble of two quite similar approaches which follow a 3-stage paradigm:\n\n1. Train a 2D classification and segmentation model for vertebrae using the provided segmentation labels\n2. Use the resulting model to predict the vertebrae class visible in each all dicom images and multiply the result with the overall fracture label to get 2D image level labels if a certain vertebrae is fractured or not.\n3. Collect all image level labels per study and train an aggregation model to predict the given study level labels. \n\nThe following figure gives a summary of our solution:\n![](https://i.imgur.com/2Q1RU35.jpeg)\n\n## Stage 1\n\nThe first stage of our approach uses only the 87 studies for which we were given 3D-segmentation masks. We used the 3D-segmentation labels to derive two type of labels for a given 2D slice. On the one hand we created 7 dimensional binary labels for each 2D slice if a certain vertebrae is visible or not and trained an EfficientNet-B5. (model S1A). On the other hand we use a vertebrae segmentation mask for each 2D slice to train a EfficientNet-B3-UNet. (model S1B)\n\nUsing model S1A we predict if the vertebrae are visible for all 2000 studys in the train set using a threshold of 0.5. We then map the target labels from study level to 2D by multiplying the prediction with the overall study labels if a vertebrae is fractured. The so created pseudo labels will be used in Stage 2. Additionally S1B is used to predict segmentation mask if any vertebrae is visible and derive a bounding box (x_min,x_max,y_min,y_max) out of it for each 2D slice. We aggregate all bounding boxes for per study using the 0.05 quantile of x_min and y_min and the 0.95 quantile for x_max and y_max to get a single box per study which will be used in stage2 for cropping the region of interest.\n\n## Stage 2\n\nFor our stage 2 models, we directly aim at predicting the probability for a fracture at each of the seven cervical vertebrae. The labels for each slice have been built by our stage 1 models.\n\nAll our models in the second stage follow a 2.5D + 3D schema. The input is a 2.5D slice with 3 channels, where the center channel is the z-dimension of interest, and the two neighboring slices form the other channels. This is then run through a 2d backend, and the last layer(s) of the backbone are transformed to 3D Convolution layers and are pooled with average across the slices. For the 3D part we take a step-size of 5 and also only take two extra slices in each direction. \n\nSo for a sample input frame with index 15, we would first stack the channels 14,15,16, and then add the frames at position -5, and +5, which also have 3 channels stacked, so we add the stacked channels 9,10,11 and 19,20,21. The input dimension is then (3,3,height,width).\n\nWe train two types of stage 2 models:\n- Full images: The input here is the full regular extracted images from the DICOM. \n- ROI cropped images: The input here is the extracted ROI from a first stage model, see above description for details.\n\nIt is very easy to overfit this data, so we apply heavy regularization, some of the most import ones:\n- Mixup\n- Random crop + resize\n- Random shift, scale, rotate\n- Random brightness\n\nAll our models are very lightweight efficientnetv2 models. For inference, it was very helpful for us to upscale the images by at least 1.125 and sometimes also apply a center crop.\n\n## Stage 3\nAs our second stage model is trained on individual slices, our models cannot fully learn the actual level of interest, which is on a study level. So we train a simple 3rd stage feed-forward neural network, that directly optimizes the competition metric, and has as input only the mean, min and maximum of all individual predictions of a study. This helps specifically to boost the overall prediction score, which we cannot optimize directly also in the second stage.\n\nThe final sub is a 30-70 blend of the max-aggregated output from 2nd stage and 3rd stage models and improves our scores by 2-3 LB points.\n\nedit: \n- We uploaded a video explaining our solution [here](https://www.youtube.com/watch?v=c_YZHwhK0Jo)\n- The inference kernel is shared [here](https://www.kaggle.com/code/ilu000/rsna2022-5th-place-solution-inference)\n- Github repository with training code and instructions can be found [here](https://github.com/pascal-pfeiffer/kaggle-rsna-2022-5th-place)",
      "votes": 51
    },
    {
      "id": 2022007,
      "postDate": "2022-11-08T16:01:31.667Z",
      "content": "<p>edit: </p>\n<ul>\n<li>We uploaded a video explaining our solution <a href=\"https://www.youtube.com/watch?v=c_YZHwhK0Jo\" target=\"_blank\">here</a></li>\n<li>The inference kernel is shared <a href=\"https://www.kaggle.com/code/ilu000/rsna2022-5th-place-solution-inference\" target=\"_blank\">here</a></li>\n<li>Github repository with training code and instructions can be found <a href=\"https://github.com/pascal-pfeiffer/kaggle-rsna-2022-5th-place\" target=\"_blank\">here</a></li>\n</ul>",
      "rawMarkdown": "edit: \n- We uploaded a video explaining our solution [here](https://www.youtube.com/watch?v=c_YZHwhK0Jo)\n- The inference kernel is shared [here](https://www.kaggle.com/code/ilu000/rsna2022-5th-place-solution-inference)\n- Github repository with training code and instructions can be found [here](https://github.com/pascal-pfeiffer/kaggle-rsna-2022-5th-place)",
      "votes": 5
    },
    {
      "id": 2011712,
      "postDate": "2022-10-31T18:12:53.297Z",
      "content": "<p>Congratulations on your speedy gold! It is visible from the solution that it was not only a really good approach but also an excellent execution to develop it in only 11 days!</p>",
      "rawMarkdown": "Congratulations on your speedy gold! It is visible from the solution that it was not only a really good approach but also an excellent execution to develop it in only 11 days!",
      "votes": 4
    },
    {
      "id": 2120969,
      "postDate": "2023-01-29T23:36:39.990Z",
      "content": "<p>Hey, I was trying to re-run the code and it's missing the <code>meta_segmentation.csv</code> file. Not sure where it's coming from either. </p>\n<p><a href=\"https://github.com/pascal-pfeiffer/kaggle-rsna-2022-5th-place/blob/main/scripts/get_meta_wirbel_dcm_v1_and_v2.py#L13\" target=\"_blank\">https://github.com/pascal-pfeiffer/kaggle-rsna-2022-5th-place/blob/main/scripts/get_meta_wirbel_dcm_v1_and_v2.py#L13</a></p>",
      "rawMarkdown": "Hey, I was trying to re-run the code and it's missing the `meta_segmentation.csv` file. Not sure where it's coming from either. \n\nhttps://github.com/pascal-pfeiffer/kaggle-rsna-2022-5th-place/blob/main/scripts/get_meta_wirbel_dcm_v1_and_v2.py#L13",
      "replies": [
        {
          "id": 2120972,
          "postDate": "2023-01-29T23:44:04.163Z",
          "content": "<p>Sorry, my bad - file is here <a href=\"https://www.kaggle.com/datasets/samuelcortinhas/rsna-2022-spine-fracture-detection-metadata\" target=\"_blank\">https://www.kaggle.com/datasets/samuelcortinhas/rsna-2022-spine-fracture-detection-metadata</a></p>",
          "rawMarkdown": "Sorry, my bad - file is here https://www.kaggle.com/datasets/samuelcortinhas/rsna-2022-spine-fracture-detection-metadata"
        }
      ]
    },
    {
      "id": 2023732,
      "postDate": "2022-11-10T01:47:24.733Z",
      "content": "<p>Thanks for sharing with us. Awesome learning here.</p>",
      "rawMarkdown": "Thanks for sharing with us. Awesome learning here."
    },
    {
      "id": 2011543,
      "postDate": "2022-10-31T16:24:36.603Z",
      "content": "<p>Tons to learn here, thanks for the write up!</p>",
      "rawMarkdown": "Tons to learn here, thanks for the write up!"
    },
    {
      "id": 2011531,
      "postDate": "2022-10-31T16:20:30.887Z",
      "content": "<p>Wonderful approach note <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>! Hearty congratulations for the result and another gold medal too!</p>",
      "rawMarkdown": "Wonderful approach note @christofhenkel! Hearty congratulations for the result and another gold medal too!"
    }
  ],
  "comments": [
    {
      "id": 2022007,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2022-11-08T16:01:31.667000",
      "content": "<p>edit: </p>\n<ul>\n<li>We uploaded a video explaining our solution <a href=\"https://www.youtube.com/watch?v=c_YZHwhK0Jo\" target=\"_blank\">here</a></li>\n<li>The inference kernel is shared <a href=\"https://www.kaggle.com/code/ilu000/rsna2022-5th-place-solution-inference\" target=\"_blank\">here</a></li>\n<li>Github repository with training code and instructions can be found <a href=\"https://github.com/pascal-pfeiffer/kaggle-rsna-2022-5th-place\" target=\"_blank\">here</a></li>\n</ul>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2011712,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2022-10-31T18:12:53.297000",
      "content": "<p>Congratulations on your speedy gold! It is visible from the solution that it was not only a really good approach but also an excellent execution to develop it in only 11 days!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2120969,
      "author_name": "Aman Arora",
      "author_url": "",
      "post_date": "2023-01-29T23:36:39.990000",
      "content": "<p>Hey, I was trying to re-run the code and it's missing the <code>meta_segmentation.csv</code> file. Not sure where it's coming from either. </p>\n<p><a href=\"https://github.com/pascal-pfeiffer/kaggle-rsna-2022-5th-place/blob/main/scripts/get_meta_wirbel_dcm_v1_and_v2.py#L13\" target=\"_blank\">https://github.com/pascal-pfeiffer/kaggle-rsna-2022-5th-place/blob/main/scripts/get_meta_wirbel_dcm_v1_and_v2.py#L13</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2120972,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2023-01-29T23:44:04.163000",
          "content": "<p>Sorry, my bad - file is here <a href=\"https://www.kaggle.com/datasets/samuelcortinhas/rsna-2022-spine-fracture-detection-metadata\" target=\"_blank\">https://www.kaggle.com/datasets/samuelcortinhas/rsna-2022-spine-fracture-detection-metadata</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2023732,
      "author_name": "Mahyar Arani",
      "author_url": "",
      "post_date": "2022-11-10T01:47:24.733000",
      "content": "<p>Thanks for sharing with us. Awesome learning here.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2011543,
      "author_name": "aspiring",
      "author_url": "",
      "post_date": "2022-10-31T16:24:36.603000",
      "content": "<p>Tons to learn here, thanks for the write up!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2011531,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2022-10-31T16:20:30.887000",
      "content": "<p>Wonderful approach note <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>! Hearty congratulations for the result and another gold medal too!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2011522": "# 5th place solution\n\nThanks to kaggle and the sponsors for organizing this interesting and leak-free competition which had a lot of different angles to explore. Coming directly from the DFL and joining quite late here, we were explicitly looking for the sprint aspect and to challenge ourselves to derive a good solution in only 11 days. Hence our team name: Speedrun ( @philippsinger, @ilu000 @christofhenkel ) we could build on ideas shared from previous 3D-RSNA competitions as well as competitions we did before and we are proud of achieving 5th rank in only this short timespan. We also want to thank authors of various github repositories which enable us to quickly iterate and try ideas, such as [timm](https://github.com/rwightman/pytorch-image-models), [albumentations](https://github.com/albumentations-team/albumentations) and [segmentation models pytorch](https://github.com/qubvel/segmentation_models.pytorch)\n\n## TL;DR\n\nOur solution is an ensemble of two quite similar approaches which follow a 3-stage paradigm:\n\n1. Train a 2D classification and segmentation model for vertebrae using the provided segmentation labels\n2. Use the resulting model to predict the vertebrae class visible in each all dicom images and multiply the result with the overall fracture label to get 2D image level labels if a certain vertebrae is fractured or not.\n3. Collect all image level labels per study and train an aggregation model to predict the given study level labels. \n\nThe following figure gives a summary of our solution:\n![](https://i.imgur.com/2Q1RU35.jpeg)\n\n## Stage 1\n\nThe first stage of our approach uses only the 87 studies for which we were given 3D-segmentation masks. We used the 3D-segmentation labels to derive two type of labels for a given 2D slice. On the one hand we created 7 dimensional binary labels for each 2D slice if a certain vertebrae is visible or not and trained an EfficientNet-B5. (model S1A). On the other hand we use a vertebrae segmentation mask for each 2D slice to train a EfficientNet-B3-UNet. (model S1B)\n\nUsing model S1A we predict if the vertebrae are visible for all 2000 studys in the train set using a threshold of 0.5. We then map the target labels from study level to 2D by multiplying the prediction with the overall study labels if a vertebrae is fractured. The so created pseudo labels will be used in Stage 2. Additionally S1B is used to predict segmentation mask if any vertebrae is visible and derive a bounding box (x_min,x_max,y_min,y_max) out of it for each 2D slice. We aggregate all bounding boxes for per study using the 0.05 quantile of x_min and y_min and the 0.95 quantile for x_max and y_max to get a single box per study which will be used in stage2 for cropping the region of interest.\n\n## Stage 2\n\nFor our stage 2 models, we directly aim at predicting the probability for a fracture at each of the seven cervical vertebrae. The labels for each slice have been built by our stage 1 models.\n\nAll our models in the second stage follow a 2.5D + 3D schema. The input is a 2.5D slice with 3 channels, where the center channel is the z-dimension of interest, and the two neighboring slices form the other channels. This is then run through a 2d backend, and the last layer(s) of the backbone are transformed to 3D Convolution layers and are pooled with average across the slices. For the 3D part we take a step-size of 5 and also only take two extra slices in each direction. \n\nSo for a sample input frame with index 15, we would first stack the channels 14,15,16, and then add the frames at position -5, and +5, which also have 3 channels stacked, so we add the stacked channels 9,10,11 and 19,20,21. The input dimension is then (3,3,height,width).\n\nWe train two types of stage 2 models:\n- Full images: The input here is the full regular extracted images from the DICOM. \n- ROI cropped images: The input here is the extracted ROI from a first stage model, see above description for details.\n\nIt is very easy to overfit this data, so we apply heavy regularization, some of the most import ones:\n- Mixup\n- Random crop + resize\n- Random shift, scale, rotate\n- Random brightness\n\nAll our models are very lightweight efficientnetv2 models. For inference, it was very helpful for us to upscale the images by at least 1.125 and sometimes also apply a center crop.\n\n## Stage 3\nAs our second stage model is trained on individual slices, our models cannot fully learn the actual level of interest, which is on a study level. So we train a simple 3rd stage feed-forward neural network, that directly optimizes the competition metric, and has as input only the mean, min and maximum of all individual predictions of a study. This helps specifically to boost the overall prediction score, which we cannot optimize directly also in the second stage.\n\nThe final sub is a 30-70 blend of the max-aggregated output from 2nd stage and 3rd stage models and improves our scores by 2-3 LB points.\n\nedit: \n- We uploaded a video explaining our solution [here](https://www.youtube.com/watch?v=c_YZHwhK0Jo)\n- The inference kernel is shared [here](https://www.kaggle.com/code/ilu000/rsna2022-5th-place-solution-inference)\n- Github repository with training code and instructions can be found [here](https://github.com/pascal-pfeiffer/kaggle-rsna-2022-5th-place)",
    "2022007": "edit: \n- We uploaded a video explaining our solution [here](https://www.youtube.com/watch?v=c_YZHwhK0Jo)\n- The inference kernel is shared [here](https://www.kaggle.com/code/ilu000/rsna2022-5th-place-solution-inference)\n- Github repository with training code and instructions can be found [here](https://github.com/pascal-pfeiffer/kaggle-rsna-2022-5th-place)",
    "2011712": "Congratulations on your speedy gold! It is visible from the solution that it was not only a really good approach but also an excellent execution to develop it in only 11 days!",
    "2120969": "Hey, I was trying to re-run the code and it's missing the `meta_segmentation.csv` file. Not sure where it's coming from either. \n\nhttps://github.com/pascal-pfeiffer/kaggle-rsna-2022-5th-place/blob/main/scripts/get_meta_wirbel_dcm_v1_and_v2.py#L13",
    "2023732": "Thanks for sharing with us. Awesome learning here.",
    "2011543": "Tons to learn here, thanks for the write up!",
    "2011531": "Wonderful approach note @christofhenkel! Hearty congratulations for the result and another gold medal too!"
  }
}