{
  "id": 447706,
  "title": "8th Place Solution & Code",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/447706",
  "author_name": "Ian Pan",
  "post_date": "2023-10-16T23:09:34.240000",
  "votes": 28,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Congratulations to all of the winners and competitors. Thank you to Kaggle and the organizers for this interesting competition.</p>\n<p>I joined relatively late after releasing my <a href=\"https://www.kaggle.com/datasets/vaillant/rsna-abdominal-trauma-extravasation-bounding-boxes\" target=\"_blank\">extravasation bounding box labels</a>. My solution was pretty similar to my prior solutions in RSNA cross-sectional imaging challenges (pulmonary embolism, cervical spine fracture). I treated the task as 3 separate subtasks: predicting solid organ injury (liver, kidney, spleen), bowel injury, and extravasation. All models were based on CNN-transformer \"2.5D\" models. </p>\n<h2>Data</h2>\n<p>I converted all DICOMs to 3-channel PNGs, where each channel was a separate CT window. I used 3 windows (soft tissue: WL=50, WW=400, liver: WL=90, WW=150, and angiography: WL=100, WW=700).</p>\n<h2>Crop Model</h2>\n<p>Using 3D connected components, I was able to generate a mask for each CT volume in order to train a model to eliminate the empty black space. This mask was converted into bounding box coordinates for each image, which allowed me to train a 2D CNN mobilenetv3_small_050 model on 256 x 256 images to predict the coordinates. For each CT volume, I took the union of the predicted bounding boxes for each individual slice and used this to crop each individual image.</p>\n<h2>Liver-Kidney-Spleen Organ Identification Model</h2>\n<p>Using the segmentation output from TotalSegmentator, I trained another mobilenetv3_small_050 model on 256 x 256 images to predict presence of these organs on individual slices. </p>\n<h2>Liver-Kidney-Spleen Injury Model</h2>\n<p>With about a week left, I decided to label some slices with injuries to the liver, kidneys, and spleen so I could train a 2D model slice-wise model. I did not release the labels since the end of the competition was near, and I did not want to create any disruption. They are now available <a href=\"https://www.kaggle.com/datasets/vaillant/rsna-abd-trauma-organ-injury-slice-labels\" target=\"_blank\">here</a>.</p>\n<p>I trained a ConvNeXt-tiny model on 2D slice labels with an additional linear layer to reduce the feature dimension to 256. This included the laterality of the kidney injury (i.e., left vs. right), though I am not sure how much this actually helped in the end. The model was trained on cropped images of size 288 x 384 from cropped volumes using the above models. This model was used to extract features from each slice; CT volumes were sampled to 128 images. Thus a CT series was converted to a sequence of shape 128 x 256. </p>\n<p>A 3-layer transformer was trained on these sequences to predict the series-level label. A weighted binary cross-entropy loss was constructed to mimic the competition metric. The validation loss was rather unstable, so I also tracked the AUC to make sure the model was learning.</p>\n<h2>Bowel Injury Model</h2>\n<p>I used the provided slice-wise bowel injury labels to train CNN-transformer model in the same manner as above, except I only cropped the individual images to remove black space, not the CT volume since bowel is present on more slices. Images were resized to 384 x 512. </p>\n<h2>Extravasation Model</h2>\n<p>Using the bounding box labels I annotated, I generated 12 nonoverlapping patches of size 128 x 128 from images of size 384 x 512 and assigned each patch with a label of injury vs. healthy. I trained a ConvNeXt-tiny model on these patches and extracted features for each patch. Thus each image was converted into a sequence of shape 12 x 256. </p>\n<p>I trained a 3-layer transformer using those sequences on slice-wise labels. I then used this transformer to extract features from each image of a resampled 128-slice volume, again resulting in a sequence of shape 128 x 256. This method improved performance over simply training a 2D CNN on whole images. A second-stage transformer was trained on these sequences to predict the series-level labels, using a weighted loss similar to the above.</p>\n<h2>Inference</h2>\n<p>5-fold ensemble of the above was used for final inference. For patient-level prediction, predictions were averaged across the series, if there were 2. Softmax activation function was applied to each label group's predictions, so all the probabilities were already normalized to 1. All probabilities were then scaled by taking the square root, which improved the private LB loss by 0.4. OOF CV was 0.375.</p>\n<h2>Additional Thoughts</h2>\n<p>I tried to incorporate 3D models, but they were taking too long to train and the performance was not as high. I was also interested in training segmentation models and training on cropped organs using the segmentation masks but did not have enough time. I tried training models on images with stacked slices as channels (i.e., each channel of the \"image\" was a separate slice), but this resulted in similar performance (slightly worse on LB). I tried training a single transformer on the concatenation of the features from the 3 types of models above, but this resulted in worse performance. Overall, I am happy to have won my 10th gold medal. </p>\n<p>Inference Notebook: <a href=\"https://www.kaggle.com/code/vaillant/rsna-trauma-submission-v2-1\" target=\"_blank\">https://www.kaggle.com/code/vaillant/rsna-trauma-submission-v2-1</a></p>\n<p>Source Code: <a href=\"https://www.kaggle.com/datasets/vaillant/rsna-trauma-src\" target=\"_blank\">https://www.kaggle.com/datasets/vaillant/rsna-trauma-src</a></p>",
  "messages": [
    {
      "id": 2485076,
      "postDate": "2023-10-16T23:09:34.240Z",
      "content": "<p>Congratulations to all of the winners and competitors. Thank you to Kaggle and the organizers for this interesting competition.</p>\n<p>I joined relatively late after releasing my <a href=\"https://www.kaggle.com/datasets/vaillant/rsna-abdominal-trauma-extravasation-bounding-boxes\" target=\"_blank\">extravasation bounding box labels</a>. My solution was pretty similar to my prior solutions in RSNA cross-sectional imaging challenges (pulmonary embolism, cervical spine fracture). I treated the task as 3 separate subtasks: predicting solid organ injury (liver, kidney, spleen), bowel injury, and extravasation. All models were based on CNN-transformer \"2.5D\" models. </p>\n<h2>Data</h2>\n<p>I converted all DICOMs to 3-channel PNGs, where each channel was a separate CT window. I used 3 windows (soft tissue: WL=50, WW=400, liver: WL=90, WW=150, and angiography: WL=100, WW=700).</p>\n<h2>Crop Model</h2>\n<p>Using 3D connected components, I was able to generate a mask for each CT volume in order to train a model to eliminate the empty black space. This mask was converted into bounding box coordinates for each image, which allowed me to train a 2D CNN mobilenetv3_small_050 model on 256 x 256 images to predict the coordinates. For each CT volume, I took the union of the predicted bounding boxes for each individual slice and used this to crop each individual image.</p>\n<h2>Liver-Kidney-Spleen Organ Identification Model</h2>\n<p>Using the segmentation output from TotalSegmentator, I trained another mobilenetv3_small_050 model on 256 x 256 images to predict presence of these organs on individual slices. </p>\n<h2>Liver-Kidney-Spleen Injury Model</h2>\n<p>With about a week left, I decided to label some slices with injuries to the liver, kidneys, and spleen so I could train a 2D model slice-wise model. I did not release the labels since the end of the competition was near, and I did not want to create any disruption. They are now available <a href=\"https://www.kaggle.com/datasets/vaillant/rsna-abd-trauma-organ-injury-slice-labels\" target=\"_blank\">here</a>.</p>\n<p>I trained a ConvNeXt-tiny model on 2D slice labels with an additional linear layer to reduce the feature dimension to 256. This included the laterality of the kidney injury (i.e., left vs. right), though I am not sure how much this actually helped in the end. The model was trained on cropped images of size 288 x 384 from cropped volumes using the above models. This model was used to extract features from each slice; CT volumes were sampled to 128 images. Thus a CT series was converted to a sequence of shape 128 x 256. </p>\n<p>A 3-layer transformer was trained on these sequences to predict the series-level label. A weighted binary cross-entropy loss was constructed to mimic the competition metric. The validation loss was rather unstable, so I also tracked the AUC to make sure the model was learning.</p>\n<h2>Bowel Injury Model</h2>\n<p>I used the provided slice-wise bowel injury labels to train CNN-transformer model in the same manner as above, except I only cropped the individual images to remove black space, not the CT volume since bowel is present on more slices. Images were resized to 384 x 512. </p>\n<h2>Extravasation Model</h2>\n<p>Using the bounding box labels I annotated, I generated 12 nonoverlapping patches of size 128 x 128 from images of size 384 x 512 and assigned each patch with a label of injury vs. healthy. I trained a ConvNeXt-tiny model on these patches and extracted features for each patch. Thus each image was converted into a sequence of shape 12 x 256. </p>\n<p>I trained a 3-layer transformer using those sequences on slice-wise labels. I then used this transformer to extract features from each image of a resampled 128-slice volume, again resulting in a sequence of shape 128 x 256. This method improved performance over simply training a 2D CNN on whole images. A second-stage transformer was trained on these sequences to predict the series-level labels, using a weighted loss similar to the above.</p>\n<h2>Inference</h2>\n<p>5-fold ensemble of the above was used for final inference. For patient-level prediction, predictions were averaged across the series, if there were 2. Softmax activation function was applied to each label group's predictions, so all the probabilities were already normalized to 1. All probabilities were then scaled by taking the square root, which improved the private LB loss by 0.4. OOF CV was 0.375.</p>\n<h2>Additional Thoughts</h2>\n<p>I tried to incorporate 3D models, but they were taking too long to train and the performance was not as high. I was also interested in training segmentation models and training on cropped organs using the segmentation masks but did not have enough time. I tried training models on images with stacked slices as channels (i.e., each channel of the \"image\" was a separate slice), but this resulted in similar performance (slightly worse on LB). I tried training a single transformer on the concatenation of the features from the 3 types of models above, but this resulted in worse performance. Overall, I am happy to have won my 10th gold medal. </p>\n<p>Inference Notebook: <a href=\"https://www.kaggle.com/code/vaillant/rsna-trauma-submission-v2-1\" target=\"_blank\">https://www.kaggle.com/code/vaillant/rsna-trauma-submission-v2-1</a></p>\n<p>Source Code: <a href=\"https://www.kaggle.com/datasets/vaillant/rsna-trauma-src\" target=\"_blank\">https://www.kaggle.com/datasets/vaillant/rsna-trauma-src</a></p>",
      "rawMarkdown": "Congratulations to all of the winners and competitors. Thank you to Kaggle and the organizers for this interesting competition.\n\nI joined relatively late after releasing my [extravasation bounding box labels](https://www.kaggle.com/datasets/vaillant/rsna-abdominal-trauma-extravasation-bounding-boxes). My solution was pretty similar to my prior solutions in RSNA cross-sectional imaging challenges (pulmonary embolism, cervical spine fracture). I treated the task as 3 separate subtasks: predicting solid organ injury (liver, kidney, spleen), bowel injury, and extravasation. All models were based on CNN-transformer \"2.5D\" models. \n\n## Data \n\nI converted all DICOMs to 3-channel PNGs, where each channel was a separate CT window. I used 3 windows (soft tissue: WL=50, WW=400, liver: WL=90, WW=150, and angiography: WL=100, WW=700).\n\n## Crop Model\n\nUsing 3D connected components, I was able to generate a mask for each CT volume in order to train a model to eliminate the empty black space. This mask was converted into bounding box coordinates for each image, which allowed me to train a 2D CNN mobilenetv3_small_050 model on 256 x 256 images to predict the coordinates. For each CT volume, I took the union of the predicted bounding boxes for each individual slice and used this to crop each individual image.\n\n## Liver-Kidney-Spleen Organ Identification Model\n\nUsing the segmentation output from TotalSegmentator, I trained another mobilenetv3_small_050 model on 256 x 256 images to predict presence of these organs on individual slices. \n\n## Liver-Kidney-Spleen Injury Model\n\nWith about a week left, I decided to label some slices with injuries to the liver, kidneys, and spleen so I could train a 2D model slice-wise model. I did not release the labels since the end of the competition was near, and I did not want to create any disruption. They are now available [here](https://www.kaggle.com/datasets/vaillant/rsna-abd-trauma-organ-injury-slice-labels).\n\nI trained a ConvNeXt-tiny model on 2D slice labels with an additional linear layer to reduce the feature dimension to 256. This included the laterality of the kidney injury (i.e., left vs. right), though I am not sure how much this actually helped in the end. The model was trained on cropped images of size 288 x 384 from cropped volumes using the above models. This model was used to extract features from each slice; CT volumes were sampled to 128 images. Thus a CT series was converted to a sequence of shape 128 x 256. \n\nA 3-layer transformer was trained on these sequences to predict the series-level label. A weighted binary cross-entropy loss was constructed to mimic the competition metric. The validation loss was rather unstable, so I also tracked the AUC to make sure the model was learning.\n\n## Bowel Injury Model\n\nI used the provided slice-wise bowel injury labels to train CNN-transformer model in the same manner as above, except I only cropped the individual images to remove black space, not the CT volume since bowel is present on more slices. Images were resized to 384 x 512. \n\n## Extravasation Model\n\nUsing the bounding box labels I annotated, I generated 12 nonoverlapping patches of size 128 x 128 from images of size 384 x 512 and assigned each patch with a label of injury vs. healthy. I trained a ConvNeXt-tiny model on these patches and extracted features for each patch. Thus each image was converted into a sequence of shape 12 x 256. \n\nI trained a 3-layer transformer using those sequences on slice-wise labels. I then used this transformer to extract features from each image of a resampled 128-slice volume, again resulting in a sequence of shape 128 x 256. This method improved performance over simply training a 2D CNN on whole images. A second-stage transformer was trained on these sequences to predict the series-level labels, using a weighted loss similar to the above.\n\n## Inference\n\n5-fold ensemble of the above was used for final inference. For patient-level prediction, predictions were averaged across the series, if there were 2. Softmax activation function was applied to each label group's predictions, so all the probabilities were already normalized to 1. All probabilities were then scaled by taking the square root, which improved the private LB loss by 0.4. OOF CV was 0.375.\n\n## Additional Thoughts\n\nI tried to incorporate 3D models, but they were taking too long to train and the performance was not as high. I was also interested in training segmentation models and training on cropped organs using the segmentation masks but did not have enough time. I tried training models on images with stacked slices as channels (i.e., each channel of the \"image\" was a separate slice), but this resulted in similar performance (slightly worse on LB). I tried training a single transformer on the concatenation of the features from the 3 types of models above, but this resulted in worse performance. Overall, I am happy to have won my 10th gold medal. \n\nInference Notebook: https://www.kaggle.com/code/vaillant/rsna-trauma-submission-v2-1\n\nSource Code: https://www.kaggle.com/datasets/vaillant/rsna-trauma-src",
      "votes": 28
    },
    {
      "id": 2489981,
      "postDate": "2023-10-20T10:50:49.220Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> that was a speed run</p>",
      "rawMarkdown": "Congratulations @vaillant that was a speed run",
      "votes": 1
    },
    {
      "id": 2485840,
      "postDate": "2023-10-17T13:43:49.147Z",
      "content": "<p>Congratulations. Great job Ian!</p>",
      "rawMarkdown": "Congratulations. Great job Ian!",
      "replies": [
        {
          "id": 2488028,
          "postDate": "2023-10-19T00:44:51.027Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    },
    {
      "id": 2485245,
      "postDate": "2023-10-17T03:54:30.883Z",
      "content": "<p>Congratulations on your 10th gold medal! It was a great pleasure to see you again at the competition.</p>",
      "rawMarkdown": "Congratulations on your 10th gold medal! It was a great pleasure to see you again at the competition.",
      "replies": [
        {
          "id": 2488027,
          "postDate": "2023-10-19T00:44:39.733Z",
          "content": "<p>Thanks! You too!</p>",
          "rawMarkdown": "Thanks! You too!"
        }
      ]
    },
    {
      "id": 2485729,
      "postDate": "2023-10-17T12:16:35.580Z",
      "content": "<p>Nice solution. Congratulations. <br>\n\"I decided to label some slices with injuries to the liver, kidneys, and spleen\". <br>\nWhat tool did you use to avoid making this task time-consuming? Med.ai ?<br>\nThank you for sharing</p>",
      "rawMarkdown": "Nice solution. Congratulations. \n\"I decided to label some slices with injuries to the liver, kidneys, and spleen\". \nWhat tool did you use to avoid making this task time-consuming? Med.ai ?\nThank you for sharing",
      "isDeleted": true,
      "replies": [
        {
          "id": 2488026,
          "postDate": "2023-10-19T00:44:27.057Z",
          "content": "<p>I just saved the PNGs of each series to a folder and deleted the ones that I wanted to label as positive injury. Then I could just see which images were deleted and mark them as positive.</p>",
          "rawMarkdown": "I just saved the PNGs of each series to a folder and deleted the ones that I wanted to label as positive injury. Then I could just see which images were deleted and mark them as positive."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2489981,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2023-10-20T10:50:49.220000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> that was a speed run</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2485840,
      "author_name": "Yee Ng",
      "author_url": "",
      "post_date": "2023-10-17T13:43:49.147000",
      "content": "<p>Congratulations. Great job Ian!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2488028,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2023-10-19T00:44:51.027000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2485245,
      "author_name": "YujiAriyasu",
      "author_url": "",
      "post_date": "2023-10-17T03:54:30.883000",
      "content": "<p>Congratulations on your 10th gold medal! It was a great pleasure to see you again at the competition.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2488027,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2023-10-19T00:44:39.733000",
          "content": "<p>Thanks! You too!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2485729,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-17T12:16:35.580000",
      "content": "<p>Nice solution. Congratulations. <br>\n\"I decided to label some slices with injuries to the liver, kidneys, and spleen\". <br>\nWhat tool did you use to avoid making this task time-consuming? Med.ai ?<br>\nThank you for sharing</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2488026,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2023-10-19T00:44:27.057000",
          "content": "<p>I just saved the PNGs of each series to a folder and deleted the ones that I wanted to label as positive injury. Then I could just see which images were deleted and mark them as positive.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2485076": "Congratulations to all of the winners and competitors. Thank you to Kaggle and the organizers for this interesting competition.\n\nI joined relatively late after releasing my [extravasation bounding box labels](https://www.kaggle.com/datasets/vaillant/rsna-abdominal-trauma-extravasation-bounding-boxes). My solution was pretty similar to my prior solutions in RSNA cross-sectional imaging challenges (pulmonary embolism, cervical spine fracture). I treated the task as 3 separate subtasks: predicting solid organ injury (liver, kidney, spleen), bowel injury, and extravasation. All models were based on CNN-transformer \"2.5D\" models. \n\n## Data \n\nI converted all DICOMs to 3-channel PNGs, where each channel was a separate CT window. I used 3 windows (soft tissue: WL=50, WW=400, liver: WL=90, WW=150, and angiography: WL=100, WW=700).\n\n## Crop Model\n\nUsing 3D connected components, I was able to generate a mask for each CT volume in order to train a model to eliminate the empty black space. This mask was converted into bounding box coordinates for each image, which allowed me to train a 2D CNN mobilenetv3_small_050 model on 256 x 256 images to predict the coordinates. For each CT volume, I took the union of the predicted bounding boxes for each individual slice and used this to crop each individual image.\n\n## Liver-Kidney-Spleen Organ Identification Model\n\nUsing the segmentation output from TotalSegmentator, I trained another mobilenetv3_small_050 model on 256 x 256 images to predict presence of these organs on individual slices. \n\n## Liver-Kidney-Spleen Injury Model\n\nWith about a week left, I decided to label some slices with injuries to the liver, kidneys, and spleen so I could train a 2D model slice-wise model. I did not release the labels since the end of the competition was near, and I did not want to create any disruption. They are now available [here](https://www.kaggle.com/datasets/vaillant/rsna-abd-trauma-organ-injury-slice-labels).\n\nI trained a ConvNeXt-tiny model on 2D slice labels with an additional linear layer to reduce the feature dimension to 256. This included the laterality of the kidney injury (i.e., left vs. right), though I am not sure how much this actually helped in the end. The model was trained on cropped images of size 288 x 384 from cropped volumes using the above models. This model was used to extract features from each slice; CT volumes were sampled to 128 images. Thus a CT series was converted to a sequence of shape 128 x 256. \n\nA 3-layer transformer was trained on these sequences to predict the series-level label. A weighted binary cross-entropy loss was constructed to mimic the competition metric. The validation loss was rather unstable, so I also tracked the AUC to make sure the model was learning.\n\n## Bowel Injury Model\n\nI used the provided slice-wise bowel injury labels to train CNN-transformer model in the same manner as above, except I only cropped the individual images to remove black space, not the CT volume since bowel is present on more slices. Images were resized to 384 x 512. \n\n## Extravasation Model\n\nUsing the bounding box labels I annotated, I generated 12 nonoverlapping patches of size 128 x 128 from images of size 384 x 512 and assigned each patch with a label of injury vs. healthy. I trained a ConvNeXt-tiny model on these patches and extracted features for each patch. Thus each image was converted into a sequence of shape 12 x 256. \n\nI trained a 3-layer transformer using those sequences on slice-wise labels. I then used this transformer to extract features from each image of a resampled 128-slice volume, again resulting in a sequence of shape 128 x 256. This method improved performance over simply training a 2D CNN on whole images. A second-stage transformer was trained on these sequences to predict the series-level labels, using a weighted loss similar to the above.\n\n## Inference\n\n5-fold ensemble of the above was used for final inference. For patient-level prediction, predictions were averaged across the series, if there were 2. Softmax activation function was applied to each label group's predictions, so all the probabilities were already normalized to 1. All probabilities were then scaled by taking the square root, which improved the private LB loss by 0.4. OOF CV was 0.375.\n\n## Additional Thoughts\n\nI tried to incorporate 3D models, but they were taking too long to train and the performance was not as high. I was also interested in training segmentation models and training on cropped organs using the segmentation masks but did not have enough time. I tried training models on images with stacked slices as channels (i.e., each channel of the \"image\" was a separate slice), but this resulted in similar performance (slightly worse on LB). I tried training a single transformer on the concatenation of the features from the 3 types of models above, but this resulted in worse performance. Overall, I am happy to have won my 10th gold medal. \n\nInference Notebook: https://www.kaggle.com/code/vaillant/rsna-trauma-submission-v2-1\n\nSource Code: https://www.kaggle.com/datasets/vaillant/rsna-trauma-src",
    "2489981": "Congratulations @vaillant that was a speed run",
    "2485840": "Congratulations. Great job Ian!",
    "2485245": "Congratulations on your 10th gold medal! It was a great pleasure to see you again at the competition.",
    "2485729": "Nice solution. Congratulations. \n\"I decided to label some slices with injuries to the liver, kidneys, and spleen\". \nWhat tool did you use to avoid making this task time-consuming? Med.ai ?\nThank you for sharing"
  }
}