{
  "id": 447848,
  "title": "4th Place Solution",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/447848",
  "author_name": "sheep",
  "post_date": "2023-10-17T13:18:36.651000",
  "votes": 21,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thank you to RSNA and Kaggle for hosting this competition. Congratulations to all the winners and participants.</p>\n<h2>Overview</h2>\n<p>I employed a 2.5 pipeline and trained for both classification and segmentation tasks.</p>\n<h2>Dataset</h2>\n<p>I utilized 3D masks from TotalSegmentator and subsequently retrained a 2D model for the liver, spleen, bowel, kidney, and body. The DICOM images were rescaled to (1, 1, 5) and stored with a 5-channel mask.<br>\nFor each epoch, I first sampled N=14 frames from the rescaled array, then used the body mask to filter out hands or other irrelevant areas.<br>\nI used the organ mask to limit the Z-axis space, as slices without the target organ might contain less valuable information.</p>\n<h2>Model</h2>\n<p>I used a Unet model integrated with Pyramid Vision Transformer V2 and MaxViT encoder. Transformers outperformed the convolutional models, especially for the extravasation target.</p>\n<p>Pretraining on the 2D mask facilitated convergence.<br>\n6 classification heads were employed for prediction.<br>\nPerhaps the lack of an RNN layer is the primary reason I didn't match the performance of the top teams.</p>\n<h2>Loss</h2>\n<p>I used the CE loss with weights identical to the metric.</p>\n<h2>Results</h2>\n<table>\n<thead>\n<tr>\n<th>Encoder</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>pvt-b2</td>\n<td>0.3783</td>\n<td>0.41</td>\n</tr>\n<tr>\n<td>pvt-b3</td>\n<td>0.3750</td>\n<td>0.41</td>\n</tr>\n<tr>\n<td>pvt-b4</td>\n<td>0.3786</td>\n<td>0.4</td>\n</tr>\n<tr>\n<td>maxvit_t</td>\n<td>0.3810</td>\n<td>0.42</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>0.3570</td>\n<td>0.4</td>\n</tr>\n<tr>\n<td>ensemble w/ scale</td>\n<td>0.3530</td>\n<td>0.39</td>\n</tr>\n</tbody>\n</table>\n<h2>Code</h2>\n<p><a href=\"https://github.com/iseekwonderful/RSNA-2023-Abdominal-Trauma-Detection-4th-Place-Code.git\" target=\"_blank\">Github link</a></p>",
  "messages": [
    {
      "id": 2485806,
      "postDate": "2023-10-17T13:18:36.653Z",
      "content": "<p>Thank you to RSNA and Kaggle for hosting this competition. Congratulations to all the winners and participants.</p>\n<h2>Overview</h2>\n<p>I employed a 2.5 pipeline and trained for both classification and segmentation tasks.</p>\n<h2>Dataset</h2>\n<p>I utilized 3D masks from TotalSegmentator and subsequently retrained a 2D model for the liver, spleen, bowel, kidney, and body. The DICOM images were rescaled to (1, 1, 5) and stored with a 5-channel mask.<br>\nFor each epoch, I first sampled N=14 frames from the rescaled array, then used the body mask to filter out hands or other irrelevant areas.<br>\nI used the organ mask to limit the Z-axis space, as slices without the target organ might contain less valuable information.</p>\n<h2>Model</h2>\n<p>I used a Unet model integrated with Pyramid Vision Transformer V2 and MaxViT encoder. Transformers outperformed the convolutional models, especially for the extravasation target.</p>\n<p>Pretraining on the 2D mask facilitated convergence.<br>\n6 classification heads were employed for prediction.<br>\nPerhaps the lack of an RNN layer is the primary reason I didn't match the performance of the top teams.</p>\n<h2>Loss</h2>\n<p>I used the CE loss with weights identical to the metric.</p>\n<h2>Results</h2>\n<table>\n<thead>\n<tr>\n<th>Encoder</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>pvt-b2</td>\n<td>0.3783</td>\n<td>0.41</td>\n</tr>\n<tr>\n<td>pvt-b3</td>\n<td>0.3750</td>\n<td>0.41</td>\n</tr>\n<tr>\n<td>pvt-b4</td>\n<td>0.3786</td>\n<td>0.4</td>\n</tr>\n<tr>\n<td>maxvit_t</td>\n<td>0.3810</td>\n<td>0.42</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>0.3570</td>\n<td>0.4</td>\n</tr>\n<tr>\n<td>ensemble w/ scale</td>\n<td>0.3530</td>\n<td>0.39</td>\n</tr>\n</tbody>\n</table>\n<h2>Code</h2>\n<p><a href=\"https://github.com/iseekwonderful/RSNA-2023-Abdominal-Trauma-Detection-4th-Place-Code.git\" target=\"_blank\">Github link</a></p>",
      "rawMarkdown": "Thank you to RSNA and Kaggle for hosting this competition. Congratulations to all the winners and participants.\n\n## Overview\nI employed a 2.5 pipeline and trained for both classification and segmentation tasks.\n\n## Dataset\nI utilized 3D masks from TotalSegmentator and subsequently retrained a 2D model for the liver, spleen, bowel, kidney, and body. The DICOM images were rescaled to (1, 1, 5) and stored with a 5-channel mask.\nFor each epoch, I first sampled N=14 frames from the rescaled array, then used the body mask to filter out hands or other irrelevant areas.\nI used the organ mask to limit the Z-axis space, as slices without the target organ might contain less valuable information.\n## Model\nI used a Unet model integrated with Pyramid Vision Transformer V2 and MaxViT encoder. Transformers outperformed the convolutional models, especially for the extravasation target.\n\nPretraining on the 2D mask facilitated convergence.\n6 classification heads were employed for prediction.\nPerhaps the lack of an RNN layer is the primary reason I didn't match the performance of the top teams.\n## Loss\nI used the CE loss with weights identical to the metric.\n\n## Results\n| Encoder | CV | LB |\n| --- | --- | --- |\n| pvt-b2 | 0.3783 | 0.41 |\n| pvt-b3 | 0.3750 | 0.41 |\n| pvt-b4 | 0.3786 | 0.4 |\n| maxvit_t | 0.3810 | 0.42 |\n| ensemble | 0.3570 | 0.4 |\n| ensemble w/ scale | 0.3530 | 0.39 |\n## Code\n[Github link](https://github.com/iseekwonderful/RSNA-2023-Abdominal-Trauma-Detection-4th-Place-Code.git)",
      "votes": 21
    },
    {
      "id": 2486576,
      "postDate": "2023-10-18T03:25:52.727Z",
      "content": "<p>wow we actually have a very similar pipeline, the only difference is we used effnet (didn't have the compute to try transformer) :(</p>",
      "rawMarkdown": "wow we actually have a very similar pipeline, the only difference is we used effnet (didn't have the compute to try transformer) :(",
      "votes": 1,
      "replies": [
        {
          "id": 2487253,
          "postDate": "2023-10-18T13:27:49.137Z",
          "content": "<p>The gap between effb0 and pvtb2 is about 0.02. btw, your 3D model performance is impressive, mine even cannot reach 0.5. Do you use any pretrain weight, roi or some other method to improve the performance?</p>",
          "rawMarkdown": "The gap between effb0 and pvtb2 is about 0.02. btw, your 3D model performance is impressive, mine even cannot reach 0.5. Do you use any pretrain weight, roi or some other method to improve the performance?",
          "replies": [
            {
              "id": 2487303,
              "postDate": "2023-10-18T14:01:12.343Z",
              "content": "<p>yea i cropped the organs, only x3d backbones worked</p>",
              "rawMarkdown": "yea i cropped the organs, only x3d backbones worked"
            },
            {
              "id": 2487330,
              "postDate": "2023-10-18T14:19:30.647Z",
              "content": "<p>Thanks, I tried several backbones, only densenet can converge. It seems crop is necessary on 3D model.</p>",
              "rawMarkdown": "Thanks, I tried several backbones, only densenet can converge. It seems crop is necessary on 3D model."
            }
          ]
        }
      ]
    },
    {
      "id": 2485839,
      "postDate": "2023-10-17T13:43:00.033Z",
      "content": "<p><a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a>  , congrats on achieving 4th position. Looking to your code .</p>",
      "rawMarkdown": "@steamedsheep  , congrats on achieving 4th position. Looking to your code .",
      "votes": 1,
      "replies": [
        {
          "id": 2487248,
          "postDate": "2023-10-18T13:23:55.250Z",
          "content": "<p>Thank you.</p>",
          "rawMarkdown": "Thank you."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2486576,
      "author_name": "Feng Qilong",
      "author_url": "",
      "post_date": "2023-10-18T03:25:52.727000",
      "content": "<p>wow we actually have a very similar pipeline, the only difference is we used effnet (didn't have the compute to try transformer) :(</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2487253,
          "author_name": "sheep",
          "author_url": "",
          "post_date": "2023-10-18T13:27:49.137000",
          "content": "<p>The gap between effb0 and pvtb2 is about 0.02. btw, your 3D model performance is impressive, mine even cannot reach 0.5. Do you use any pretrain weight, roi or some other method to improve the performance?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2487303,
              "author_name": "Feng Qilong",
              "author_url": "",
              "post_date": "2023-10-18T14:01:12.343000",
              "content": "<p>yea i cropped the organs, only x3d backbones worked</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2487330,
              "author_name": "sheep",
              "author_url": "",
              "post_date": "2023-10-18T14:19:30.647000",
              "content": "<p>Thanks, I tried several backbones, only densenet can converge. It seems crop is necessary on 3D model.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2485839,
      "author_name": "SOUMENDRA PRASAD MOHANTY",
      "author_url": "",
      "post_date": "2023-10-17T13:43:00.033000",
      "content": "<p><a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a>  , congrats on achieving 4th position. Looking to your code .</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2487248,
          "author_name": "sheep",
          "author_url": "",
          "post_date": "2023-10-18T13:23:55.250000",
          "content": "<p>Thank you.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2485806": "Thank you to RSNA and Kaggle for hosting this competition. Congratulations to all the winners and participants.\n\n## Overview\nI employed a 2.5 pipeline and trained for both classification and segmentation tasks.\n\n## Dataset\nI utilized 3D masks from TotalSegmentator and subsequently retrained a 2D model for the liver, spleen, bowel, kidney, and body. The DICOM images were rescaled to (1, 1, 5) and stored with a 5-channel mask.\nFor each epoch, I first sampled N=14 frames from the rescaled array, then used the body mask to filter out hands or other irrelevant areas.\nI used the organ mask to limit the Z-axis space, as slices without the target organ might contain less valuable information.\n## Model\nI used a Unet model integrated with Pyramid Vision Transformer V2 and MaxViT encoder. Transformers outperformed the convolutional models, especially for the extravasation target.\n\nPretraining on the 2D mask facilitated convergence.\n6 classification heads were employed for prediction.\nPerhaps the lack of an RNN layer is the primary reason I didn't match the performance of the top teams.\n## Loss\nI used the CE loss with weights identical to the metric.\n\n## Results\n| Encoder | CV | LB |\n| --- | --- | --- |\n| pvt-b2 | 0.3783 | 0.41 |\n| pvt-b3 | 0.3750 | 0.41 |\n| pvt-b4 | 0.3786 | 0.4 |\n| maxvit_t | 0.3810 | 0.42 |\n| ensemble | 0.3570 | 0.4 |\n| ensemble w/ scale | 0.3530 | 0.39 |\n## Code\n[Github link](https://github.com/iseekwonderful/RSNA-2023-Abdominal-Trauma-Detection-4th-Place-Code.git)",
    "2486576": "wow we actually have a very similar pipeline, the only difference is we used effnet (didn't have the compute to try transformer) :(",
    "2485839": "@steamedsheep  , congrats on achieving 4th position. Looking to your code ."
  }
}