{
  "id": 362593,
  "title": "32nd place solution",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362593",
  "author_name": "yu4u",
  "post_date": "2022-10-28T00:13:21.561000",
  "votes": 43,
  "comment_count": 15,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2Fcf946877ffa2055047fade8fb8a8808f%2Frsna2022.png?generation=1666915657071769&amp;alt=media\" alt=\"\"></p>\n<p>Congrats to all prize and medal winners!<br>\nI enjoyed this competition because there are a great variety of options for solving this task. I could not try all of them of course, but for example:</p>\n<ul>\n<li>classifier vs. object detection approach</li>\n<li>3D model vs. 2.5D model</li>\n<li>simultaneously detect vertebrae and fracture vs. first detect vertebrae and then classify them (or detect bboxes)</li>\n<li>train segmentation mask to extract cervical vertebrae vs. use raw voxel (slices)</li>\n<li>calculate overall probability from C1-C7 probabilities vs. train a dedicated model for overall fracture prediction</li>\n</ul>\n<p>My brief summary of solution is:</p>\n<ul>\n<li>pipeline<ul>\n<li>use segmentation model to extract cervical vertebrae</li>\n<li>crop C1-C7 regions as voxels</li>\n<li>predict fracture probability for each cervical vertebra (C1-C7) with classification model</li>\n<li>use the sum of C1-C7 fracture probabilities as a overall fracture probability (with clipping)</li></ul></li>\n<li>segmentation model<ul>\n<li>MONAI 3D UNet</li>\n<li>surprisingly works well with a small amount of training data</li></ul></li>\n<li>classification model<ul>\n<li>2.5D CNN (EfficientNetV2-L) + Transformer encoder</li>\n<li>use different positional embedding for each cervical vertebra (C1-C7)</li></ul></li>\n</ul>\n<p>Does not work for me</p>\n<ul>\n<li>3D CNN classifier<ul>\n<li>I prefer 3D CNN approach to 2.5D approach ;(</li>\n<li>same as UW-Madison GI Tract Image Segmentation competition (3D CNN works but 2.5D was better)</li></ul></li>\n<li>simultaneously predict C1-C7 probabilities with attention (without segmentation)</li>\n<li>dedicated overall model</li>\n</ul>\n<p>I look forward to seeing the other teams' solutions!</p>",
  "messages": [
    {
      "id": 2006995,
      "postDate": "2022-10-28T00:13:21.560Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2Fcf946877ffa2055047fade8fb8a8808f%2Frsna2022.png?generation=1666915657071769&amp;alt=media\" alt=\"\"></p>\n<p>Congrats to all prize and medal winners!<br>\nI enjoyed this competition because there are a great variety of options for solving this task. I could not try all of them of course, but for example:</p>\n<ul>\n<li>classifier vs. object detection approach</li>\n<li>3D model vs. 2.5D model</li>\n<li>simultaneously detect vertebrae and fracture vs. first detect vertebrae and then classify them (or detect bboxes)</li>\n<li>train segmentation mask to extract cervical vertebrae vs. use raw voxel (slices)</li>\n<li>calculate overall probability from C1-C7 probabilities vs. train a dedicated model for overall fracture prediction</li>\n</ul>\n<p>My brief summary of solution is:</p>\n<ul>\n<li>pipeline<ul>\n<li>use segmentation model to extract cervical vertebrae</li>\n<li>crop C1-C7 regions as voxels</li>\n<li>predict fracture probability for each cervical vertebra (C1-C7) with classification model</li>\n<li>use the sum of C1-C7 fracture probabilities as a overall fracture probability (with clipping)</li></ul></li>\n<li>segmentation model<ul>\n<li>MONAI 3D UNet</li>\n<li>surprisingly works well with a small amount of training data</li></ul></li>\n<li>classification model<ul>\n<li>2.5D CNN (EfficientNetV2-L) + Transformer encoder</li>\n<li>use different positional embedding for each cervical vertebra (C1-C7)</li></ul></li>\n</ul>\n<p>Does not work for me</p>\n<ul>\n<li>3D CNN classifier<ul>\n<li>I prefer 3D CNN approach to 2.5D approach ;(</li>\n<li>same as UW-Madison GI Tract Image Segmentation competition (3D CNN works but 2.5D was better)</li></ul></li>\n<li>simultaneously predict C1-C7 probabilities with attention (without segmentation)</li>\n<li>dedicated overall model</li>\n</ul>\n<p>I look forward to seeing the other teams' solutions!</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2Fcf946877ffa2055047fade8fb8a8808f%2Frsna2022.png?generation=1666915657071769&alt=media)\n\nCongrats to all prize and medal winners!\nI enjoyed this competition because there are a great variety of options for solving this task. I could not try all of them of course, but for example:\n\n- classifier vs. object detection approach\n- 3D model vs. 2.5D model\n- simultaneously detect vertebrae and fracture vs. first detect vertebrae and then classify them (or detect bboxes)\n- train segmentation mask to extract cervical vertebrae vs. use raw voxel (slices)\n- calculate overall probability from C1-C7 probabilities vs. train a dedicated model for overall fracture prediction\n\nMy brief summary of solution is:\n\n- pipeline\n  - use segmentation model to extract cervical vertebrae\n  - crop C1-C7 regions as voxels\n  - predict fracture probability for each cervical vertebra (C1-C7) with classification model\n  - use the sum of C1-C7 fracture probabilities as a overall fracture probability (with clipping)\n- segmentation model\n  - MONAI 3D UNet\n  - surprisingly works well with a small amount of training data\n- classification model\n  - 2.5D CNN (EfficientNetV2-L) + Transformer encoder\n  - use different positional embedding for each cervical vertebra (C1-C7)\n\nDoes not work for me\n\n- 3D CNN classifier\n  - I prefer 3D CNN approach to 2.5D approach ;(\n  - same as UW-Madison GI Tract Image Segmentation competition (3D CNN works but 2.5D was better)\n- simultaneously predict C1-C7 probabilities with attention (without segmentation)\n- dedicated overall model\n\n\nI look forward to seeing the other teams' solutions!\n",
      "votes": 42
    },
    {
      "id": 2007056,
      "postDate": "2022-10-28T01:50:12.510Z",
      "content": "<p>An elegant approach and implementation :)</p>",
      "rawMarkdown": "An elegant approach and implementation :)",
      "votes": 1
    },
    {
      "id": 2007082,
      "postDate": "2022-10-28T02:46:42.650Z",
      "content": "<p>Oh your pipeline is quite similar to mine 😃</p>",
      "rawMarkdown": "Oh your pipeline is quite similar to mine 😃",
      "votes": 2,
      "replies": [
        {
          "id": 2007278,
          "postDate": "2022-10-28T06:13:57.410Z",
          "content": "<p>Wow, I read your solution and realized it. The solution is quite similar but the score is twice as different lol<br>\nI should do some additional experiments to figure out what made the difference.</p>",
          "rawMarkdown": "Wow, I read your solution and realized it. The solution is quite similar but the score is twice as different lol\nI should do some additional experiments to figure out what made the difference."
        },
        {
          "id": 2007291,
          "postDate": "2022-10-28T06:25:00.883Z",
          "content": "<p>yes it is very interesting that we can now compare your solution with Qishen's solution. Qishen embeds the mask as 6th channel to the 2.5D data. It may help. His type2 model is innovative but according to him just improved the CV by 0.01. I am still trying to understand the logic behind his LSTM type2 model. You are not using LSTM, but I do not think that it is the main score differentiation factor.</p>",
          "rawMarkdown": "yes it is very interesting that we can now compare your solution with Qishen's solution. Qishen embeds the mask as 6th channel to the 2.5D data. It may help. His type2 model is innovative but according to him just improved the CV by 0.01. I am still trying to understand the logic behind his LSTM type2 model. You are not using LSTM, but I do not think that it is the main score differentiation factor."
        }
      ]
    },
    {
      "id": 2007001,
      "postDate": "2022-10-28T00:18:59.487Z",
      "content": "<p>Nice work.  How did you feed the 4x256x256  (x8) into EffNet?  Is that a 4 channel images somehow concatenated together?  Or a 4x512x1024 image&gt;?</p>",
      "rawMarkdown": "Nice work.  How did you feed the 4x256x256  (x8) into EffNet?  Is that a 4 channel images somehow concatenated together?  Or a 4x512x1024 image>?\n",
      "votes": 2,
      "replies": [
        {
          "id": 2007005,
          "postDate": "2022-10-28T00:26:47.957Z",
          "content": "<p>Thx! 4 channel 256x256 images (8 images) are fed into the same CNN backbone independently.</p>",
          "rawMarkdown": "Thx! 4 channel 256x256 images (8 images) are fed into the same CNN backbone independently.",
          "votes": 2
        },
        {
          "id": 2007119,
          "postDate": "2022-10-28T03:25:54.927Z",
          "content": "<p>Thanks for sharing your solution <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> . In this way, how do you set the labels for these 8 4-channel images? Do they all use the original labels?</p>",
          "rawMarkdown": "Thanks for sharing your solution @ren4yu . In this way, how do you set the labels for these 8 4-channel images? Do they all use the original labels?",
          "votes": 1
        },
        {
          "id": 2007283,
          "postDate": "2022-10-28T06:18:55.223Z",
          "content": "<p>I used original labels. \"8 4-channel images\" are finally encoded into a single output via global average pooling. It is not the case that 8 images are independently used to predict (eight) labels.</p>",
          "rawMarkdown": "I used original labels. \"8 4-channel images\" are finally encoded into a single output via global average pooling. It is not the case that 8 images are independently used to predict (eight) labels."
        }
      ]
    },
    {
      "id": 2184068,
      "postDate": "2023-03-16T05:46:00.327Z",
      "content": "<p>Can you provide your code and elaborate on your extensive description of your solution?</p>",
      "rawMarkdown": "Can you provide your code and elaborate on your extensive description of your solution?"
    },
    {
      "id": 2007464,
      "postDate": "2022-10-28T08:35:56.060Z",
      "content": "<p>Great Solution! <br>\nWhat was your idea behind using transformer encoder after the classification model? you could just use the predictions from the classification model</p>",
      "rawMarkdown": "Great Solution! \nWhat was your idea behind using transformer encoder after the classification model? you could just use the predictions from the classification model"
    },
    {
      "id": 2007121,
      "postDate": "2022-10-28T03:26:56.640Z",
      "content": "<p>Interesting solution indeed. Bravo. How do you obtain the 32x.x. images? I imagine that the output of the cropping phase are some overlapped volumes between adjacent vertebrates. Right? Do you interpolate each volume to a fixed size 32x.x. volume?   </p>",
      "rawMarkdown": "Interesting solution indeed. Bravo. How do you obtain the 32x.x. images? I imagine that the output of the cropping phase are some overlapped volumes between adjacent vertebrates. Right? Do you interpolate each volume to a fixed size 32x.x. volume?   ",
      "replies": [
        {
          "id": 2007285,
          "postDate": "2022-10-28T06:22:02.520Z",
          "content": "<p>Yes, the cropped voxels have large overlap between them. Cropped voxels are then resized into the fixed sizes (32x256x256) via F.interpolate function.</p>",
          "rawMarkdown": "Yes, the cropped voxels have large overlap between them. Cropped voxels are then resized into the fixed sizes (32x256x256) via F.interpolate function.",
          "votes": 1
        },
        {
          "id": 2007295,
          "postDate": "2022-10-28T06:27:16.890Z",
          "content": "<p>thanks for the clarification</p>",
          "rawMarkdown": "thanks for the clarification"
        }
      ]
    },
    {
      "id": 2007098,
      "postDate": "2022-10-28T02:57:43.797Z",
      "content": "<p>Very intuitive pipeline. Thanks for sharing.</p>",
      "rawMarkdown": "Very intuitive pipeline. Thanks for sharing."
    },
    {
      "id": 2007114,
      "postDate": "2022-10-28T03:18:21.953Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 2007056,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2022-10-28T01:50:12.510000",
      "content": "<p>An elegant approach and implementation :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2007082,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2022-10-28T02:46:42.650000",
      "content": "<p>Oh your pipeline is quite similar to mine 😃</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2007278,
          "author_name": "yu4u",
          "author_url": "",
          "post_date": "2022-10-28T06:13:57.410000",
          "content": "<p>Wow, I read your solution and realized it. The solution is quite similar but the score is twice as different lol<br>\nI should do some additional experiments to figure out what made the difference.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2007291,
          "author_name": "Behnam Molaee",
          "author_url": "",
          "post_date": "2022-10-28T06:25:00.883000",
          "content": "<p>yes it is very interesting that we can now compare your solution with Qishen's solution. Qishen embeds the mask as 6th channel to the 2.5D data. It may help. His type2 model is innovative but according to him just improved the CV by 0.01. I am still trying to understand the logic behind his LSTM type2 model. You are not using LSTM, but I do not think that it is the main score differentiation factor.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2007001,
      "author_name": "SolverWorld",
      "author_url": "",
      "post_date": "2022-10-28T00:18:59.487000",
      "content": "<p>Nice work.  How did you feed the 4x256x256  (x8) into EffNet?  Is that a 4 channel images somehow concatenated together?  Or a 4x512x1024 image&gt;?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2007005,
          "author_name": "yu4u",
          "author_url": "",
          "post_date": "2022-10-28T00:26:47.957000",
          "content": "<p>Thx! 4 channel 256x256 images (8 images) are fed into the same CNN backbone independently.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2007119,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2022-10-28T03:25:54.927000",
          "content": "<p>Thanks for sharing your solution <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> . In this way, how do you set the labels for these 8 4-channel images? Do they all use the original labels?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2007283,
          "author_name": "yu4u",
          "author_url": "",
          "post_date": "2022-10-28T06:18:55.223000",
          "content": "<p>I used original labels. \"8 4-channel images\" are finally encoded into a single output via global average pooling. It is not the case that 8 images are independently used to predict (eight) labels.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2184068,
      "author_name": "Gayathri Mahesh",
      "author_url": "",
      "post_date": "2023-03-16T05:46:00.327000",
      "content": "<p>Can you provide your code and elaborate on your extensive description of your solution?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2007464,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2022-10-28T08:35:56.060000",
      "content": "<p>Great Solution! <br>\nWhat was your idea behind using transformer encoder after the classification model? you could just use the predictions from the classification model</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2007121,
      "author_name": "Behnam Molaee",
      "author_url": "",
      "post_date": "2022-10-28T03:26:56.640000",
      "content": "<p>Interesting solution indeed. Bravo. How do you obtain the 32x.x. images? I imagine that the output of the cropping phase are some overlapped volumes between adjacent vertebrates. Right? Do you interpolate each volume to a fixed size 32x.x. volume?   </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2007285,
          "author_name": "yu4u",
          "author_url": "",
          "post_date": "2022-10-28T06:22:02.520000",
          "content": "<p>Yes, the cropped voxels have large overlap between them. Cropped voxels are then resized into the fixed sizes (32x256x256) via F.interpolate function.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2007295,
          "author_name": "Behnam Molaee",
          "author_url": "",
          "post_date": "2022-10-28T06:27:16.890000",
          "content": "<p>thanks for the clarification</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2007098,
      "author_name": "Bardia Khosravi",
      "author_url": "",
      "post_date": "2022-10-28T02:57:43.797000",
      "content": "<p>Very intuitive pipeline. Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2007114,
      "author_name": "Doktoroso ",
      "author_url": "",
      "post_date": "2022-10-28T03:18:21.953000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2006995": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2Fcf946877ffa2055047fade8fb8a8808f%2Frsna2022.png?generation=1666915657071769&alt=media)\n\nCongrats to all prize and medal winners!\nI enjoyed this competition because there are a great variety of options for solving this task. I could not try all of them of course, but for example:\n\n- classifier vs. object detection approach\n- 3D model vs. 2.5D model\n- simultaneously detect vertebrae and fracture vs. first detect vertebrae and then classify them (or detect bboxes)\n- train segmentation mask to extract cervical vertebrae vs. use raw voxel (slices)\n- calculate overall probability from C1-C7 probabilities vs. train a dedicated model for overall fracture prediction\n\nMy brief summary of solution is:\n\n- pipeline\n  - use segmentation model to extract cervical vertebrae\n  - crop C1-C7 regions as voxels\n  - predict fracture probability for each cervical vertebra (C1-C7) with classification model\n  - use the sum of C1-C7 fracture probabilities as a overall fracture probability (with clipping)\n- segmentation model\n  - MONAI 3D UNet\n  - surprisingly works well with a small amount of training data\n- classification model\n  - 2.5D CNN (EfficientNetV2-L) + Transformer encoder\n  - use different positional embedding for each cervical vertebra (C1-C7)\n\nDoes not work for me\n\n- 3D CNN classifier\n  - I prefer 3D CNN approach to 2.5D approach ;(\n  - same as UW-Madison GI Tract Image Segmentation competition (3D CNN works but 2.5D was better)\n- simultaneously predict C1-C7 probabilities with attention (without segmentation)\n- dedicated overall model\n\n\nI look forward to seeing the other teams' solutions!\n",
    "2007056": "An elegant approach and implementation :)",
    "2007082": "Oh your pipeline is quite similar to mine 😃",
    "2007001": "Nice work.  How did you feed the 4x256x256  (x8) into EffNet?  Is that a 4 channel images somehow concatenated together?  Or a 4x512x1024 image>?\n",
    "2184068": "Can you provide your code and elaborate on your extensive description of your solution?",
    "2007464": "Great Solution! \nWhat was your idea behind using transformer encoder after the classification model? you could just use the predictions from the classification model",
    "2007121": "Interesting solution indeed. Bravo. How do you obtain the 32x.x. images? I imagine that the output of the cropping phase are some overlapped volumes between adjacent vertebrates. Right? Do you interpolate each volume to a fixed size 32x.x. volume?   ",
    "2007098": "Very intuitive pipeline. Thanks for sharing.",
    "2007114": "Thanks for sharing!"
  }
}