{
  "id": 362651,
  "title": "[6th place] Solution Overview: 3D CNN + TD-CNN",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362651",
  "author_name": "Ian Pan",
  "post_date": "2022-10-28T10:47:41.942000",
  "votes": 59,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Thank you to the organizers for putting together this interesting competition, and congratulations to all of the competitors on their efforts. I'm glad that shakeup was kind to me on private LB. One of my goals after becoming GM was to win a gold model with &lt;10 submissions, and I'm happy to accomplish that with this challenge. This is a brief overview of my solution. </p>\n<p>A more detailed writeup is available here: <a href=\"https://docs.google.com/document/d/1GcHeUDks2dnECKmJn97nF2djFRTVVnGpORFpUH0nMuU/edit\" target=\"_blank\">https://docs.google.com/document/d/1GcHeUDks2dnECKmJn97nF2djFRTVVnGpORFpUH0nMuU/edit</a></p>\n<h2>Summary</h2>\n<p>-3D cervical spine segmentation<br>\n-Individual vertebra extraction<br>\n-Stage 1 models: 3D CNN and TD-CNN for vertebra-level fracture classification<br>\n-Stage 2 models: transformers for final exam-level classification using features extracted from stage 1 models </p>\n<h2>Segmentation</h2>\n<p>I used a 3D DeepLabV3+ encoder-decoder architecture with X3D 3D CNN backbone (<a href=\"https://pytorchvideo.readthedocs.io/en/latest/api/models/x3d.html\" target=\"_blank\">https://pytorchvideo.readthedocs.io/en/latest/api/models/x3d.html</a>) trained at 192x192x192 resolution to generate segmentations of C1-C7. </p>\n<p>Models were first trained on the 87 semantic segmentation labels provided by the organizers. I then pseudolabeled the remaining studies with these models and retrained on the whole dataset. </p>\n<p>Using the output of the segmentation models, one can extract a cube containing the vertebra of interest. Inevitably, there is some overlap with other vertebra levels due to their orientation within the scan. </p>\n<h2>Stage 1: 3D CNN</h2>\n<p>Once a cube is extracted for each vertebra, I trained a X3D-L 3D CNN to perform binary fracture classification for each level. It was straightforward to map each vertebra to a binary label, as we're provided with the fractured levels for each study. The input size was 64x288x288 for vertebra. I also modified the model so that the z-stride was 1 for all layers. Thus, the first dimension (64) was not downsampled until final pooling. </p>\n<p>After training the 3D classification model, features (432-D) were extracted from each vertebra. Thus, each study was represented by a 7x432 sequence. </p>\n<h2>Stage 2: TD-CNN</h2>\n<p>A TD-CNN is simply a 2D CNN feature extractor with a sequence model head. In my case, I used a 2-layer transformer.  For a given volume NxHxW, the 2D CNN extracts a D-dimensional feature for each individual image (HxW) in the volume, which are then input (NxD) to the sequence model head. Ultimately, the model is trained to end-to-end. </p>\n<p>I trained this model in 3 parts. First, I trained a TF-EfficientNetV2-S model to act as the 2D CNN feature extractor. </p>\n<p>In order to train this model, I needed image-level labels. While there was a subset of the dataset which did have those labels, I was able to pseudolabel the entire dataset by a method I call class activation sequence. </p>\n<p>Using the 3D CNN fracture classification model, I could generate a 3D class activation map for each vertebra by removing the pooling and classification layers and then obtaining a weighted average of the 3D feature maps using the classification layer weights. For each z-axis image in the 3D feature map, the max value was taken, converting the 3D feature map into a 1D sequence. This sequence was then rescaled to [0, 1] and resampled back to the original number of slices so a single value could be corresponded to each slice.</p>\n<p>e.g., (64, 288, 288) -&gt; (432, 64, 9, 9) -&gt; (64, 9, 9) -&gt; (64, ) -&gt; (original # of slices, )</p>\n<p>Because one vertebra-level input usually contained more than 1 level, there was overlap between these sequences. To deal with this, I just averaged the overlapping values. </p>\n<p>I thresholded the values at 0.5 to generate pseudolabels for each image in the dataset. I did not use the provided image-level labels for training. </p>\n<p>Stepwise training occurred as follows:<br>\n1) Train 2D CNN feature extractor as binary classifier on pseudolabels (288x288)<br>\n2) Add transformer head, freeze feature extractor, train head <br>\n3) Fine-tune entire model end-to-end (32x288x288)</p>\n<p>Similar to the 3D CNN model, a feature was extracted for each vertebra (256-D) after training, resulting in a 7x256 sequence for each vertebra. </p>\n<h2>Stage 2: Transformers</h2>\n<p>Now that we have essentially converted each study into a sequence, we can use transformers to model the final output (C1-C7 fractures and overall fracture). </p>\n<p>I used a 3-layer transformer and trained on 3 separate inputs:<br>\n1) 7x432 input from 3D CNN (5-fold CV: 0.3205)<br>\n2) 7x256 input from TD-CNN (5-fold CV: 0.3369)<br>\n3) 7x688 input by fusing the above 2 sequences (5-fold CV: 0.2962) </p>\n<p>The final submission was an ensemble of the outputs from the above 3 models (0.25, 0.25, 0.5). </p>\n<h2>Things that did not help:</h2>\n<p>-Adding mask as separate channel in image<br>\n-Masking out the non-segmented parts of the image<br>\n-Extracting slice-wise features and training a 2D-CNN + sequence model on study-wise features- doing the vertebra-level approach was better</p>\n<h2>Code</h2>\n<p><a href=\"https://github.com/i-pan/kaggle-rsna-cspine\" target=\"_blank\">https://github.com/i-pan/kaggle-rsna-cspine</a></p>",
  "messages": [
    {
      "id": 2007593,
      "postDate": "2022-10-28T10:47:41.943Z",
      "content": "<p>Thank you to the organizers for putting together this interesting competition, and congratulations to all of the competitors on their efforts. I'm glad that shakeup was kind to me on private LB. One of my goals after becoming GM was to win a gold model with &lt;10 submissions, and I'm happy to accomplish that with this challenge. This is a brief overview of my solution. </p>\n<p>A more detailed writeup is available here: <a href=\"https://docs.google.com/document/d/1GcHeUDks2dnECKmJn97nF2djFRTVVnGpORFpUH0nMuU/edit\" target=\"_blank\">https://docs.google.com/document/d/1GcHeUDks2dnECKmJn97nF2djFRTVVnGpORFpUH0nMuU/edit</a></p>\n<h2>Summary</h2>\n<p>-3D cervical spine segmentation<br>\n-Individual vertebra extraction<br>\n-Stage 1 models: 3D CNN and TD-CNN for vertebra-level fracture classification<br>\n-Stage 2 models: transformers for final exam-level classification using features extracted from stage 1 models </p>\n<h2>Segmentation</h2>\n<p>I used a 3D DeepLabV3+ encoder-decoder architecture with X3D 3D CNN backbone (<a href=\"https://pytorchvideo.readthedocs.io/en/latest/api/models/x3d.html\" target=\"_blank\">https://pytorchvideo.readthedocs.io/en/latest/api/models/x3d.html</a>) trained at 192x192x192 resolution to generate segmentations of C1-C7. </p>\n<p>Models were first trained on the 87 semantic segmentation labels provided by the organizers. I then pseudolabeled the remaining studies with these models and retrained on the whole dataset. </p>\n<p>Using the output of the segmentation models, one can extract a cube containing the vertebra of interest. Inevitably, there is some overlap with other vertebra levels due to their orientation within the scan. </p>\n<h2>Stage 1: 3D CNN</h2>\n<p>Once a cube is extracted for each vertebra, I trained a X3D-L 3D CNN to perform binary fracture classification for each level. It was straightforward to map each vertebra to a binary label, as we're provided with the fractured levels for each study. The input size was 64x288x288 for vertebra. I also modified the model so that the z-stride was 1 for all layers. Thus, the first dimension (64) was not downsampled until final pooling. </p>\n<p>After training the 3D classification model, features (432-D) were extracted from each vertebra. Thus, each study was represented by a 7x432 sequence. </p>\n<h2>Stage 2: TD-CNN</h2>\n<p>A TD-CNN is simply a 2D CNN feature extractor with a sequence model head. In my case, I used a 2-layer transformer.  For a given volume NxHxW, the 2D CNN extracts a D-dimensional feature for each individual image (HxW) in the volume, which are then input (NxD) to the sequence model head. Ultimately, the model is trained to end-to-end. </p>\n<p>I trained this model in 3 parts. First, I trained a TF-EfficientNetV2-S model to act as the 2D CNN feature extractor. </p>\n<p>In order to train this model, I needed image-level labels. While there was a subset of the dataset which did have those labels, I was able to pseudolabel the entire dataset by a method I call class activation sequence. </p>\n<p>Using the 3D CNN fracture classification model, I could generate a 3D class activation map for each vertebra by removing the pooling and classification layers and then obtaining a weighted average of the 3D feature maps using the classification layer weights. For each z-axis image in the 3D feature map, the max value was taken, converting the 3D feature map into a 1D sequence. This sequence was then rescaled to [0, 1] and resampled back to the original number of slices so a single value could be corresponded to each slice.</p>\n<p>e.g., (64, 288, 288) -&gt; (432, 64, 9, 9) -&gt; (64, 9, 9) -&gt; (64, ) -&gt; (original # of slices, )</p>\n<p>Because one vertebra-level input usually contained more than 1 level, there was overlap between these sequences. To deal with this, I just averaged the overlapping values. </p>\n<p>I thresholded the values at 0.5 to generate pseudolabels for each image in the dataset. I did not use the provided image-level labels for training. </p>\n<p>Stepwise training occurred as follows:<br>\n1) Train 2D CNN feature extractor as binary classifier on pseudolabels (288x288)<br>\n2) Add transformer head, freeze feature extractor, train head <br>\n3) Fine-tune entire model end-to-end (32x288x288)</p>\n<p>Similar to the 3D CNN model, a feature was extracted for each vertebra (256-D) after training, resulting in a 7x256 sequence for each vertebra. </p>\n<h2>Stage 2: Transformers</h2>\n<p>Now that we have essentially converted each study into a sequence, we can use transformers to model the final output (C1-C7 fractures and overall fracture). </p>\n<p>I used a 3-layer transformer and trained on 3 separate inputs:<br>\n1) 7x432 input from 3D CNN (5-fold CV: 0.3205)<br>\n2) 7x256 input from TD-CNN (5-fold CV: 0.3369)<br>\n3) 7x688 input by fusing the above 2 sequences (5-fold CV: 0.2962) </p>\n<p>The final submission was an ensemble of the outputs from the above 3 models (0.25, 0.25, 0.5). </p>\n<h2>Things that did not help:</h2>\n<p>-Adding mask as separate channel in image<br>\n-Masking out the non-segmented parts of the image<br>\n-Extracting slice-wise features and training a 2D-CNN + sequence model on study-wise features- doing the vertebra-level approach was better</p>\n<h2>Code</h2>\n<p><a href=\"https://github.com/i-pan/kaggle-rsna-cspine\" target=\"_blank\">https://github.com/i-pan/kaggle-rsna-cspine</a></p>",
      "rawMarkdown": "Thank you to the organizers for putting together this interesting competition, and congratulations to all of the competitors on their efforts. I'm glad that shakeup was kind to me on private LB. One of my goals after becoming GM was to win a gold model with <10 submissions, and I'm happy to accomplish that with this challenge. This is a brief overview of my solution. \n\nA more detailed writeup is available here: https://docs.google.com/document/d/1GcHeUDks2dnECKmJn97nF2djFRTVVnGpORFpUH0nMuU/edit\n\n## Summary\n-3D cervical spine segmentation\n-Individual vertebra extraction\n-Stage 1 models: 3D CNN and TD-CNN for vertebra-level fracture classification\n-Stage 2 models: transformers for final exam-level classification using features extracted from stage 1 models \n\n## Segmentation\nI used a 3D DeepLabV3+ encoder-decoder architecture with X3D 3D CNN backbone (https://pytorchvideo.readthedocs.io/en/latest/api/models/x3d.html) trained at 192x192x192 resolution to generate segmentations of C1-C7. \n\nModels were first trained on the 87 semantic segmentation labels provided by the organizers. I then pseudolabeled the remaining studies with these models and retrained on the whole dataset. \n\nUsing the output of the segmentation models, one can extract a cube containing the vertebra of interest. Inevitably, there is some overlap with other vertebra levels due to their orientation within the scan. \n\n## Stage 1: 3D CNN\nOnce a cube is extracted for each vertebra, I trained a X3D-L 3D CNN to perform binary fracture classification for each level. It was straightforward to map each vertebra to a binary label, as we're provided with the fractured levels for each study. The input size was 64x288x288 for vertebra. I also modified the model so that the z-stride was 1 for all layers. Thus, the first dimension (64) was not downsampled until final pooling. \n\nAfter training the 3D classification model, features (432-D) were extracted from each vertebra. Thus, each study was represented by a 7x432 sequence. \n\n## Stage 2: TD-CNN\nA TD-CNN is simply a 2D CNN feature extractor with a sequence model head. In my case, I used a 2-layer transformer.  For a given volume NxHxW, the 2D CNN extracts a D-dimensional feature for each individual image (HxW) in the volume, which are then input (NxD) to the sequence model head. Ultimately, the model is trained to end-to-end. \n\nI trained this model in 3 parts. First, I trained a TF-EfficientNetV2-S model to act as the 2D CNN feature extractor. \n\nIn order to train this model, I needed image-level labels. While there was a subset of the dataset which did have those labels, I was able to pseudolabel the entire dataset by a method I call class activation sequence. \n\nUsing the 3D CNN fracture classification model, I could generate a 3D class activation map for each vertebra by removing the pooling and classification layers and then obtaining a weighted average of the 3D feature maps using the classification layer weights. For each z-axis image in the 3D feature map, the max value was taken, converting the 3D feature map into a 1D sequence. This sequence was then rescaled to [0, 1] and resampled back to the original number of slices so a single value could be corresponded to each slice.\n\ne.g., (64, 288, 288) -> (432, 64, 9, 9) -> (64, 9, 9) -> (64, ) -> (original # of slices, )\n\nBecause one vertebra-level input usually contained more than 1 level, there was overlap between these sequences. To deal with this, I just averaged the overlapping values. \n\nI thresholded the values at 0.5 to generate pseudolabels for each image in the dataset. I did not use the provided image-level labels for training. \n\nStepwise training occurred as follows:\n1) Train 2D CNN feature extractor as binary classifier on pseudolabels (288x288)\n2) Add transformer head, freeze feature extractor, train head \n3) Fine-tune entire model end-to-end (32x288x288)\n\nSimilar to the 3D CNN model, a feature was extracted for each vertebra (256-D) after training, resulting in a 7x256 sequence for each vertebra. \n\n## Stage 2: Transformers\nNow that we have essentially converted each study into a sequence, we can use transformers to model the final output (C1-C7 fractures and overall fracture). \n\nI used a 3-layer transformer and trained on 3 separate inputs:\n1) 7x432 input from 3D CNN (5-fold CV: 0.3205)\n2) 7x256 input from TD-CNN (5-fold CV: 0.3369)\n3) 7x688 input by fusing the above 2 sequences (5-fold CV: 0.2962) \n\nThe final submission was an ensemble of the outputs from the above 3 models (0.25, 0.25, 0.5). \n\n## Things that did not help:\n-Adding mask as separate channel in image\n-Masking out the non-segmented parts of the image\n-Extracting slice-wise features and training a 2D-CNN + sequence model on study-wise features- doing the vertebra-level approach was better\n\n## Code\nhttps://github.com/i-pan/kaggle-rsna-cspine",
      "votes": 59
    },
    {
      "id": 2007630,
      "postDate": "2022-10-28T11:31:37.573Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>  Your model is robust! Very little difference between private and public. </p>",
      "rawMarkdown": "Congratulations @vaillant  Your model is robust! Very little difference between private and public. ",
      "votes": 3,
      "replies": [
        {
          "id": 2010065,
          "postDate": "2022-10-30T13:23:04.383Z",
          "content": "<p>Thanks, and congratulations to you as well! </p>",
          "rawMarkdown": "Thanks, and congratulations to you as well! "
        }
      ]
    },
    {
      "id": 2007699,
      "postDate": "2022-10-28T12:38:08.147Z",
      "content": "<p>Congrats! Great work. And as usual very nice write up.</p>\n<p>That's amazing that you got the 3D CNN to work. I tried really hard to make 3D CNN work but my model didn't learn to pick up fractures. Instead it learned to recognize the level of the vertebral body and give the probability of the fracture based on which cervical level it is (I think… you never know what the model is truly doing but I notice all my C7 levels have similar probabilities).</p>",
      "rawMarkdown": "Congrats! Great work. And as usual very nice write up.\n\nThat's amazing that you got the 3D CNN to work. I tried really hard to make 3D CNN work but my model didn't learn to pick up fractures. Instead it learned to recognize the level of the vertebral body and give the probability of the fracture based on which cervical level it is (I think... you never know what the model is truly doing but I notice all my C7 levels have similar probabilities).",
      "votes": 4,
      "replies": [
        {
          "id": 2010064,
          "postDate": "2022-10-30T13:22:43.927Z",
          "content": "<p>I've had issues with 3D CNNs in the past, but the pretrained X3D models are surprisingly good for a variety of 3D tasks. </p>",
          "rawMarkdown": "I've had issues with 3D CNNs in the past, but the pretrained X3D models are surprisingly good for a variety of 3D tasks. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 2009137,
      "postDate": "2022-10-29T17:22:46Z",
      "content": "<p>Congrats! Great solution as always Ian.</p>\n<p>The 3D pseudolabeling at image-level is very unique. You said you didn't use the provided image-level labels. Does that mean you tried and didn't like their performance, or that you intuitively didn't trust them?</p>",
      "rawMarkdown": "Congrats! Great solution as always Ian.\n\nThe 3D pseudolabeling at image-level is very unique. You said you didn't use the provided image-level labels. Does that mean you tried and didn't like their performance, or that you intuitively didn't trust them?",
      "votes": 1,
      "replies": [
        {
          "id": 2010063,
          "postDate": "2022-10-30T13:21:59.647Z",
          "content": "<p>Thanks!</p>\n<p>I tried training feature extractors using only the image-level labeled subset, but it didn't seem to work as well as when I trained on the pseudo-labeled whole dataset. </p>\n<p>I didn't try using the both (e.g., provided labels when available, pseudolabels when not) since I didn't have much time, and things seemed to work well using only the pseudolabels. </p>",
          "rawMarkdown": "Thanks!\n\nI tried training feature extractors using only the image-level labeled subset, but it didn't seem to work as well as when I trained on the pseudo-labeled whole dataset. \n\nI didn't try using the both (e.g., provided labels when available, pseudolabels when not) since I didn't have much time, and things seemed to work well using only the pseudolabels. ",
          "votes": 1
        },
        {
          "id": 2011466,
          "postDate": "2022-10-31T15:29:57.353Z",
          "content": "<p>Makes sense. Would be a cool performance comparison to see. pseudolabel all vs pseudolabel+provided labels if available vs provided labels using only provided subset.</p>",
          "rawMarkdown": "Makes sense. Would be a cool performance comparison to see. pseudolabel all vs pseudolabel+provided labels if available vs provided labels using only provided subset."
        }
      ]
    },
    {
      "id": 2007617,
      "postDate": "2022-10-28T11:12:19.310Z",
      "content": "<p>Congratulations, both for solo gold and your personal single-digit sub goal, got to learn about both new models and new ways :)</p>",
      "rawMarkdown": "Congratulations, both for solo gold and your personal single-digit sub goal, got to learn about both new models and new ways :)",
      "votes": 2,
      "replies": [
        {
          "id": 2010067,
          "postDate": "2022-10-30T13:23:22.130Z",
          "content": "<p>Thanks and congrats as well!</p>",
          "rawMarkdown": "Thanks and congrats as well!"
        }
      ]
    },
    {
      "id": 2008942,
      "postDate": "2022-10-29T13:30:27.250Z",
      "content": "<p>Amazing work </p>",
      "rawMarkdown": "Amazing work "
    },
    {
      "id": 2164729,
      "postDate": "2023-03-01T18:54:31.257Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2164411,
      "postDate": "2023-03-01T14:47:02.993Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2164408,
      "postDate": "2023-03-01T14:41:56.810Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2007630,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2022-10-28T11:31:37.573000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>  Your model is robust! Very little difference between private and public. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 2010065,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2022-10-30T13:23:04.383000",
          "content": "<p>Thanks, and congratulations to you as well! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2007699,
      "author_name": "Yee Ng",
      "author_url": "",
      "post_date": "2022-10-28T12:38:08.147000",
      "content": "<p>Congrats! Great work. And as usual very nice write up.</p>\n<p>That's amazing that you got the 3D CNN to work. I tried really hard to make 3D CNN work but my model didn't learn to pick up fractures. Instead it learned to recognize the level of the vertebral body and give the probability of the fracture based on which cervical level it is (I think… you never know what the model is truly doing but I notice all my C7 levels have similar probabilities).</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2010064,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2022-10-30T13:22:43.927000",
          "content": "<p>I've had issues with 3D CNNs in the past, but the pretrained X3D models are surprisingly good for a variety of 3D tasks. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2009137,
      "author_name": "Jesse",
      "author_url": "",
      "post_date": "2022-10-29T17:22:46",
      "content": "<p>Congrats! Great solution as always Ian.</p>\n<p>The 3D pseudolabeling at image-level is very unique. You said you didn't use the provided image-level labels. Does that mean you tried and didn't like their performance, or that you intuitively didn't trust them?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2010063,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2022-10-30T13:21:59.647000",
          "content": "<p>Thanks!</p>\n<p>I tried training feature extractors using only the image-level labeled subset, but it didn't seem to work as well as when I trained on the pseudo-labeled whole dataset. </p>\n<p>I didn't try using the both (e.g., provided labels when available, pseudolabels when not) since I didn't have much time, and things seemed to work well using only the pseudolabels. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2011466,
          "author_name": "Jesse",
          "author_url": "",
          "post_date": "2022-10-31T15:29:57.353000",
          "content": "<p>Makes sense. Would be a cool performance comparison to see. pseudolabel all vs pseudolabel+provided labels if available vs provided labels using only provided subset.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2007617,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2022-10-28T11:12:19.310000",
      "content": "<p>Congratulations, both for solo gold and your personal single-digit sub goal, got to learn about both new models and new ways :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2010067,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2022-10-30T13:23:22.130000",
          "content": "<p>Thanks and congrats as well!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2008942,
      "author_name": "Arnab_Dey",
      "author_url": "",
      "post_date": "2022-10-29T13:30:27.250000",
      "content": "<p>Amazing work </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2164729,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-01T18:54:31.257000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2164411,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-01T14:47:02.993000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2164408,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-01T14:41:56.810000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2007593": "Thank you to the organizers for putting together this interesting competition, and congratulations to all of the competitors on their efforts. I'm glad that shakeup was kind to me on private LB. One of my goals after becoming GM was to win a gold model with <10 submissions, and I'm happy to accomplish that with this challenge. This is a brief overview of my solution. \n\nA more detailed writeup is available here: https://docs.google.com/document/d/1GcHeUDks2dnECKmJn97nF2djFRTVVnGpORFpUH0nMuU/edit\n\n## Summary\n-3D cervical spine segmentation\n-Individual vertebra extraction\n-Stage 1 models: 3D CNN and TD-CNN for vertebra-level fracture classification\n-Stage 2 models: transformers for final exam-level classification using features extracted from stage 1 models \n\n## Segmentation\nI used a 3D DeepLabV3+ encoder-decoder architecture with X3D 3D CNN backbone (https://pytorchvideo.readthedocs.io/en/latest/api/models/x3d.html) trained at 192x192x192 resolution to generate segmentations of C1-C7. \n\nModels were first trained on the 87 semantic segmentation labels provided by the organizers. I then pseudolabeled the remaining studies with these models and retrained on the whole dataset. \n\nUsing the output of the segmentation models, one can extract a cube containing the vertebra of interest. Inevitably, there is some overlap with other vertebra levels due to their orientation within the scan. \n\n## Stage 1: 3D CNN\nOnce a cube is extracted for each vertebra, I trained a X3D-L 3D CNN to perform binary fracture classification for each level. It was straightforward to map each vertebra to a binary label, as we're provided with the fractured levels for each study. The input size was 64x288x288 for vertebra. I also modified the model so that the z-stride was 1 for all layers. Thus, the first dimension (64) was not downsampled until final pooling. \n\nAfter training the 3D classification model, features (432-D) were extracted from each vertebra. Thus, each study was represented by a 7x432 sequence. \n\n## Stage 2: TD-CNN\nA TD-CNN is simply a 2D CNN feature extractor with a sequence model head. In my case, I used a 2-layer transformer.  For a given volume NxHxW, the 2D CNN extracts a D-dimensional feature for each individual image (HxW) in the volume, which are then input (NxD) to the sequence model head. Ultimately, the model is trained to end-to-end. \n\nI trained this model in 3 parts. First, I trained a TF-EfficientNetV2-S model to act as the 2D CNN feature extractor. \n\nIn order to train this model, I needed image-level labels. While there was a subset of the dataset which did have those labels, I was able to pseudolabel the entire dataset by a method I call class activation sequence. \n\nUsing the 3D CNN fracture classification model, I could generate a 3D class activation map for each vertebra by removing the pooling and classification layers and then obtaining a weighted average of the 3D feature maps using the classification layer weights. For each z-axis image in the 3D feature map, the max value was taken, converting the 3D feature map into a 1D sequence. This sequence was then rescaled to [0, 1] and resampled back to the original number of slices so a single value could be corresponded to each slice.\n\ne.g., (64, 288, 288) -> (432, 64, 9, 9) -> (64, 9, 9) -> (64, ) -> (original # of slices, )\n\nBecause one vertebra-level input usually contained more than 1 level, there was overlap between these sequences. To deal with this, I just averaged the overlapping values. \n\nI thresholded the values at 0.5 to generate pseudolabels for each image in the dataset. I did not use the provided image-level labels for training. \n\nStepwise training occurred as follows:\n1) Train 2D CNN feature extractor as binary classifier on pseudolabels (288x288)\n2) Add transformer head, freeze feature extractor, train head \n3) Fine-tune entire model end-to-end (32x288x288)\n\nSimilar to the 3D CNN model, a feature was extracted for each vertebra (256-D) after training, resulting in a 7x256 sequence for each vertebra. \n\n## Stage 2: Transformers\nNow that we have essentially converted each study into a sequence, we can use transformers to model the final output (C1-C7 fractures and overall fracture). \n\nI used a 3-layer transformer and trained on 3 separate inputs:\n1) 7x432 input from 3D CNN (5-fold CV: 0.3205)\n2) 7x256 input from TD-CNN (5-fold CV: 0.3369)\n3) 7x688 input by fusing the above 2 sequences (5-fold CV: 0.2962) \n\nThe final submission was an ensemble of the outputs from the above 3 models (0.25, 0.25, 0.5). \n\n## Things that did not help:\n-Adding mask as separate channel in image\n-Masking out the non-segmented parts of the image\n-Extracting slice-wise features and training a 2D-CNN + sequence model on study-wise features- doing the vertebra-level approach was better\n\n## Code\nhttps://github.com/i-pan/kaggle-rsna-cspine",
    "2007630": "Congratulations @vaillant  Your model is robust! Very little difference between private and public. ",
    "2007699": "Congrats! Great work. And as usual very nice write up.\n\nThat's amazing that you got the 3D CNN to work. I tried really hard to make 3D CNN work but my model didn't learn to pick up fractures. Instead it learned to recognize the level of the vertebral body and give the probability of the fracture based on which cervical level it is (I think... you never know what the model is truly doing but I notice all my C7 levels have similar probabilities).",
    "2009137": "Congrats! Great solution as always Ian.\n\nThe 3D pseudolabeling at image-level is very unique. You said you didn't use the provided image-level labels. Does that mean you tried and didn't like their performance, or that you intuitively didn't trust them?",
    "2007617": "Congratulations, both for solo gold and your personal single-digit sub goal, got to learn about both new models and new ways :)",
    "2008942": "Amazing work ",
    "2164729": "",
    "2164411": "",
    "2164408": ""
  }
}