{
  "id": 362643,
  "title": "3rd place solution",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362643",
  "author_name": "Darragh",
  "post_date": "2022-10-28T09:32:07.625000",
  "votes": 59,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Congrats all and thank you RSNA for another great challenge. Special kudos to <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> (and <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a>) and others who showed model performance can really be pushed more here. </p>\n<p>Here is my solution, I may fill out a bit more detail, or correct some of it in time. </p>\n<p>Training code : <a href=\"https://github.com/darraghdog/RSNA22\" target=\"_blank\">https://github.com/darraghdog/RSNA22</a><br>\nInference code : <a href=\"https://www.kaggle.com/code/darraghdog/rsna-2022-3rd-place-solution-inference\" target=\"_blank\">https://www.kaggle.com/code/darraghdog/rsna-2022-3rd-place-solution-inference</a><br>\nSlides located <a href=\"https://docs.google.com/presentation/d/1lS4yOTJT4EyaCjODGIO811RGex9jKQdZypzOjY_jxDA/edit?usp=sharing\" target=\"_blank\">here</a><br>\nVideo of solution : <a href=\"https://youtu.be/f-QA5MLN16Q\" target=\"_blank\">here</a></p>\n<h3>TL;DR</h3>\n<p>Adjacent to the problem, I hoped to get a solution which required less intense work on labelling, so this could be more easily scaled by the spine radiology specialists from the ASNR and ASSR. <br>\nI did not use the fracture bounding boxes, and only used highlevel data from the segmentation maps. This consisted of two things from the segmentation maps (1) bounding box of the C1-C7 vertebrae by taking the outer limit of the segmentation, (2) ratio of vertebrae volume in the slice divided by max vertebrae volume seen in any slice. This second point is explained in a bit more detail below. If we had bounding box labels, instead of segmentations, for the vertebrae, these could also be used to calculate the ratios. <br>\nFor point (1) the individual slice level bounding boxes were not used downstream. Instead a study level bounding box for the C1-C7 vertebrae, which was taken from the rollmean max of the individual slice level bounding boxes. <br>\nFrom point (2), in combination with the fracture labels in <code>train.csv</code> we can get an approximate label for fracture and vertebrae type in each slice to fit models. I only used 2.5D CNN + 1D RNN which is pretty much lifted straight from <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> ‘s <a href=\"https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/discussion/194145\" target=\"_blank\">first place solution</a> to RSNA two years ago.  </p>\n<h3>Bounding box preprocessing</h3>\n<p>Again lifted from first place approach two years ago, I found it increased performance to zoom in on the vertebrae. I used efficientnet-v2 to predict five labels - <code>x0, y0, x1, y1, has_bbox</code>. The four corners of the bounding box, and a probability if the slice has a vertebrae or not. In downstream models, the range of slices before C1-7 vertebrae began in the z-axis, and after they ended, were excluded for training and inference. This range was found by using z-axis wise rollmean of <code>has_bbox</code> probability.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F213493%2Fb811845bd8a6afd471f1453695189e65%2Fbbox.png?generation=1666948850696604&amp;alt=media\" alt=\"\"><br>\nAlso in downstream models the vertebrae were cropped in preprocessing. I used the outer bbox of all boxes in the z-axis to crop the whole study in one go with a single box per study. The cropped study was then resized to 512 * 512, and the same augmentation (shiftscale, cutout, etc) was applied to the study. I also cropped and resized the dicom before converting to uint8 in the hope that the interpolating the raw dicom values to a larger size would give greater resolution - not sure if this helped or not, but it gave me peace of mind 😊 </p>\n<h3>Model 1 : Slice level vertebrae labels</h3>\n<p>A 2.5D CNN+1d RNN (same as described below) was used to train a model using the 87study level segmentation maps. As mentioned the label was the ratio of vertebrae volume in the slice divided by max vertebrae volume seen in any slice. This was trained on z-axis windows of studies - so 32 * 3 slices per sample, aggregated to 32 2.5D images passed through the CNN and then the RNN predicts the vertebrae ratio. RMSE loss was used. <br>\nThe slice level CV predictions for C1-C7 vertebrae were multiplied by the study level fracture labels to give slice level fracture labels as seen below. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F213493%2F263b4ef7609e970ee0810902307a0589%2Fpreds.png?generation=1666948962692906&amp;alt=media\" alt=\"\"></p>\n<h3>Model 2 : Initial sequential model</h3>\n<p>The same architecture again (2.5D CNN+1d RNN) was used to train on z-axis windows of studies with RMSE loss on three different targets - 1) Slice level vertebrae ratio 2) Slice level fracture label (seen above) and 3) Max fracture value in the window.<br>\nI trained on random 32 * 3 slice windows of studies. A lot of these windows had no fracture so I used a batchsize of 48 (using accumulation 16) to ensure some fractures are in each batch. Experiments with undersampling did not help.  <br>\nThe model was simple enough, but the addition of the attention mechanism over the 1d RNN output helped a lot. Transformers or other architectures did not improve it. Two backbone’s were used from timm - resnest50d and seresnext50. Resnest50d was used by <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> in the RSNA STR competition two years ago and performed best. The model trained for a long time - ~9 hours per fold. </p>\n<h3>Model 3 : Final sequential model</h3>\n<p>The same architecture again (2.5D CNN+1d RNN) was used to train the full study over on the final study level labels, found in <code>train.csv</code>. The CNN backbone was loaded from the checkpoint weights of model 2 and set to no gradients, and a new 1d RNN and attention mechanism were initialized and trained on top of this. With the final labels, the competition metric as loss was used. The only additional difference in the model was squeezing the embedding layers in very long studies - again, similar to <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> ‘s - if there were more than 192 * 3 slices outputted, torch functional interpolation was used to reshape them to a max sequence of 192. <br>\nAs mentioned before, for model 2 and 3, a single bounding box was used to crop all slices in a study, and the same augmentation used across the study. For model 3 only, the <code>has_bbox</code> range was used to exclude slices before the vertebrae started and ended. And for model3, the CNN outputted embeddings were extracted in chunks of 32 * 3 2.5d images. </p>\n<h3>Final Submission</h3>\n<p>Weights from the bounding box model and the model3 were used for inference. A combination of the  resnest50d and seresnext50, which had 6 and 4 respectively sets of weights from model 3 trained over the full dataset. <br>\nI found it kind on RAM to collect batches of studies in uint8 format, and normalize on GPU; and inference through the CNN was performed in chunks of 32 * 2.5d images at a time. </p>\n<p>Inference time was ~ 4.5 hours. </p>",
  "messages": [
    {
      "id": 2007514,
      "postDate": "2022-10-28T09:32:07.627Z",
      "content": "<p>Congrats all and thank you RSNA for another great challenge. Special kudos to <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> (and <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a>) and others who showed model performance can really be pushed more here. </p>\n<p>Here is my solution, I may fill out a bit more detail, or correct some of it in time. </p>\n<p>Training code : <a href=\"https://github.com/darraghdog/RSNA22\" target=\"_blank\">https://github.com/darraghdog/RSNA22</a><br>\nInference code : <a href=\"https://www.kaggle.com/code/darraghdog/rsna-2022-3rd-place-solution-inference\" target=\"_blank\">https://www.kaggle.com/code/darraghdog/rsna-2022-3rd-place-solution-inference</a><br>\nSlides located <a href=\"https://docs.google.com/presentation/d/1lS4yOTJT4EyaCjODGIO811RGex9jKQdZypzOjY_jxDA/edit?usp=sharing\" target=\"_blank\">here</a><br>\nVideo of solution : <a href=\"https://youtu.be/f-QA5MLN16Q\" target=\"_blank\">here</a></p>\n<h3>TL;DR</h3>\n<p>Adjacent to the problem, I hoped to get a solution which required less intense work on labelling, so this could be more easily scaled by the spine radiology specialists from the ASNR and ASSR. <br>\nI did not use the fracture bounding boxes, and only used highlevel data from the segmentation maps. This consisted of two things from the segmentation maps (1) bounding box of the C1-C7 vertebrae by taking the outer limit of the segmentation, (2) ratio of vertebrae volume in the slice divided by max vertebrae volume seen in any slice. This second point is explained in a bit more detail below. If we had bounding box labels, instead of segmentations, for the vertebrae, these could also be used to calculate the ratios. <br>\nFor point (1) the individual slice level bounding boxes were not used downstream. Instead a study level bounding box for the C1-C7 vertebrae, which was taken from the rollmean max of the individual slice level bounding boxes. <br>\nFrom point (2), in combination with the fracture labels in <code>train.csv</code> we can get an approximate label for fracture and vertebrae type in each slice to fit models. I only used 2.5D CNN + 1D RNN which is pretty much lifted straight from <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> ‘s <a href=\"https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/discussion/194145\" target=\"_blank\">first place solution</a> to RSNA two years ago.  </p>\n<h3>Bounding box preprocessing</h3>\n<p>Again lifted from first place approach two years ago, I found it increased performance to zoom in on the vertebrae. I used efficientnet-v2 to predict five labels - <code>x0, y0, x1, y1, has_bbox</code>. The four corners of the bounding box, and a probability if the slice has a vertebrae or not. In downstream models, the range of slices before C1-7 vertebrae began in the z-axis, and after they ended, were excluded for training and inference. This range was found by using z-axis wise rollmean of <code>has_bbox</code> probability.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F213493%2Fb811845bd8a6afd471f1453695189e65%2Fbbox.png?generation=1666948850696604&amp;alt=media\" alt=\"\"><br>\nAlso in downstream models the vertebrae were cropped in preprocessing. I used the outer bbox of all boxes in the z-axis to crop the whole study in one go with a single box per study. The cropped study was then resized to 512 * 512, and the same augmentation (shiftscale, cutout, etc) was applied to the study. I also cropped and resized the dicom before converting to uint8 in the hope that the interpolating the raw dicom values to a larger size would give greater resolution - not sure if this helped or not, but it gave me peace of mind 😊 </p>\n<h3>Model 1 : Slice level vertebrae labels</h3>\n<p>A 2.5D CNN+1d RNN (same as described below) was used to train a model using the 87study level segmentation maps. As mentioned the label was the ratio of vertebrae volume in the slice divided by max vertebrae volume seen in any slice. This was trained on z-axis windows of studies - so 32 * 3 slices per sample, aggregated to 32 2.5D images passed through the CNN and then the RNN predicts the vertebrae ratio. RMSE loss was used. <br>\nThe slice level CV predictions for C1-C7 vertebrae were multiplied by the study level fracture labels to give slice level fracture labels as seen below. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F213493%2F263b4ef7609e970ee0810902307a0589%2Fpreds.png?generation=1666948962692906&amp;alt=media\" alt=\"\"></p>\n<h3>Model 2 : Initial sequential model</h3>\n<p>The same architecture again (2.5D CNN+1d RNN) was used to train on z-axis windows of studies with RMSE loss on three different targets - 1) Slice level vertebrae ratio 2) Slice level fracture label (seen above) and 3) Max fracture value in the window.<br>\nI trained on random 32 * 3 slice windows of studies. A lot of these windows had no fracture so I used a batchsize of 48 (using accumulation 16) to ensure some fractures are in each batch. Experiments with undersampling did not help.  <br>\nThe model was simple enough, but the addition of the attention mechanism over the 1d RNN output helped a lot. Transformers or other architectures did not improve it. Two backbone’s were used from timm - resnest50d and seresnext50. Resnest50d was used by <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> in the RSNA STR competition two years ago and performed best. The model trained for a long time - ~9 hours per fold. </p>\n<h3>Model 3 : Final sequential model</h3>\n<p>The same architecture again (2.5D CNN+1d RNN) was used to train the full study over on the final study level labels, found in <code>train.csv</code>. The CNN backbone was loaded from the checkpoint weights of model 2 and set to no gradients, and a new 1d RNN and attention mechanism were initialized and trained on top of this. With the final labels, the competition metric as loss was used. The only additional difference in the model was squeezing the embedding layers in very long studies - again, similar to <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> ‘s - if there were more than 192 * 3 slices outputted, torch functional interpolation was used to reshape them to a max sequence of 192. <br>\nAs mentioned before, for model 2 and 3, a single bounding box was used to crop all slices in a study, and the same augmentation used across the study. For model 3 only, the <code>has_bbox</code> range was used to exclude slices before the vertebrae started and ended. And for model3, the CNN outputted embeddings were extracted in chunks of 32 * 3 2.5d images. </p>\n<h3>Final Submission</h3>\n<p>Weights from the bounding box model and the model3 were used for inference. A combination of the  resnest50d and seresnext50, which had 6 and 4 respectively sets of weights from model 3 trained over the full dataset. <br>\nI found it kind on RAM to collect batches of studies in uint8 format, and normalize on GPU; and inference through the CNN was performed in chunks of 32 * 2.5d images at a time. </p>\n<p>Inference time was ~ 4.5 hours. </p>",
      "rawMarkdown": "Congrats all and thank you RSNA for another great challenge. Special kudos to @haqishen (and @selimsef) and others who showed model performance can really be pushed more here. \n\nHere is my solution, I may fill out a bit more detail, or correct some of it in time. \n\nTraining code : https://github.com/darraghdog/RSNA22\nInference code : https://www.kaggle.com/code/darraghdog/rsna-2022-3rd-place-solution-inference\nSlides located [here](https://docs.google.com/presentation/d/1lS4yOTJT4EyaCjODGIO811RGex9jKQdZypzOjY_jxDA/edit?usp=sharing)\nVideo of solution : [here](https://youtu.be/f-QA5MLN16Q)\n\n### TL;DR\nAdjacent to the problem, I hoped to get a solution which required less intense work on labelling, so this could be more easily scaled by the spine radiology specialists from the ASNR and ASSR. \nI did not use the fracture bounding boxes, and only used highlevel data from the segmentation maps. This consisted of two things from the segmentation maps (1) bounding box of the C1-C7 vertebrae by taking the outer limit of the segmentation, (2) ratio of vertebrae volume in the slice divided by max vertebrae volume seen in any slice. This second point is explained in a bit more detail below. If we had bounding box labels, instead of segmentations, for the vertebrae, these could also be used to calculate the ratios. \nFor point (1) the individual slice level bounding boxes were not used downstream. Instead a study level bounding box for the C1-C7 vertebrae, which was taken from the rollmean max of the individual slice level bounding boxes. \nFrom point (2), in combination with the fracture labels in `train.csv` we can get an approximate label for fracture and vertebrae type in each slice to fit models. I only used 2.5D CNN + 1D RNN which is pretty much lifted straight from @wowfattie ‘s [first place solution](https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/discussion/194145) to RSNA two years ago.  \n\n### Bounding box preprocessing\nAgain lifted from first place approach two years ago, I found it increased performance to zoom in on the vertebrae. I used efficientnet-v2 to predict five labels - `x0, y0, x1, y1, has_bbox`. The four corners of the bounding box, and a probability if the slice has a vertebrae or not. In downstream models, the range of slices before C1-7 vertebrae began in the z-axis, and after they ended, were excluded for training and inference. This range was found by using z-axis wise rollmean of `has_bbox` probability.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F213493%2Fb811845bd8a6afd471f1453695189e65%2Fbbox.png?generation=1666948850696604&alt=media)\nAlso in downstream models the vertebrae were cropped in preprocessing. I used the outer bbox of all boxes in the z-axis to crop the whole study in one go with a single box per study. The cropped study was then resized to 512 * 512, and the same augmentation (shiftscale, cutout, etc) was applied to the study. I also cropped and resized the dicom before converting to uint8 in the hope that the interpolating the raw dicom values to a larger size would give greater resolution - not sure if this helped or not, but it gave me peace of mind 😊 \n\n### Model 1 : Slice level vertebrae labels\nA 2.5D CNN+1d RNN (same as described below) was used to train a model using the 87study level segmentation maps. As mentioned the label was the ratio of vertebrae volume in the slice divided by max vertebrae volume seen in any slice. This was trained on z-axis windows of studies - so 32 * 3 slices per sample, aggregated to 32 2.5D images passed through the CNN and then the RNN predicts the vertebrae ratio. RMSE loss was used. \nThe slice level CV predictions for C1-C7 vertebrae were multiplied by the study level fracture labels to give slice level fracture labels as seen below. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F213493%2F263b4ef7609e970ee0810902307a0589%2Fpreds.png?generation=1666948962692906&alt=media)\n\n### Model 2 : Initial sequential model\nThe same architecture again (2.5D CNN+1d RNN) was used to train on z-axis windows of studies with RMSE loss on three different targets - 1) Slice level vertebrae ratio 2) Slice level fracture label (seen above) and 3) Max fracture value in the window.\nI trained on random 32 * 3 slice windows of studies. A lot of these windows had no fracture so I used a batchsize of 48 (using accumulation 16) to ensure some fractures are in each batch. Experiments with undersampling did not help.  \nThe model was simple enough, but the addition of the attention mechanism over the 1d RNN output helped a lot. Transformers or other architectures did not improve it. Two backbone’s were used from timm - resnest50d and seresnext50. Resnest50d was used by @vaillant in the RSNA STR competition two years ago and performed best. The model trained for a long time - ~9 hours per fold. \n\n### Model 3 : Final sequential model\nThe same architecture again (2.5D CNN+1d RNN) was used to train the full study over on the final study level labels, found in `train.csv`. The CNN backbone was loaded from the checkpoint weights of model 2 and set to no gradients, and a new 1d RNN and attention mechanism were initialized and trained on top of this. With the final labels, the competition metric as loss was used. The only additional difference in the model was squeezing the embedding layers in very long studies - again, similar to @wowfattie ‘s - if there were more than 192 * 3 slices outputted, torch functional interpolation was used to reshape them to a max sequence of 192. \nAs mentioned before, for model 2 and 3, a single bounding box was used to crop all slices in a study, and the same augmentation used across the study. For model 3 only, the `has_bbox` range was used to exclude slices before the vertebrae started and ended. And for model3, the CNN outputted embeddings were extracted in chunks of 32 * 3 2.5d images. \n\n### Final Submission\nWeights from the bounding box model and the model3 were used for inference. A combination of the  resnest50d and seresnext50, which had 6 and 4 respectively sets of weights from model 3 trained over the full dataset. \nI found it kind on RAM to collect batches of studies in uint8 format, and normalize on GPU; and inference through the CNN was performed in chunks of 32 * 2.5d images at a time. \n\nInference time was ~ 4.5 hours. \n",
      "votes": 59
    },
    {
      "id": 2007684,
      "postDate": "2022-10-28T12:26:16.710Z",
      "content": "<p>Thanks for the write up, congrats! </p>\n<p>So when you are creating the bounding boxes of the C1-C7 vertebrae by taking the outer limit of the segmentation, was the algorithm ever thrown off by stray pixels that can appear anywhere on the image? Do you have a preprocessing step to ignore pixel islands that are less than a number? I did something similar but I calculated the center of mass and standard deviation of the distribution of the mask and then use it to create bounding boxes (bounding cubes actually). But the boxes are not as tight  around the spine as the images you showed above.</p>",
      "rawMarkdown": "Thanks for the write up, congrats! \n\nSo when you are creating the bounding boxes of the C1-C7 vertebrae by taking the outer limit of the segmentation, was the algorithm ever thrown off by stray pixels that can appear anywhere on the image? Do you have a preprocessing step to ignore pixel islands that are less than a number? I did something similar but I calculated the center of mass and standard deviation of the distribution of the mask and then use it to create bounding boxes (bounding cubes actually). But the boxes are not as tight  around the spine as the images you showed above.",
      "votes": 3,
      "replies": [
        {
          "id": 2007748,
          "postDate": "2022-10-28T13:05:58.203Z",
          "content": "<blockquote>\n  <p>was the algorithm ever thrown off by stray pixels that can appear anywhere on the image? Do you have a preprocessing step to ignore pixel islands that are less than a number? </p>\n</blockquote>\n<p>I was just thinking about this today, it probably was - I took a z-axis rolling mean of bounding box coordinates which would somehow counteract this, but not fully. </p>\n<blockquote>\n  <p>I did something similar but I calculated the center of mass and standard deviation of the distribution of the mask and then use it to create bounding boxes (bounding cubes actually). But the boxes are not as tight around the spine as the images you showed above.</p>\n</blockquote>\n<p>Nice approach, it probably would have helped me. </p>",
          "rawMarkdown": "> was the algorithm ever thrown off by stray pixels that can appear anywhere on the image? Do you have a preprocessing step to ignore pixel islands that are less than a number? \n\nI was just thinking about this today, it probably was - I took a z-axis rolling mean of bounding box coordinates which would somehow counteract this, but not fully. \n\n>I did something similar but I calculated the center of mass and standard deviation of the distribution of the mask and then use it to create bounding boxes (bounding cubes actually). But the boxes are not as tight around the spine as the images you showed above.\n\nNice approach, it probably would have helped me. "
        }
      ]
    },
    {
      "id": 2007995,
      "postDate": "2022-10-28T16:22:10.703Z",
      "content": "<p>Congratulations! Thank you for sharing your idea!!!</p>",
      "rawMarkdown": "Congratulations! Thank you for sharing your idea!!!",
      "votes": 1
    },
    {
      "id": 2007873,
      "postDate": "2022-10-28T14:58:45.293Z",
      "content": "<p>Congratulations! Amazing Work</p>",
      "rawMarkdown": "Congratulations! Amazing Work",
      "votes": 1
    },
    {
      "id": 2007784,
      "postDate": "2022-10-28T13:41:32.827Z",
      "content": "<p>Congratulations!   In your Model 2, do you train end-to-end a  single 2.5D CNN (32 copies of a CNN with same weights feeding to 1d RNN)?  Thx.</p>",
      "rawMarkdown": "Congratulations!   In your Model 2, do you train end-to-end a  single 2.5D CNN (32 copies of a CNN with same weights feeding to 1d RNN)?  Thx.",
      "votes": 1,
      "replies": [
        {
          "id": 2007808,
          "postDate": "2022-10-28T14:02:45.307Z",
          "content": "<p>Models were initialised roughly like, </p>\n<pre><code>self.backbone = timm.create_model('resnest50d')\nhidden_size = self.backbone.fc.in_features # 2048\nself.backbone.fc = torch.nn.Identity()\nself.rnn = nn.LSTM(hidden_size, hidden_size, batch_first=True,  bidirectional=True)\n</code></pre>\n<p>Then forward step (with shapes for illustration)</p>\n<pre><code>batchsize, seqlen, ch, h, w = batch['image'].shape # (2,32,3,512,512)\nx = batch['image'].view(-1, ch, h, w) # (64,3,512,512)\nemb = self.backbone(x) #  (64,2048)\nemb = emb.view(batchsize, seqlen, -1) #&amp;nbsp;(2, 32 ,2048)\nlogits = self.rnn(emb)[0] # (2, 32 , 4096)\n</code></pre>",
          "rawMarkdown": "Models were initialised roughly like, \n```        \nself.backbone = timm.create_model('resnest50d')\nhidden_size = self.backbone.fc.in_features # 2048\nself.backbone.fc = torch.nn.Identity()\nself.rnn = nn.LSTM(hidden_size, hidden_size, batch_first=True,  bidirectional=True)\n```\nThen forward step (with shapes for illustration)\n\n```\nbatchsize, seqlen, ch, h, w = batch['image'].shape # (2,32,3,512,512)\nx = batch['image'].view(-1, ch, h, w) # (64,3,512,512)\nemb = self.backbone(x) #  (64,2048)\nemb = emb.view(batchsize, seqlen, -1) # (2, 32 ,2048)\nlogits = self.rnn(emb)[0] # (2, 32 , 4096)\n```\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 2007607,
      "postDate": "2022-10-28T11:07:28.323Z",
      "content": "<p>Congratulations!, Honestly, I did not get much of it, but seems like a really cool solution</p>",
      "rawMarkdown": "Congratulations!, Honestly, I did not get much of it, but seems like a really cool solution",
      "votes": 1
    },
    {
      "id": 2007577,
      "postDate": "2022-10-28T10:21:12.793Z",
      "content": "<p>Congrats! You have a completely different pipeline, I learned a lot!</p>",
      "rawMarkdown": "Congrats! You have a completely different pipeline, I learned a lot!",
      "votes": 1
    },
    {
      "id": 2008159,
      "postDate": "2022-10-28T19:36:12.547Z",
      "content": "<p>Amazing work and a great result to culminate the effort <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a>! Keep it up and hearty congratulations!</p>",
      "rawMarkdown": "Amazing work and a great result to culminate the effort @darraghdog! Keep it up and hearty congratulations!",
      "votes": 2
    },
    {
      "id": 2007595,
      "postDate": "2022-10-28T10:54:23.217Z",
      "content": "<p>Congratulations! The vertebra ratio idea is very cool. </p>",
      "rawMarkdown": "Congratulations! The vertebra ratio idea is very cool. ",
      "votes": 2
    },
    {
      "id": 2008430,
      "postDate": "2022-10-29T04:06:31.147Z",
      "content": "<p>Congratulations!! Thanks for sharing</p>",
      "rawMarkdown": "Congratulations!! Thanks for sharing\n",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2007684,
      "author_name": "Yee Ng",
      "author_url": "",
      "post_date": "2022-10-28T12:26:16.710000",
      "content": "<p>Thanks for the write up, congrats! </p>\n<p>So when you are creating the bounding boxes of the C1-C7 vertebrae by taking the outer limit of the segmentation, was the algorithm ever thrown off by stray pixels that can appear anywhere on the image? Do you have a preprocessing step to ignore pixel islands that are less than a number? I did something similar but I calculated the center of mass and standard deviation of the distribution of the mask and then use it to create bounding boxes (bounding cubes actually). But the boxes are not as tight  around the spine as the images you showed above.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2007748,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2022-10-28T13:05:58.203000",
          "content": "<blockquote>\n  <p>was the algorithm ever thrown off by stray pixels that can appear anywhere on the image? Do you have a preprocessing step to ignore pixel islands that are less than a number? </p>\n</blockquote>\n<p>I was just thinking about this today, it probably was - I took a z-axis rolling mean of bounding box coordinates which would somehow counteract this, but not fully. </p>\n<blockquote>\n  <p>I did something similar but I calculated the center of mass and standard deviation of the distribution of the mask and then use it to create bounding boxes (bounding cubes actually). But the boxes are not as tight around the spine as the images you showed above.</p>\n</blockquote>\n<p>Nice approach, it probably would have helped me. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2007995,
      "author_name": "Markrvtx",
      "author_url": "",
      "post_date": "2022-10-28T16:22:10.703000",
      "content": "<p>Congratulations! Thank you for sharing your idea!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2007873,
      "author_name": "Arnab_Dey",
      "author_url": "",
      "post_date": "2022-10-28T14:58:45.293000",
      "content": "<p>Congratulations! Amazing Work</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2007784,
      "author_name": "SolverWorld",
      "author_url": "",
      "post_date": "2022-10-28T13:41:32.827000",
      "content": "<p>Congratulations!   In your Model 2, do you train end-to-end a  single 2.5D CNN (32 copies of a CNN with same weights feeding to 1d RNN)?  Thx.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2007808,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2022-10-28T14:02:45.307000",
          "content": "<p>Models were initialised roughly like, </p>\n<pre><code>self.backbone = timm.create_model('resnest50d')\nhidden_size = self.backbone.fc.in_features # 2048\nself.backbone.fc = torch.nn.Identity()\nself.rnn = nn.LSTM(hidden_size, hidden_size, batch_first=True,  bidirectional=True)\n</code></pre>\n<p>Then forward step (with shapes for illustration)</p>\n<pre><code>batchsize, seqlen, ch, h, w = batch['image'].shape # (2,32,3,512,512)\nx = batch['image'].view(-1, ch, h, w) # (64,3,512,512)\nemb = self.backbone(x) #  (64,2048)\nemb = emb.view(batchsize, seqlen, -1) #&amp;nbsp;(2, 32 ,2048)\nlogits = self.rnn(emb)[0] # (2, 32 , 4096)\n</code></pre>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2007607,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2022-10-28T11:07:28.323000",
      "content": "<p>Congratulations!, Honestly, I did not get much of it, but seems like a really cool solution</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2007577,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2022-10-28T10:21:12.793000",
      "content": "<p>Congrats! You have a completely different pipeline, I learned a lot!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2008159,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2022-10-28T19:36:12.547000",
      "content": "<p>Amazing work and a great result to culminate the effort <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a>! Keep it up and hearty congratulations!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2007595,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2022-10-28T10:54:23.217000",
      "content": "<p>Congratulations! The vertebra ratio idea is very cool. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2008430,
      "author_name": "Shreya Mishra 0307",
      "author_url": "",
      "post_date": "2022-10-29T04:06:31.147000",
      "content": "<p>Congratulations!! Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2007514": "Congrats all and thank you RSNA for another great challenge. Special kudos to @haqishen (and @selimsef) and others who showed model performance can really be pushed more here. \n\nHere is my solution, I may fill out a bit more detail, or correct some of it in time. \n\nTraining code : https://github.com/darraghdog/RSNA22\nInference code : https://www.kaggle.com/code/darraghdog/rsna-2022-3rd-place-solution-inference\nSlides located [here](https://docs.google.com/presentation/d/1lS4yOTJT4EyaCjODGIO811RGex9jKQdZypzOjY_jxDA/edit?usp=sharing)\nVideo of solution : [here](https://youtu.be/f-QA5MLN16Q)\n\n### TL;DR\nAdjacent to the problem, I hoped to get a solution which required less intense work on labelling, so this could be more easily scaled by the spine radiology specialists from the ASNR and ASSR. \nI did not use the fracture bounding boxes, and only used highlevel data from the segmentation maps. This consisted of two things from the segmentation maps (1) bounding box of the C1-C7 vertebrae by taking the outer limit of the segmentation, (2) ratio of vertebrae volume in the slice divided by max vertebrae volume seen in any slice. This second point is explained in a bit more detail below. If we had bounding box labels, instead of segmentations, for the vertebrae, these could also be used to calculate the ratios. \nFor point (1) the individual slice level bounding boxes were not used downstream. Instead a study level bounding box for the C1-C7 vertebrae, which was taken from the rollmean max of the individual slice level bounding boxes. \nFrom point (2), in combination with the fracture labels in `train.csv` we can get an approximate label for fracture and vertebrae type in each slice to fit models. I only used 2.5D CNN + 1D RNN which is pretty much lifted straight from @wowfattie ‘s [first place solution](https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/discussion/194145) to RSNA two years ago.  \n\n### Bounding box preprocessing\nAgain lifted from first place approach two years ago, I found it increased performance to zoom in on the vertebrae. I used efficientnet-v2 to predict five labels - `x0, y0, x1, y1, has_bbox`. The four corners of the bounding box, and a probability if the slice has a vertebrae or not. In downstream models, the range of slices before C1-7 vertebrae began in the z-axis, and after they ended, were excluded for training and inference. This range was found by using z-axis wise rollmean of `has_bbox` probability.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F213493%2Fb811845bd8a6afd471f1453695189e65%2Fbbox.png?generation=1666948850696604&alt=media)\nAlso in downstream models the vertebrae were cropped in preprocessing. I used the outer bbox of all boxes in the z-axis to crop the whole study in one go with a single box per study. The cropped study was then resized to 512 * 512, and the same augmentation (shiftscale, cutout, etc) was applied to the study. I also cropped and resized the dicom before converting to uint8 in the hope that the interpolating the raw dicom values to a larger size would give greater resolution - not sure if this helped or not, but it gave me peace of mind 😊 \n\n### Model 1 : Slice level vertebrae labels\nA 2.5D CNN+1d RNN (same as described below) was used to train a model using the 87study level segmentation maps. As mentioned the label was the ratio of vertebrae volume in the slice divided by max vertebrae volume seen in any slice. This was trained on z-axis windows of studies - so 32 * 3 slices per sample, aggregated to 32 2.5D images passed through the CNN and then the RNN predicts the vertebrae ratio. RMSE loss was used. \nThe slice level CV predictions for C1-C7 vertebrae were multiplied by the study level fracture labels to give slice level fracture labels as seen below. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F213493%2F263b4ef7609e970ee0810902307a0589%2Fpreds.png?generation=1666948962692906&alt=media)\n\n### Model 2 : Initial sequential model\nThe same architecture again (2.5D CNN+1d RNN) was used to train on z-axis windows of studies with RMSE loss on three different targets - 1) Slice level vertebrae ratio 2) Slice level fracture label (seen above) and 3) Max fracture value in the window.\nI trained on random 32 * 3 slice windows of studies. A lot of these windows had no fracture so I used a batchsize of 48 (using accumulation 16) to ensure some fractures are in each batch. Experiments with undersampling did not help.  \nThe model was simple enough, but the addition of the attention mechanism over the 1d RNN output helped a lot. Transformers or other architectures did not improve it. Two backbone’s were used from timm - resnest50d and seresnext50. Resnest50d was used by @vaillant in the RSNA STR competition two years ago and performed best. The model trained for a long time - ~9 hours per fold. \n\n### Model 3 : Final sequential model\nThe same architecture again (2.5D CNN+1d RNN) was used to train the full study over on the final study level labels, found in `train.csv`. The CNN backbone was loaded from the checkpoint weights of model 2 and set to no gradients, and a new 1d RNN and attention mechanism were initialized and trained on top of this. With the final labels, the competition metric as loss was used. The only additional difference in the model was squeezing the embedding layers in very long studies - again, similar to @wowfattie ‘s - if there were more than 192 * 3 slices outputted, torch functional interpolation was used to reshape them to a max sequence of 192. \nAs mentioned before, for model 2 and 3, a single bounding box was used to crop all slices in a study, and the same augmentation used across the study. For model 3 only, the `has_bbox` range was used to exclude slices before the vertebrae started and ended. And for model3, the CNN outputted embeddings were extracted in chunks of 32 * 3 2.5d images. \n\n### Final Submission\nWeights from the bounding box model and the model3 were used for inference. A combination of the  resnest50d and seresnext50, which had 6 and 4 respectively sets of weights from model 3 trained over the full dataset. \nI found it kind on RAM to collect batches of studies in uint8 format, and normalize on GPU; and inference through the CNN was performed in chunks of 32 * 2.5d images at a time. \n\nInference time was ~ 4.5 hours. \n",
    "2007684": "Thanks for the write up, congrats! \n\nSo when you are creating the bounding boxes of the C1-C7 vertebrae by taking the outer limit of the segmentation, was the algorithm ever thrown off by stray pixels that can appear anywhere on the image? Do you have a preprocessing step to ignore pixel islands that are less than a number? I did something similar but I calculated the center of mass and standard deviation of the distribution of the mask and then use it to create bounding boxes (bounding cubes actually). But the boxes are not as tight  around the spine as the images you showed above.",
    "2007995": "Congratulations! Thank you for sharing your idea!!!",
    "2007873": "Congratulations! Amazing Work",
    "2007784": "Congratulations!   In your Model 2, do you train end-to-end a  single 2.5D CNN (32 copies of a CNN with same weights feeding to 1d RNN)?  Thx.",
    "2007607": "Congratulations!, Honestly, I did not get much of it, but seems like a really cool solution",
    "2007577": "Congrats! You have a completely different pipeline, I learned a lot!",
    "2008159": "Amazing work and a great result to culminate the effort @darraghdog! Keep it up and hearty congratulations!",
    "2007595": "Congratulations! The vertebra ratio idea is very cool. ",
    "2008430": "Congratulations!! Thanks for sharing\n"
  }
}