{
  "id": 362687,
  "title": "13th Place Solution",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362687",
  "author_name": "Bardia Khosravi",
  "post_date": "2022-10-28T14:50:24.985000",
  "votes": 32,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Thanks to all the organizers for collecting and annotating this rich pool of C-spine CTs and also congratulations to all the winners for their amazing work. I should also thank my teammate <a href=\"https://www.kaggle.com/braderickson\" target=\"_blank\">@braderickson</a> for all his clinical and technical insights; this solution would not have been possible without him.</p>\n<p>It was my first serious Kaggle competition, and I really enjoyed it while making some big mistakes, and hopefully, I can learn from them for my future competitions [more at the end].</p>\n<h2>Summary</h2>\n<p>From the beginning, we decided to focus on a 3D pipeline as there are some queues in each slice (like the nutrient canals) that might throw off a 2D model (just our initial idea). Our pipeline had a very simple flow, segment C1-C7 vertebrae, and a 3D classifier that gets each vertebra volume and predicts if it is fractured or not. </p>\n<h2>Segmentation</h2>\n<p>This was a bit challenging; only 87 samples had vertebra segmentation masks and based on the literature around VerSe challenge this was not enough to have a robust segmentation model. We used VerSe challenge dataset in addition to CTSpine1K dataset to train a 3D SwinUNETR model (from MONAI). For the segmentation task, we used patches of [128, 128, 128] with [1, 1, 1] pixel spacing and upsampled the masks to the images' original dimensions. We used the 87 annotated images as the test set to evaluate the robustness of our model. This model reached a mean DSC of ~0.90 on the test set, making it a viable option to be used as a standalone model (no CV for segmentation). </p>\n<p>One thing that we did, was to rotate the vertebrae to have the anterior and posterior center of masses align in a horizontal line; this helped fix the hyper-flexed neck positioning that some patients had.</p>\n<p>Mask post-processing was applied as serial closing and dilation (to get rid of stray pixels) and getting the 7 largest connected components, one for each vertebra. At this stage, we were able to get a clean vertebra volume. But as others mentioned, there are some overlaps, and two consecutive vertebrae might be visible at the top or bottom of the volume. To overcome this, we blacked out other vertebrae. For example, in the following image of a C1 vertebra, we blacked out the odontoid process of the C2 vertebra. This, in our experiments, had the best performance compared to other approaches.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5189553%2Fbe1f7efe85201de67301b964d03d1c5c%2FScreenshot_4.png?generation=1666964924580012&amp;alt=media\" alt=\"\"></p>\n<p>With all the resampling, upsampling of the mask, and post-processing this stage took ~4.5 hours of the total runtime.</p>\n<h2>Classifier</h2>\n<p>We got the volume dimension of each vertebra and found that &gt;90% of samples had a volume less than 256x256x64. So we chose this size as the input to our 3D classifier. The most challenging part of the pipeline was training a robust 3D classifier. MONAI has a great API for creating 3D versions of famous 2D models like EfficientNets (only v1s), DensNets and ResNeXts. We finally decided to use an ensemble of EfficientNetB4s and DenseNet209 models as they had the best performance with our CV and reached an AUC of &gt;.90 on our CV. We added horizontal flipping as TTA, and it helped boost the performance by 0.02.</p>\n<p>For calculating the patient overall labels, we used conditional probabilities of the highest three fracture probabilities (as more than 95% of samples had &lt;=3 fractures). Compared to setting the max value as the patients' overall label, this helped by 0.02 score improvement. We rounded the probabilities &lt;0.05a nd &gt;0.90 to zero and one respectively as it did help the public score by 0.02. This was our biggest mistake! After finishing the competition we ran our models without this thresholding, and we got a public score of 0.25 (0.03 improvement).</p>\n<p>The 10 classifiers took ~2.5 hours to run, which was pretty fast.</p>\n<h2>Lessons Learnt</h2>\n<p>Here are some remarks about my first Kaggle competition experience, hope this might help others:</p>\n<ol>\n<li>Start Early: If this is your first time, starting soon helps with iterating with different ideas and models.</li>\n<li>Do Your Research: Search how others have done this before and how you can improve on that. Also, search for related external datasets.</li>\n<li>Implement and Use the Competition Metric: This will give you an overall idea of where you are in the leaderboard without needing to submit so often.</li>\n<li>Do CV: Though it is tempting to have your hyperparameters optimized in a single fold and get the \"Best\" results on that before testing it on others, resist that temptation. Ensembling all folds is almost always helpful, so the sooner you reach this point, the better.</li>\n<li>Don't Overfit: This is the hardest in my opinion. Try to optimize everything on your CV rather than the public test set. Don't submit just to see if changing a small parameter (like thresholds) will improve your public test score.</li>\n<li>Have Fun!</li>\n</ol>\n<h2>Acknowledgements</h2>\n<p>I used MONAI and Pytorch-Lightning for all the experiments, and it made my life much easier as iterating between models and hyperparameters was made so simple. Also, thanks to all the Kagglers who participated in this competition; I learned a lot from their codes and discussions. </p>",
  "messages": [
    {
      "id": 2007859,
      "postDate": "2022-10-28T14:50:24.987Z",
      "content": "<p>Thanks to all the organizers for collecting and annotating this rich pool of C-spine CTs and also congratulations to all the winners for their amazing work. I should also thank my teammate <a href=\"https://www.kaggle.com/braderickson\" target=\"_blank\">@braderickson</a> for all his clinical and technical insights; this solution would not have been possible without him.</p>\n<p>It was my first serious Kaggle competition, and I really enjoyed it while making some big mistakes, and hopefully, I can learn from them for my future competitions [more at the end].</p>\n<h2>Summary</h2>\n<p>From the beginning, we decided to focus on a 3D pipeline as there are some queues in each slice (like the nutrient canals) that might throw off a 2D model (just our initial idea). Our pipeline had a very simple flow, segment C1-C7 vertebrae, and a 3D classifier that gets each vertebra volume and predicts if it is fractured or not. </p>\n<h2>Segmentation</h2>\n<p>This was a bit challenging; only 87 samples had vertebra segmentation masks and based on the literature around VerSe challenge this was not enough to have a robust segmentation model. We used VerSe challenge dataset in addition to CTSpine1K dataset to train a 3D SwinUNETR model (from MONAI). For the segmentation task, we used patches of [128, 128, 128] with [1, 1, 1] pixel spacing and upsampled the masks to the images' original dimensions. We used the 87 annotated images as the test set to evaluate the robustness of our model. This model reached a mean DSC of ~0.90 on the test set, making it a viable option to be used as a standalone model (no CV for segmentation). </p>\n<p>One thing that we did, was to rotate the vertebrae to have the anterior and posterior center of masses align in a horizontal line; this helped fix the hyper-flexed neck positioning that some patients had.</p>\n<p>Mask post-processing was applied as serial closing and dilation (to get rid of stray pixels) and getting the 7 largest connected components, one for each vertebra. At this stage, we were able to get a clean vertebra volume. But as others mentioned, there are some overlaps, and two consecutive vertebrae might be visible at the top or bottom of the volume. To overcome this, we blacked out other vertebrae. For example, in the following image of a C1 vertebra, we blacked out the odontoid process of the C2 vertebra. This, in our experiments, had the best performance compared to other approaches.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5189553%2Fbe1f7efe85201de67301b964d03d1c5c%2FScreenshot_4.png?generation=1666964924580012&amp;alt=media\" alt=\"\"></p>\n<p>With all the resampling, upsampling of the mask, and post-processing this stage took ~4.5 hours of the total runtime.</p>\n<h2>Classifier</h2>\n<p>We got the volume dimension of each vertebra and found that &gt;90% of samples had a volume less than 256x256x64. So we chose this size as the input to our 3D classifier. The most challenging part of the pipeline was training a robust 3D classifier. MONAI has a great API for creating 3D versions of famous 2D models like EfficientNets (only v1s), DensNets and ResNeXts. We finally decided to use an ensemble of EfficientNetB4s and DenseNet209 models as they had the best performance with our CV and reached an AUC of &gt;.90 on our CV. We added horizontal flipping as TTA, and it helped boost the performance by 0.02.</p>\n<p>For calculating the patient overall labels, we used conditional probabilities of the highest three fracture probabilities (as more than 95% of samples had &lt;=3 fractures). Compared to setting the max value as the patients' overall label, this helped by 0.02 score improvement. We rounded the probabilities &lt;0.05a nd &gt;0.90 to zero and one respectively as it did help the public score by 0.02. This was our biggest mistake! After finishing the competition we ran our models without this thresholding, and we got a public score of 0.25 (0.03 improvement).</p>\n<p>The 10 classifiers took ~2.5 hours to run, which was pretty fast.</p>\n<h2>Lessons Learnt</h2>\n<p>Here are some remarks about my first Kaggle competition experience, hope this might help others:</p>\n<ol>\n<li>Start Early: If this is your first time, starting soon helps with iterating with different ideas and models.</li>\n<li>Do Your Research: Search how others have done this before and how you can improve on that. Also, search for related external datasets.</li>\n<li>Implement and Use the Competition Metric: This will give you an overall idea of where you are in the leaderboard without needing to submit so often.</li>\n<li>Do CV: Though it is tempting to have your hyperparameters optimized in a single fold and get the \"Best\" results on that before testing it on others, resist that temptation. Ensembling all folds is almost always helpful, so the sooner you reach this point, the better.</li>\n<li>Don't Overfit: This is the hardest in my opinion. Try to optimize everything on your CV rather than the public test set. Don't submit just to see if changing a small parameter (like thresholds) will improve your public test score.</li>\n<li>Have Fun!</li>\n</ol>\n<h2>Acknowledgements</h2>\n<p>I used MONAI and Pytorch-Lightning for all the experiments, and it made my life much easier as iterating between models and hyperparameters was made so simple. Also, thanks to all the Kagglers who participated in this competition; I learned a lot from their codes and discussions. </p>",
      "rawMarkdown": "Thanks to all the organizers for collecting and annotating this rich pool of C-spine CTs and also congratulations to all the winners for their amazing work. I should also thank my teammate @braderickson for all his clinical and technical insights; this solution would not have been possible without him.\n\nIt was my first serious Kaggle competition, and I really enjoyed it while making some big mistakes, and hopefully, I can learn from them for my future competitions [more at the end].\n\n## Summary\n\nFrom the beginning, we decided to focus on a 3D pipeline as there are some queues in each slice (like the nutrient canals) that might throw off a 2D model (just our initial idea). Our pipeline had a very simple flow, segment C1-C7 vertebrae, and a 3D classifier that gets each vertebra volume and predicts if it is fractured or not. \n\n## Segmentation\n\nThis was a bit challenging; only 87 samples had vertebra segmentation masks and based on the literature around VerSe challenge this was not enough to have a robust segmentation model. We used VerSe challenge dataset in addition to CTSpine1K dataset to train a 3D SwinUNETR model (from MONAI). For the segmentation task, we used patches of [128, 128, 128] with [1, 1, 1] pixel spacing and upsampled the masks to the images' original dimensions. We used the 87 annotated images as the test set to evaluate the robustness of our model. This model reached a mean DSC of ~0.90 on the test set, making it a viable option to be used as a standalone model (no CV for segmentation). \n\nOne thing that we did, was to rotate the vertebrae to have the anterior and posterior center of masses align in a horizontal line; this helped fix the hyper-flexed neck positioning that some patients had.\n\nMask post-processing was applied as serial closing and dilation (to get rid of stray pixels) and getting the 7 largest connected components, one for each vertebra. At this stage, we were able to get a clean vertebra volume. But as others mentioned, there are some overlaps, and two consecutive vertebrae might be visible at the top or bottom of the volume. To overcome this, we blacked out other vertebrae. For example, in the following image of a C1 vertebra, we blacked out the odontoid process of the C2 vertebra. This, in our experiments, had the best performance compared to other approaches.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5189553%2Fbe1f7efe85201de67301b964d03d1c5c%2FScreenshot_4.png?generation=1666964924580012&alt=media)\n\nWith all the resampling, upsampling of the mask, and post-processing this stage took ~4.5 hours of the total runtime.\n\n## Classifier\n\nWe got the volume dimension of each vertebra and found that >90% of samples had a volume less than 256x256x64. So we chose this size as the input to our 3D classifier. The most challenging part of the pipeline was training a robust 3D classifier. MONAI has a great API for creating 3D versions of famous 2D models like EfficientNets (only v1s), DensNets and ResNeXts. We finally decided to use an ensemble of EfficientNetB4s and DenseNet209 models as they had the best performance with our CV and reached an AUC of >.90 on our CV. We added horizontal flipping as TTA, and it helped boost the performance by 0.02.\n\nFor calculating the patient overall labels, we used conditional probabilities of the highest three fracture probabilities (as more than 95% of samples had <=3 fractures). Compared to setting the max value as the patients' overall label, this helped by 0.02 score improvement. We rounded the probabilities <0.05a nd >0.90 to zero and one respectively as it did help the public score by 0.02. This was our biggest mistake! After finishing the competition we ran our models without this thresholding, and we got a public score of 0.25 (0.03 improvement).\n\nThe 10 classifiers took ~2.5 hours to run, which was pretty fast.\n\n## Lessons Learnt\n\nHere are some remarks about my first Kaggle competition experience, hope this might help others:\n\n1. Start Early: If this is your first time, starting soon helps with iterating with different ideas and models.\n2. Do Your Research: Search how others have done this before and how you can improve on that. Also, search for related external datasets.\n3. Implement and Use the Competition Metric: This will give you an overall idea of where you are in the leaderboard without needing to submit so often.\n4. Do CV: Though it is tempting to have your hyperparameters optimized in a single fold and get the \"Best\" results on that before testing it on others, resist that temptation. Ensembling all folds is almost always helpful, so the sooner you reach this point, the better.\n5. Don't Overfit: This is the hardest in my opinion. Try to optimize everything on your CV rather than the public test set. Don't submit just to see if changing a small parameter (like thresholds) will improve your public test score.\n6. Have Fun!\n\n## Acknowledgements\n\nI used MONAI and Pytorch-Lightning for all the experiments, and it made my life much easier as iterating between models and hyperparameters was made so simple. Also, thanks to all the Kagglers who participated in this competition; I learned a lot from their codes and discussions. ",
      "votes": 32
    },
    {
      "id": 2010967,
      "postDate": "2022-10-31T09:29:36.590Z",
      "content": "<p>That's a pity with the overfit; with the benefit of this learning I'm sure you'll be coming back stronger next time 💪💪💪</p>",
      "rawMarkdown": "That's a pity with the overfit; with the benefit of this learning I'm sure you'll be coming back stronger next time 💪💪💪",
      "votes": 1
    },
    {
      "id": 2010018,
      "postDate": "2022-10-30T13:03:38.737Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": 1
    },
    {
      "id": 2008385,
      "postDate": "2022-10-29T02:25:44.697Z",
      "content": "<p>Congratulations B. Khosravi and thanks for sharing your 13th solution.</p>",
      "rawMarkdown": "Congratulations B. Khosravi and thanks for sharing your 13th solution.",
      "votes": 1
    },
    {
      "id": 2008223,
      "postDate": "2022-10-28T20:47:29.843Z",
      "content": "<p>Thanks Bardia for the clear explanation and congrat for the medal.</p>",
      "rawMarkdown": "Thanks Bardia for the clear explanation and congrat for the medal.",
      "votes": 1
    },
    {
      "id": 2008146,
      "postDate": "2022-10-28T19:23:43.283Z",
      "content": "<p>Great result <a href=\"https://www.kaggle.com/bardiakh\" target=\"_blank\">@bardiakh</a>, appreciate the efforts and congratulations for the result! Well explained too!</p>",
      "rawMarkdown": "Great result @bardiakh, appreciate the efforts and congratulations for the result! Well explained too!",
      "votes": 1
    },
    {
      "id": 2008097,
      "postDate": "2022-10-28T18:35:51.400Z",
      "content": "<p>Great work ! Congrats</p>",
      "rawMarkdown": "Great work ! Congrats",
      "votes": 1
    },
    {
      "id": 2007971,
      "postDate": "2022-10-28T16:06:19.797Z",
      "content": "<p>Congrats! Nice work and great write up!</p>",
      "rawMarkdown": "Congrats! Nice work and great write up!",
      "votes": 1
    },
    {
      "id": 2007956,
      "postDate": "2022-10-28T15:45:23.870Z",
      "content": "<p>Thanks for the writeup. Interesting how the individual and conditional probabilities worked so well. I didn't try that and just assumed having a second stage model to optimize the loss directly would be better. </p>\n<blockquote>\n  <p>We rounded the probabilities &lt;0.05a nd &gt;0.90 to zero and one respectively as it did help the public score by 0.02. This was our biggest mistake!</p>\n</blockquote>\n<p>Yes, quite risky to do this because, because the decrease in loss if you're right is much less than the increase in loss if you're wrong. Interesting that you got a rather significant decrease on public LB. </p>",
      "rawMarkdown": "Thanks for the writeup. Interesting how the individual and conditional probabilities worked so well. I didn't try that and just assumed having a second stage model to optimize the loss directly would be better. \n\n> We rounded the probabilities <0.05a nd >0.90 to zero and one respectively as it did help the public score by 0.02. This was our biggest mistake!\n\nYes, quite risky to do this because, because the decrease in loss if you're right is much less than the increase in loss if you're wrong. Interesting that you got a rather significant decrease on public LB. ",
      "votes": 1,
      "replies": [
        {
          "id": 2007963,
          "postDate": "2022-10-28T15:55:19.543Z",
          "content": "<p>Thanks. </p>\n<p>Learned this the hard way :)</p>",
          "rawMarkdown": "Thanks. \n\nLearned this the hard way :)"
        }
      ]
    },
    {
      "id": 2010721,
      "postDate": "2022-10-31T04:37:05.910Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/bardiakh\" target=\"_blank\">@bardiakh</a> ! For your 3d classifier, could you share your training method? I took lots of time in trying, but failed to get a reasonable 3d classifier.</p>",
      "rawMarkdown": "Thanks for sharing @bardiakh ! For your 3d classifier, could you share your training method? I took lots of time in trying, but failed to get a reasonable 3d classifier."
    },
    {
      "id": 2007911,
      "postDate": "2022-10-28T15:14:12.797Z",
      "content": "<p>Very nice work!  <br>\nWhen you say</p>\n<pre><code>For calculating the patient overall labels, we used conditional probabilities of the highest three fracture probabilities \n</code></pre>\n<p>what do you mean?  The patient overall labels are given and you know which vertebrae you have clipped out.  Or is it the probability that you have identified the vertebrae correctly? </p>",
      "rawMarkdown": "Very nice work!  \nWhen you say\n```\nFor calculating the patient overall labels, we used conditional probabilities of the highest three fracture probabilities \n```\nwhat do you mean?  The patient overall labels are given and you know which vertebrae you have clipped out.  Or is it the probability that you have identified the vertebrae correctly? ",
      "replies": [
        {
          "id": 2007924,
          "postDate": "2022-10-28T15:21:14.950Z",
          "content": "<p>Thanks.</p>\n<p>We have fracture probability for each vertebra. Suppose for one patient that is <code>[0.1, 0.2, 0.1, 0.2, 0.3, 0.6, 0.7]</code>. The three highest values are <code>[03, 0.6, 0.7]</code>. The patient's overall probability was calculated as follows:</p>\n<pre><code>1 - ([1-0.3] * [1-0.6] * [1-0.7]) = 0.916\n</code></pre>",
          "rawMarkdown": "Thanks.\n\nWe have fracture probability for each vertebra. Suppose for one patient that is `[0.1, 0.2, 0.1, 0.2, 0.3, 0.6, 0.7]`. The three highest values are `[03, 0.6, 0.7]`. The patient's overall probability was calculated as follows:\n\n```\n1 - ([1-0.3] * [1-0.6] * [1-0.7]) = 0.916\n```",
          "votes": 2
        },
        {
          "id": 2007940,
          "postDate": "2022-10-28T15:31:02.607Z",
          "content": "<p>Thank you for explaining 😃</p>",
          "rawMarkdown": "Thank you for explaining 😃",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2010967,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2022-10-31T09:29:36.590000",
      "content": "<p>That's a pity with the overfit; with the benefit of this learning I'm sure you'll be coming back stronger next time 💪💪💪</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2010018,
      "author_name": "Samuel Cortinhas",
      "author_url": "",
      "post_date": "2022-10-30T13:03:38.737000",
      "content": "<p>Congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2008385,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2022-10-29T02:25:44.697000",
      "content": "<p>Congratulations B. Khosravi and thanks for sharing your 13th solution.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2008223,
      "author_name": "Behnam Molaee",
      "author_url": "",
      "post_date": "2022-10-28T20:47:29.843000",
      "content": "<p>Thanks Bardia for the clear explanation and congrat for the medal.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2008146,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2022-10-28T19:23:43.283000",
      "content": "<p>Great result <a href=\"https://www.kaggle.com/bardiakh\" target=\"_blank\">@bardiakh</a>, appreciate the efforts and congratulations for the result! Well explained too!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2008097,
      "author_name": "Arnab_Dey",
      "author_url": "",
      "post_date": "2022-10-28T18:35:51.400000",
      "content": "<p>Great work ! Congrats</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2007971,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2022-10-28T16:06:19.797000",
      "content": "<p>Congrats! Nice work and great write up!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2007956,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2022-10-28T15:45:23.870000",
      "content": "<p>Thanks for the writeup. Interesting how the individual and conditional probabilities worked so well. I didn't try that and just assumed having a second stage model to optimize the loss directly would be better. </p>\n<blockquote>\n  <p>We rounded the probabilities &lt;0.05a nd &gt;0.90 to zero and one respectively as it did help the public score by 0.02. This was our biggest mistake!</p>\n</blockquote>\n<p>Yes, quite risky to do this because, because the decrease in loss if you're right is much less than the increase in loss if you're wrong. Interesting that you got a rather significant decrease on public LB. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2007963,
          "author_name": "Bardia Khosravi",
          "author_url": "",
          "post_date": "2022-10-28T15:55:19.543000",
          "content": "<p>Thanks. </p>\n<p>Learned this the hard way :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2010721,
      "author_name": "Yiheng Wang",
      "author_url": "",
      "post_date": "2022-10-31T04:37:05.910000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/bardiakh\" target=\"_blank\">@bardiakh</a> ! For your 3d classifier, could you share your training method? I took lots of time in trying, but failed to get a reasonable 3d classifier.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2007911,
      "author_name": "SolverWorld",
      "author_url": "",
      "post_date": "2022-10-28T15:14:12.797000",
      "content": "<p>Very nice work!  <br>\nWhen you say</p>\n<pre><code>For calculating the patient overall labels, we used conditional probabilities of the highest three fracture probabilities \n</code></pre>\n<p>what do you mean?  The patient overall labels are given and you know which vertebrae you have clipped out.  Or is it the probability that you have identified the vertebrae correctly? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2007924,
          "author_name": "Bardia Khosravi",
          "author_url": "",
          "post_date": "2022-10-28T15:21:14.950000",
          "content": "<p>Thanks.</p>\n<p>We have fracture probability for each vertebra. Suppose for one patient that is <code>[0.1, 0.2, 0.1, 0.2, 0.3, 0.6, 0.7]</code>. The three highest values are <code>[03, 0.6, 0.7]</code>. The patient's overall probability was calculated as follows:</p>\n<pre><code>1 - ([1-0.3] * [1-0.6] * [1-0.7]) = 0.916\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2007940,
          "author_name": "SolverWorld",
          "author_url": "",
          "post_date": "2022-10-28T15:31:02.607000",
          "content": "<p>Thank you for explaining 😃</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2007859": "Thanks to all the organizers for collecting and annotating this rich pool of C-spine CTs and also congratulations to all the winners for their amazing work. I should also thank my teammate @braderickson for all his clinical and technical insights; this solution would not have been possible without him.\n\nIt was my first serious Kaggle competition, and I really enjoyed it while making some big mistakes, and hopefully, I can learn from them for my future competitions [more at the end].\n\n## Summary\n\nFrom the beginning, we decided to focus on a 3D pipeline as there are some queues in each slice (like the nutrient canals) that might throw off a 2D model (just our initial idea). Our pipeline had a very simple flow, segment C1-C7 vertebrae, and a 3D classifier that gets each vertebra volume and predicts if it is fractured or not. \n\n## Segmentation\n\nThis was a bit challenging; only 87 samples had vertebra segmentation masks and based on the literature around VerSe challenge this was not enough to have a robust segmentation model. We used VerSe challenge dataset in addition to CTSpine1K dataset to train a 3D SwinUNETR model (from MONAI). For the segmentation task, we used patches of [128, 128, 128] with [1, 1, 1] pixel spacing and upsampled the masks to the images' original dimensions. We used the 87 annotated images as the test set to evaluate the robustness of our model. This model reached a mean DSC of ~0.90 on the test set, making it a viable option to be used as a standalone model (no CV for segmentation). \n\nOne thing that we did, was to rotate the vertebrae to have the anterior and posterior center of masses align in a horizontal line; this helped fix the hyper-flexed neck positioning that some patients had.\n\nMask post-processing was applied as serial closing and dilation (to get rid of stray pixels) and getting the 7 largest connected components, one for each vertebra. At this stage, we were able to get a clean vertebra volume. But as others mentioned, there are some overlaps, and two consecutive vertebrae might be visible at the top or bottom of the volume. To overcome this, we blacked out other vertebrae. For example, in the following image of a C1 vertebra, we blacked out the odontoid process of the C2 vertebra. This, in our experiments, had the best performance compared to other approaches.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5189553%2Fbe1f7efe85201de67301b964d03d1c5c%2FScreenshot_4.png?generation=1666964924580012&alt=media)\n\nWith all the resampling, upsampling of the mask, and post-processing this stage took ~4.5 hours of the total runtime.\n\n## Classifier\n\nWe got the volume dimension of each vertebra and found that >90% of samples had a volume less than 256x256x64. So we chose this size as the input to our 3D classifier. The most challenging part of the pipeline was training a robust 3D classifier. MONAI has a great API for creating 3D versions of famous 2D models like EfficientNets (only v1s), DensNets and ResNeXts. We finally decided to use an ensemble of EfficientNetB4s and DenseNet209 models as they had the best performance with our CV and reached an AUC of >.90 on our CV. We added horizontal flipping as TTA, and it helped boost the performance by 0.02.\n\nFor calculating the patient overall labels, we used conditional probabilities of the highest three fracture probabilities (as more than 95% of samples had <=3 fractures). Compared to setting the max value as the patients' overall label, this helped by 0.02 score improvement. We rounded the probabilities <0.05a nd >0.90 to zero and one respectively as it did help the public score by 0.02. This was our biggest mistake! After finishing the competition we ran our models without this thresholding, and we got a public score of 0.25 (0.03 improvement).\n\nThe 10 classifiers took ~2.5 hours to run, which was pretty fast.\n\n## Lessons Learnt\n\nHere are some remarks about my first Kaggle competition experience, hope this might help others:\n\n1. Start Early: If this is your first time, starting soon helps with iterating with different ideas and models.\n2. Do Your Research: Search how others have done this before and how you can improve on that. Also, search for related external datasets.\n3. Implement and Use the Competition Metric: This will give you an overall idea of where you are in the leaderboard without needing to submit so often.\n4. Do CV: Though it is tempting to have your hyperparameters optimized in a single fold and get the \"Best\" results on that before testing it on others, resist that temptation. Ensembling all folds is almost always helpful, so the sooner you reach this point, the better.\n5. Don't Overfit: This is the hardest in my opinion. Try to optimize everything on your CV rather than the public test set. Don't submit just to see if changing a small parameter (like thresholds) will improve your public test score.\n6. Have Fun!\n\n## Acknowledgements\n\nI used MONAI and Pytorch-Lightning for all the experiments, and it made my life much easier as iterating between models and hyperparameters was made so simple. Also, thanks to all the Kagglers who participated in this competition; I learned a lot from their codes and discussions. ",
    "2010967": "That's a pity with the overfit; with the benefit of this learning I'm sure you'll be coming back stronger next time 💪💪💪",
    "2010018": "Congratulations!",
    "2008385": "Congratulations B. Khosravi and thanks for sharing your 13th solution.",
    "2008223": "Thanks Bardia for the clear explanation and congrat for the medal.",
    "2008146": "Great result @bardiakh, appreciate the efforts and congratulations for the result! Well explained too!",
    "2008097": "Great work ! Congrats",
    "2007971": "Congrats! Nice work and great write up!",
    "2007956": "Thanks for the writeup. Interesting how the individual and conditional probabilities worked so well. I didn't try that and just assumed having a second stage model to optimize the loss directly would be better. \n\n> We rounded the probabilities <0.05a nd >0.90 to zero and one respectively as it did help the public score by 0.02. This was our biggest mistake!\n\nYes, quite risky to do this because, because the decrease in loss if you're right is much less than the increase in loss if you're wrong. Interesting that you got a rather significant decrease on public LB. ",
    "2010721": "Thanks for sharing @bardiakh ! For your 3d classifier, could you share your training method? I took lots of time in trying, but failed to get a reasonable 3d classifier.",
    "2007911": "Very nice work!  \nWhen you say\n```\nFor calculating the patient overall labels, we used conditional probabilities of the highest three fracture probabilities \n```\nwhat do you mean?  The patient overall labels are given and you know which vertebrae you have clipped out.  Or is it the probability that you have identified the vertebrae correctly? "
  }
}