{
  "id": 362844,
  "title": "12th Place Solution",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362844",
  "author_name": "Takashi Someya",
  "post_date": "2022-10-29T12:11:59.265000",
  "votes": 26,
  "comment_count": 5,
  "views": 0,
  "content": "<p>First of all, I would like to thank the kaggle staff and hosts for organizing and hosting this competition.<br>\nI would also like to thank my teammate <a href=\"https://www.kaggle.com/yuyuki11235\" target=\"_blank\">@yuyuki11235</a>. He suggested and implemented a lot of ideas and was really helpful.<br>\nWe hope to try again to win the gold medal next time.</p>\n<h2>Summary</h2>\n<p>We used a 2-stage approach, extracting the image-level features with CNN models and inferring the patient-level fracture probabilities with sequential models.<br>\nThe slice images were preprocessed with YOLOX cropping, 2.5D approach and a windowing method.<br>\nFor final submission, we used efficietnet-v2-l as a feature extractor and LSTM, GRU, and Conv1d for the sequence model.</p>\n<h2>Preprocess</h2>\n<h5>Bounding box cropping</h5>\n<p>We trained YOLOX-l model to detect vertebral positions.<br>\nBbox labels were created from the outer frame of the segmentation masks (similar to the <a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362643\" target=\"_blank\">3rd place solution</a>).<br>\nFor cropping, we only used the highest confidence bbox in the images with a confidence level of 0.7 or higher.<br>\n20 pixel margin was added to bbox when cropping.</p>\n<h5>Use neighbor slices and windowing</h5>\n<p>As in previous RSNA competitions, we used a 2.5D approach using neighboring slices and a windowing technique.<br>\nFor windowing, we applied the values used in <a href=\"https://arxiv.org/abs/2010.13336\" target=\"_blank\">this paper</a>.</p>\n<ul>\n<li>Standard bone window (w=500, c=1800)</li>\n<li>Gross bone window (w=650, c=400)</li>\n<li>Soft tissue window (w=300, c=80)</li>\n</ul>\n<p>We combined 2.5D and windowing as follows (s denotes a slice index).<br>\nch1 : s-1 slice with Standard bone window<br>\nch2 : s slice with Gross bone window<br>\nch3 : s+1 slice with Soft tissue window</p>\n<p>An example of input images created by preprocessing is shown below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6102861%2Ff1139ecd39345e6e4992a1039aee5797%2Finput_images.png?generation=1667030306195626&amp;alt=media\" alt=\"\"></p>\n<h2>CNN Model (1st stage)</h2>\n<p>Our CNN models were based on <a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">@vslaykovsky</a>'s <a href=\"https://www.kaggle.com/code/vslaykovsky/train-pytorch-effnetv2-baseline-cv-0-49\" target=\"_blank\">public baseline</a>.<br>\nWe trained our models to infer fracture and vertebral positions.<br>\nThe main tips related to training are listed below.</p>\n<h5>Augmentation</h5>\n<ul>\n<li>HorizontalFlip</li>\n<li>GaussianBlur</li>\n<li>GaussNoise</li>\n<li>ShiftScaleRotate</li>\n<li>CoarseDropout</li>\n<li>RandomBrightnessContrast</li>\n<li>OneOf([GridDistortion, ElasticTransform])</li>\n</ul>\n<h5>Label cleaning</h5>\n<p>Since the fracture and vertebral positions per image were not given in this competition, we tried to create the accurate image-level labels by multiple pseudo-labeling steps.<br>\nStarting with <a href=\"https://www.kaggle.com/code/vslaykovsky/train-pytorch-effnetv2-baseline-cv-0-49\" target=\"_blank\">this notebook</a>'s labels, we first made pseudo-labels of fracture and vertebral positions from the oof predictions of efficientnet-b4 and use them as the labels for the next training (with some post-processings such as thresholding).<br>\nBy repeating the pseudo-labeling with larger models, we were able to achieve a better CV.</p>\n<h2>Sequential Model (2nd stage)</h2>\n<p>We trained patient-level classification models using image-level embeddings.<br>\nOur final ensemble were LSTM, GRU, and Conv1d models.<br>\nThe main processing steps are as follows:</p>\n<ul>\n<li>concatenate image-level embedding for each patient (NxD, N: number of slices per patient, D: embedding dimension)</li>\n<li>apply cv2.resize function to fix the length of the sequence (MxD, M: fixed sequence length) (use <a href=\"https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/discussion/194145\" target=\"_blank\">this notebook</a> idea)</li>\n<li>input above features to each model</li>\n</ul>\n<p>Other tricks are as follows:</p>\n<ul>\n<li>Mask only the C1~C7 areas by inferring the bone position (NxD→N'xD, N': number of slices of C1~C7)</li>\n<li>Attention pooling</li>\n<li>Randomly skip slices during training (fixed at skip 3 during inference, see Final Prediction section bellow)</li>\n</ul>\n<h2>Final Prediction</h2>\n<ul>\n<li>Use 1/3 slices as df[df[\"Slices\"] % 3 == 0] to reduce inference time</li>\n<li>YOLOX-l bbox prediciotn (~1h inference time)</li>\n<li>4 out of 5 folds of efficienetnet-v2-l (below 2h inference time per fold)</li>\n<li>image size = (512, 512)</li>\n<li>3 sequential models (LSTM, GRU, Conv1d) with difference seeds (4 seeds for each model)</li>\n<li>fixed sequence length = 224</li>\n</ul>",
  "messages": [
    {
      "id": 2008869,
      "postDate": "2022-10-29T12:11:59.267Z",
      "content": "<p>First of all, I would like to thank the kaggle staff and hosts for organizing and hosting this competition.<br>\nI would also like to thank my teammate <a href=\"https://www.kaggle.com/yuyuki11235\" target=\"_blank\">@yuyuki11235</a>. He suggested and implemented a lot of ideas and was really helpful.<br>\nWe hope to try again to win the gold medal next time.</p>\n<h2>Summary</h2>\n<p>We used a 2-stage approach, extracting the image-level features with CNN models and inferring the patient-level fracture probabilities with sequential models.<br>\nThe slice images were preprocessed with YOLOX cropping, 2.5D approach and a windowing method.<br>\nFor final submission, we used efficietnet-v2-l as a feature extractor and LSTM, GRU, and Conv1d for the sequence model.</p>\n<h2>Preprocess</h2>\n<h5>Bounding box cropping</h5>\n<p>We trained YOLOX-l model to detect vertebral positions.<br>\nBbox labels were created from the outer frame of the segmentation masks (similar to the <a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362643\" target=\"_blank\">3rd place solution</a>).<br>\nFor cropping, we only used the highest confidence bbox in the images with a confidence level of 0.7 or higher.<br>\n20 pixel margin was added to bbox when cropping.</p>\n<h5>Use neighbor slices and windowing</h5>\n<p>As in previous RSNA competitions, we used a 2.5D approach using neighboring slices and a windowing technique.<br>\nFor windowing, we applied the values used in <a href=\"https://arxiv.org/abs/2010.13336\" target=\"_blank\">this paper</a>.</p>\n<ul>\n<li>Standard bone window (w=500, c=1800)</li>\n<li>Gross bone window (w=650, c=400)</li>\n<li>Soft tissue window (w=300, c=80)</li>\n</ul>\n<p>We combined 2.5D and windowing as follows (s denotes a slice index).<br>\nch1 : s-1 slice with Standard bone window<br>\nch2 : s slice with Gross bone window<br>\nch3 : s+1 slice with Soft tissue window</p>\n<p>An example of input images created by preprocessing is shown below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6102861%2Ff1139ecd39345e6e4992a1039aee5797%2Finput_images.png?generation=1667030306195626&amp;alt=media\" alt=\"\"></p>\n<h2>CNN Model (1st stage)</h2>\n<p>Our CNN models were based on <a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">@vslaykovsky</a>'s <a href=\"https://www.kaggle.com/code/vslaykovsky/train-pytorch-effnetv2-baseline-cv-0-49\" target=\"_blank\">public baseline</a>.<br>\nWe trained our models to infer fracture and vertebral positions.<br>\nThe main tips related to training are listed below.</p>\n<h5>Augmentation</h5>\n<ul>\n<li>HorizontalFlip</li>\n<li>GaussianBlur</li>\n<li>GaussNoise</li>\n<li>ShiftScaleRotate</li>\n<li>CoarseDropout</li>\n<li>RandomBrightnessContrast</li>\n<li>OneOf([GridDistortion, ElasticTransform])</li>\n</ul>\n<h5>Label cleaning</h5>\n<p>Since the fracture and vertebral positions per image were not given in this competition, we tried to create the accurate image-level labels by multiple pseudo-labeling steps.<br>\nStarting with <a href=\"https://www.kaggle.com/code/vslaykovsky/train-pytorch-effnetv2-baseline-cv-0-49\" target=\"_blank\">this notebook</a>'s labels, we first made pseudo-labels of fracture and vertebral positions from the oof predictions of efficientnet-b4 and use them as the labels for the next training (with some post-processings such as thresholding).<br>\nBy repeating the pseudo-labeling with larger models, we were able to achieve a better CV.</p>\n<h2>Sequential Model (2nd stage)</h2>\n<p>We trained patient-level classification models using image-level embeddings.<br>\nOur final ensemble were LSTM, GRU, and Conv1d models.<br>\nThe main processing steps are as follows:</p>\n<ul>\n<li>concatenate image-level embedding for each patient (NxD, N: number of slices per patient, D: embedding dimension)</li>\n<li>apply cv2.resize function to fix the length of the sequence (MxD, M: fixed sequence length) (use <a href=\"https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/discussion/194145\" target=\"_blank\">this notebook</a> idea)</li>\n<li>input above features to each model</li>\n</ul>\n<p>Other tricks are as follows:</p>\n<ul>\n<li>Mask only the C1~C7 areas by inferring the bone position (NxD→N'xD, N': number of slices of C1~C7)</li>\n<li>Attention pooling</li>\n<li>Randomly skip slices during training (fixed at skip 3 during inference, see Final Prediction section bellow)</li>\n</ul>\n<h2>Final Prediction</h2>\n<ul>\n<li>Use 1/3 slices as df[df[\"Slices\"] % 3 == 0] to reduce inference time</li>\n<li>YOLOX-l bbox prediciotn (~1h inference time)</li>\n<li>4 out of 5 folds of efficienetnet-v2-l (below 2h inference time per fold)</li>\n<li>image size = (512, 512)</li>\n<li>3 sequential models (LSTM, GRU, Conv1d) with difference seeds (4 seeds for each model)</li>\n<li>fixed sequence length = 224</li>\n</ul>",
      "rawMarkdown": "First of all, I would like to thank the kaggle staff and hosts for organizing and hosting this competition.\nI would also like to thank my teammate @yuyuki11235. He suggested and implemented a lot of ideas and was really helpful.\nWe hope to try again to win the gold medal next time.\n\n## Summary\nWe used a 2-stage approach, extracting the image-level features with CNN models and inferring the patient-level fracture probabilities with sequential models.\nThe slice images were preprocessed with YOLOX cropping, 2.5D approach and a windowing method.\nFor final submission, we used efficietnet-v2-l as a feature extractor and LSTM, GRU, and Conv1d for the sequence model.\n\n## Preprocess\n##### Bounding box cropping\nWe trained YOLOX-l model to detect vertebral positions.\nBbox labels were created from the outer frame of the segmentation masks (similar to the [3rd place solution](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362643)).\nFor cropping, we only used the highest confidence bbox in the images with a confidence level of 0.7 or higher.\n20 pixel margin was added to bbox when cropping.\n\n##### Use neighbor slices and windowing\nAs in previous RSNA competitions, we used a 2.5D approach using neighboring slices and a windowing technique.\nFor windowing, we applied the values used in [this paper](https://arxiv.org/abs/2010.13336).\n- Standard bone window (w=500, c=1800)\n- Gross bone window (w=650, c=400)\n- Soft tissue window (w=300, c=80)\n\nWe combined 2.5D and windowing as follows (s denotes a slice index).\nch1 : s-1 slice with Standard bone window\nch2 : s slice with Gross bone window\nch3 : s+1 slice with Soft tissue window\n\nAn example of input images created by preprocessing is shown below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6102861%2Ff1139ecd39345e6e4992a1039aee5797%2Finput_images.png?generation=1667030306195626&alt=media)\n\n## CNN Model (1st stage)\nOur CNN models were based on @vslaykovsky's [public baseline](https://www.kaggle.com/code/vslaykovsky/train-pytorch-effnetv2-baseline-cv-0-49).\nWe trained our models to infer fracture and vertebral positions.\nThe main tips related to training are listed below.\n\n##### Augmentation\n- HorizontalFlip\n- GaussianBlur\n- GaussNoise\n- ShiftScaleRotate\n- CoarseDropout\n- RandomBrightnessContrast\n- OneOf([GridDistortion, ElasticTransform])\n\n##### Label cleaning\nSince the fracture and vertebral positions per image were not given in this competition, we tried to create the accurate image-level labels by multiple pseudo-labeling steps.\nStarting with [this notebook](https://www.kaggle.com/code/vslaykovsky/train-pytorch-effnetv2-baseline-cv-0-49)'s labels, we first made pseudo-labels of fracture and vertebral positions from the oof predictions of efficientnet-b4 and use them as the labels for the next training (with some post-processings such as thresholding).\nBy repeating the pseudo-labeling with larger models, we were able to achieve a better CV.\n\n## Sequential Model (2nd stage)\nWe trained patient-level classification models using image-level embeddings.\nOur final ensemble were LSTM, GRU, and Conv1d models.\nThe main processing steps are as follows:\n- concatenate image-level embedding for each patient (NxD, N: number of slices per patient, D: embedding dimension)\n- apply cv2.resize function to fix the length of the sequence (MxD, M: fixed sequence length) (use [this notebook](https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/discussion/194145) idea)\n- input above features to each model\n\nOther tricks are as follows:\n- Mask only the C1~C7 areas by inferring the bone position (NxD→N'xD, N': number of slices of C1~C7)\n- Attention pooling\n- Randomly skip slices during training (fixed at skip 3 during inference, see Final Prediction section bellow)\n\n## Final Prediction\n- Use 1/3 slices as df[df[\"Slices\"] % 3 == 0] to reduce inference time\n- YOLOX-l bbox prediciotn (~1h inference time)\n- 4 out of 5 folds of efficienetnet-v2-l (below 2h inference time per fold)\n- image size = (512, 512)\n- 3 sequential models (LSTM, GRU, Conv1d) with difference seeds (4 seeds for each model)\n- fixed sequence length = 224\n",
      "votes": 26
    },
    {
      "id": 2009081,
      "postDate": "2022-10-29T16:10:19.153Z",
      "content": "<p>Congrats on your great finish!<br>\nI had a similar approach except for YOLO box cropping. Did you try without that? How much of an improvement do you get from it?</p>",
      "rawMarkdown": "Congrats on your great finish!\nI had a similar approach except for YOLO box cropping. Did you try without that? How much of an improvement do you get from it?",
      "votes": 1,
      "replies": [
        {
          "id": 2009113,
          "postDate": "2022-10-29T16:47:39.087Z",
          "content": "<p>Congrats on your silver medal too!<br>\nIn our experiments, bbox cropping improved CV by 0.05 or more.</p>",
          "rawMarkdown": "Congrats on your silver medal too!\nIn our experiments, bbox cropping improved CV by 0.05 or more.",
          "votes": 2,
          "replies": [
            {
              "id": 2180902,
              "postDate": "2023-03-14T06:55:28.053Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2046283,
      "postDate": "2022-11-28T04:41:08.913Z",
      "content": "<p>congratulations! How much did label cleaning increase LB? My team overfitted the CV due to cleaning, so please let me know if there is anything I should be careful about.</p>",
      "rawMarkdown": "congratulations! How much did label cleaning increase LB? My team overfitted the CV due to cleaning, so please let me know if there is anything I should be careful about."
    },
    {
      "id": 2009156,
      "postDate": "2022-10-29T17:40:48.857Z",
      "content": "<p>Congratulations!  Can you give more details on your Phase 2 sequential models?</p>",
      "rawMarkdown": "Congratulations!  Can you give more details on your Phase 2 sequential models?"
    }
  ],
  "comments": [
    {
      "id": 2009081,
      "author_name": "Yerram Varun",
      "author_url": "",
      "post_date": "2022-10-29T16:10:19.153000",
      "content": "<p>Congrats on your great finish!<br>\nI had a similar approach except for YOLO box cropping. Did you try without that? How much of an improvement do you get from it?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2009113,
          "author_name": "Takashi Someya",
          "author_url": "",
          "post_date": "2022-10-29T16:47:39.087000",
          "content": "<p>Congrats on your silver medal too!<br>\nIn our experiments, bbox cropping improved CV by 0.05 or more.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2180902,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-14T06:55:28.053000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2046283,
      "author_name": "patriot",
      "author_url": "",
      "post_date": "2022-11-28T04:41:08.913000",
      "content": "<p>congratulations! How much did label cleaning increase LB? My team overfitted the CV due to cleaning, so please let me know if there is anything I should be careful about.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2009156,
      "author_name": "SolverWorld",
      "author_url": "",
      "post_date": "2022-10-29T17:40:48.857000",
      "content": "<p>Congratulations!  Can you give more details on your Phase 2 sequential models?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2008869": "First of all, I would like to thank the kaggle staff and hosts for organizing and hosting this competition.\nI would also like to thank my teammate @yuyuki11235. He suggested and implemented a lot of ideas and was really helpful.\nWe hope to try again to win the gold medal next time.\n\n## Summary\nWe used a 2-stage approach, extracting the image-level features with CNN models and inferring the patient-level fracture probabilities with sequential models.\nThe slice images were preprocessed with YOLOX cropping, 2.5D approach and a windowing method.\nFor final submission, we used efficietnet-v2-l as a feature extractor and LSTM, GRU, and Conv1d for the sequence model.\n\n## Preprocess\n##### Bounding box cropping\nWe trained YOLOX-l model to detect vertebral positions.\nBbox labels were created from the outer frame of the segmentation masks (similar to the [3rd place solution](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362643)).\nFor cropping, we only used the highest confidence bbox in the images with a confidence level of 0.7 or higher.\n20 pixel margin was added to bbox when cropping.\n\n##### Use neighbor slices and windowing\nAs in previous RSNA competitions, we used a 2.5D approach using neighboring slices and a windowing technique.\nFor windowing, we applied the values used in [this paper](https://arxiv.org/abs/2010.13336).\n- Standard bone window (w=500, c=1800)\n- Gross bone window (w=650, c=400)\n- Soft tissue window (w=300, c=80)\n\nWe combined 2.5D and windowing as follows (s denotes a slice index).\nch1 : s-1 slice with Standard bone window\nch2 : s slice with Gross bone window\nch3 : s+1 slice with Soft tissue window\n\nAn example of input images created by preprocessing is shown below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6102861%2Ff1139ecd39345e6e4992a1039aee5797%2Finput_images.png?generation=1667030306195626&alt=media)\n\n## CNN Model (1st stage)\nOur CNN models were based on @vslaykovsky's [public baseline](https://www.kaggle.com/code/vslaykovsky/train-pytorch-effnetv2-baseline-cv-0-49).\nWe trained our models to infer fracture and vertebral positions.\nThe main tips related to training are listed below.\n\n##### Augmentation\n- HorizontalFlip\n- GaussianBlur\n- GaussNoise\n- ShiftScaleRotate\n- CoarseDropout\n- RandomBrightnessContrast\n- OneOf([GridDistortion, ElasticTransform])\n\n##### Label cleaning\nSince the fracture and vertebral positions per image were not given in this competition, we tried to create the accurate image-level labels by multiple pseudo-labeling steps.\nStarting with [this notebook](https://www.kaggle.com/code/vslaykovsky/train-pytorch-effnetv2-baseline-cv-0-49)'s labels, we first made pseudo-labels of fracture and vertebral positions from the oof predictions of efficientnet-b4 and use them as the labels for the next training (with some post-processings such as thresholding).\nBy repeating the pseudo-labeling with larger models, we were able to achieve a better CV.\n\n## Sequential Model (2nd stage)\nWe trained patient-level classification models using image-level embeddings.\nOur final ensemble were LSTM, GRU, and Conv1d models.\nThe main processing steps are as follows:\n- concatenate image-level embedding for each patient (NxD, N: number of slices per patient, D: embedding dimension)\n- apply cv2.resize function to fix the length of the sequence (MxD, M: fixed sequence length) (use [this notebook](https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/discussion/194145) idea)\n- input above features to each model\n\nOther tricks are as follows:\n- Mask only the C1~C7 areas by inferring the bone position (NxD→N'xD, N': number of slices of C1~C7)\n- Attention pooling\n- Randomly skip slices during training (fixed at skip 3 during inference, see Final Prediction section bellow)\n\n## Final Prediction\n- Use 1/3 slices as df[df[\"Slices\"] % 3 == 0] to reduce inference time\n- YOLOX-l bbox prediciotn (~1h inference time)\n- 4 out of 5 folds of efficienetnet-v2-l (below 2h inference time per fold)\n- image size = (512, 512)\n- 3 sequential models (LSTM, GRU, Conv1d) with difference seeds (4 seeds for each model)\n- fixed sequence length = 224\n",
    "2009081": "Congrats on your great finish!\nI had a similar approach except for YOLO box cropping. Did you try without that? How much of an improvement do you get from it?",
    "2046283": "congratulations! How much did label cleaning increase LB? My team overfitted the CV due to cleaning, so please let me know if there is anything I should be careful about.",
    "2009156": "Congratulations!  Can you give more details on your Phase 2 sequential models?"
  }
}