{
  "id": 364837,
  "title": "4th place solution, CSN is all you need for 3D",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/364837",
  "author_name": "Selim Seferbekov",
  "post_date": "2022-11-08T14:28:16.586000",
  "votes": 34,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I started working on this challenge really late, made first sub 10 days before deadline, so could not run a lot of experiments and used codebase/tricks from  <a href=\"https://www.kaggle.com/competitions/uw-madison-gi-tract-image-segmentation\" target=\"_blank\">GI tract segmentation challenge</a> and NASA comet detection challenge.</p>\n<h3>TLDR</h3>\n<p>Two stage approach, 3D segmentation + 3D classification for each vertebra crop</p>\n<h3>The main trick</h3>\n<p>A lot of  failed attempts to use 3D networks for classification are related to overfitting . That's the case for pure 3D convolutions which usually are not needed and we can easily work with architectures that achieve state of the art performance on video/action classification datasets. <br>\nSo the options are:</p>\n<ul>\n<li>inflate some state of the art 2d nets, like effnets but start using 3Ds from 2nd or 3rd block . It works but still tends to overfit and  imagenet pretraining is not the best in this case</li>\n<li>use 2D networks with LSTM/ConvLSTM/Transformer heads. It works and allows to use a lot of architectures, the main issue here is that 3D nature is considered only on the latest stage.</li>\n<li>use pretrained networks from ig65m/kinetics datasest. That gave the best results for me. </li>\n</ul>\n<p>From <a href=\"https://paperswithcode.com/sota/action-classification-on-kinetics-400\" target=\"_blank\">https://paperswithcode.com/sota/action-classification-on-kinetics-400</a> it is clearly seen that the best convolutional architecture for video classifcation in 2022 is still <a href=\"https://arxiv.org/abs/1904.02811\" target=\"_blank\">ir-CSN-152</a>. Even though transformers achieve higher scores on Kinetics datasets they don't perform as good on small datasets like RSNA and CSNs are the best for small sized datasets and if one needs fast and accurate 3D segmentation and/or classification.</p>\n<p>I used mmaction2's implementation of CSNs.</p>\n<h3>Segmentation</h3>\n<ul>\n<li>UNet like multiclass segmentor</li>\n<li>encoder: <strong>ir-CSN-50</strong></li>\n<li>decoder: standard unet decoder with nn upsampling but with <a href=\"https://arxiv.org/abs/1711.11248v3\" target=\"_blank\">(2+1)d convolutions</a></li>\n<li>pure 3d convolutions in decoder lead to NaNs in amp training</li>\n</ul>\n<p><strong>Training</strong>:</p>\n<ul>\n<li>2x subsampling(just slice ::2) by z axis</li>\n<li>2x linear downsampled images</li>\n<li>memory mapping to reduce IO/CPU overhead</li>\n<li>AdamW + wd, cosine LR annealing</li>\n<li>Loss: focal-jaccard optimized loss for faster multiclass jaccard computation</li>\n<li>2D augs, ReplayCompose in Albumentations library, lightweight geometric augs + hflip</li>\n</ul>\n<h3>Classification</h3>\n<ul>\n<li>4 folds of ir(ip)-CSN-152 with global max pooling</li>\n<li>multilabel 8 classes, if less than 30% of vertebra is visible then make vertebra's label 0 and recompute overall label, otherwise keep the same as provided</li>\n</ul>\n<p><strong>Training</strong></p>\n<ul>\n<li>3 channel input (img, img, integer encoded segmentation mask</li>\n<li>2x subsampling(just slice ::2) by z axis</li>\n<li>using 40 slices(80 in original data) around each vertebra</li>\n<li>crops were resized to 256x256</li>\n<li>BCE loss</li>\n<li>target metric as validation</li>\n<li>AdamW + wd, cosine LR annealing</li>\n<li>Augmentations: 2D augs, replay compose in Albumentations library, flips, rotations, geometric. Needed to make more augs compared to segmentation pipeline</li>\n</ul>\n<h3>Code</h3>\n<ul>\n<li>inference kernel and weights: <a href=\"https://www.kaggle.com/code/selimsef/rsna-csn-segmentor-classifier\" target=\"_blank\">https://www.kaggle.com/code/selimsef/rsna-csn-segmentor-classifier</a></li>\n<li>github code for training: <a href=\"https://github.com/selimsef/rsna_cervical_fracture/\" target=\"_blank\">https://github.com/selimsef/rsna_cervical_fracture/</a></li>\n</ul>",
  "messages": [
    {
      "id": 2021867,
      "postDate": "2022-11-08T14:28:16.587Z",
      "content": "<p>I started working on this challenge really late, made first sub 10 days before deadline, so could not run a lot of experiments and used codebase/tricks from  <a href=\"https://www.kaggle.com/competitions/uw-madison-gi-tract-image-segmentation\" target=\"_blank\">GI tract segmentation challenge</a> and NASA comet detection challenge.</p>\n<h3>TLDR</h3>\n<p>Two stage approach, 3D segmentation + 3D classification for each vertebra crop</p>\n<h3>The main trick</h3>\n<p>A lot of  failed attempts to use 3D networks for classification are related to overfitting . That's the case for pure 3D convolutions which usually are not needed and we can easily work with architectures that achieve state of the art performance on video/action classification datasets. <br>\nSo the options are:</p>\n<ul>\n<li>inflate some state of the art 2d nets, like effnets but start using 3Ds from 2nd or 3rd block . It works but still tends to overfit and  imagenet pretraining is not the best in this case</li>\n<li>use 2D networks with LSTM/ConvLSTM/Transformer heads. It works and allows to use a lot of architectures, the main issue here is that 3D nature is considered only on the latest stage.</li>\n<li>use pretrained networks from ig65m/kinetics datasest. That gave the best results for me. </li>\n</ul>\n<p>From <a href=\"https://paperswithcode.com/sota/action-classification-on-kinetics-400\" target=\"_blank\">https://paperswithcode.com/sota/action-classification-on-kinetics-400</a> it is clearly seen that the best convolutional architecture for video classifcation in 2022 is still <a href=\"https://arxiv.org/abs/1904.02811\" target=\"_blank\">ir-CSN-152</a>. Even though transformers achieve higher scores on Kinetics datasets they don't perform as good on small datasets like RSNA and CSNs are the best for small sized datasets and if one needs fast and accurate 3D segmentation and/or classification.</p>\n<p>I used mmaction2's implementation of CSNs.</p>\n<h3>Segmentation</h3>\n<ul>\n<li>UNet like multiclass segmentor</li>\n<li>encoder: <strong>ir-CSN-50</strong></li>\n<li>decoder: standard unet decoder with nn upsampling but with <a href=\"https://arxiv.org/abs/1711.11248v3\" target=\"_blank\">(2+1)d convolutions</a></li>\n<li>pure 3d convolutions in decoder lead to NaNs in amp training</li>\n</ul>\n<p><strong>Training</strong>:</p>\n<ul>\n<li>2x subsampling(just slice ::2) by z axis</li>\n<li>2x linear downsampled images</li>\n<li>memory mapping to reduce IO/CPU overhead</li>\n<li>AdamW + wd, cosine LR annealing</li>\n<li>Loss: focal-jaccard optimized loss for faster multiclass jaccard computation</li>\n<li>2D augs, ReplayCompose in Albumentations library, lightweight geometric augs + hflip</li>\n</ul>\n<h3>Classification</h3>\n<ul>\n<li>4 folds of ir(ip)-CSN-152 with global max pooling</li>\n<li>multilabel 8 classes, if less than 30% of vertebra is visible then make vertebra's label 0 and recompute overall label, otherwise keep the same as provided</li>\n</ul>\n<p><strong>Training</strong></p>\n<ul>\n<li>3 channel input (img, img, integer encoded segmentation mask</li>\n<li>2x subsampling(just slice ::2) by z axis</li>\n<li>using 40 slices(80 in original data) around each vertebra</li>\n<li>crops were resized to 256x256</li>\n<li>BCE loss</li>\n<li>target metric as validation</li>\n<li>AdamW + wd, cosine LR annealing</li>\n<li>Augmentations: 2D augs, replay compose in Albumentations library, flips, rotations, geometric. Needed to make more augs compared to segmentation pipeline</li>\n</ul>\n<h3>Code</h3>\n<ul>\n<li>inference kernel and weights: <a href=\"https://www.kaggle.com/code/selimsef/rsna-csn-segmentor-classifier\" target=\"_blank\">https://www.kaggle.com/code/selimsef/rsna-csn-segmentor-classifier</a></li>\n<li>github code for training: <a href=\"https://github.com/selimsef/rsna_cervical_fracture/\" target=\"_blank\">https://github.com/selimsef/rsna_cervical_fracture/</a></li>\n</ul>",
      "rawMarkdown": "I started working on this challenge really late, made first sub 10 days before deadline, so could not run a lot of experiments and used codebase/tricks from  [GI tract segmentation challenge](https://www.kaggle.com/competitions/uw-madison-gi-tract-image-segmentation) and NASA comet detection challenge.\n\n###TLDR\nTwo stage approach, 3D segmentation + 3D classification for each vertebra crop\n\n###The main trick\nA lot of  failed attempts to use 3D networks for classification are related to overfitting . That's the case for pure 3D convolutions which usually are not needed and we can easily work with architectures that achieve state of the art performance on video/action classification datasets. \nSo the options are:\n- inflate some state of the art 2d nets, like effnets but start using 3Ds from 2nd or 3rd block . It works but still tends to overfit and  imagenet pretraining is not the best in this case\n- use 2D networks with LSTM/ConvLSTM/Transformer heads. It works and allows to use a lot of architectures, the main issue here is that 3D nature is considered only on the latest stage.\n- use pretrained networks from ig65m/kinetics datasest. That gave the best results for me. \n\n\n\n \nFrom https://paperswithcode.com/sota/action-classification-on-kinetics-400 it is clearly seen that the best convolutional architecture for video classifcation in 2022 is still [ir-CSN-152](https://arxiv.org/abs/1904.02811). Even though transformers achieve higher scores on Kinetics datasets they don't perform as good on small datasets like RSNA and CSNs are the best for small sized datasets and if one needs fast and accurate 3D segmentation and/or classification.\n\nI used mmaction2's implementation of CSNs.\n\n### Segmentation\n- UNet like multiclass segmentor\n- encoder: **ir-CSN-50**\n- decoder: standard unet decoder with nn upsampling but with [(2+1)d convolutions](https://arxiv.org/abs/1711.11248v3)\n- pure 3d convolutions in decoder lead to NaNs in amp training\n\n**Training**:\n\n- 2x subsampling(just slice ::2) by z axis\n- 2x linear downsampled images\n- memory mapping to reduce IO/CPU overhead\n- AdamW + wd, cosine LR annealing\n- Loss: focal-jaccard optimized loss for faster multiclass jaccard computation\n- 2D augs, ReplayCompose in Albumentations library, lightweight geometric augs + hflip\n\n### Classification\n\n- 4 folds of ir(ip)-CSN-152 with global max pooling\n- multilabel 8 classes, if less than 30% of vertebra is visible then make vertebra's label 0 and recompute overall label, otherwise keep the same as provided\n\n**Training**\n\n- 3 channel input (img, img, integer encoded segmentation mask\n- 2x subsampling(just slice ::2) by z axis\n- using 40 slices(80 in original data) around each vertebra\n- crops were resized to 256x256\n- BCE loss\n- target metric as validation\n- AdamW + wd, cosine LR annealing\n- Augmentations: 2D augs, replay compose in Albumentations library, flips, rotations, geometric. Needed to make more augs compared to segmentation pipeline\n\n### Code\n- inference kernel and weights: https://www.kaggle.com/code/selimsef/rsna-csn-segmentor-classifier\n- github code for training: https://github.com/selimsef/rsna_cervical_fracture/",
      "votes": 34
    },
    {
      "id": 2060896,
      "postDate": "2022-12-10T14:36:34.997Z",
      "content": "<p><a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a>, could you explain why the input of the classifiers is in the form of 2x images + masks? Below is the part of the code from <code>DatasetCrops.__getitem__()</code> I am asking about.</p>\n<pre><code>images = np.concatenate([images, images, masks], axis=-)\n</code></pre>",
      "rawMarkdown": "@selimsef, could you explain why the input of the classifiers is in the form of 2x images + masks? Below is the part of the code from `DatasetCrops.__getitem__()` I am asking about.\n\n```py\nimages = np.concatenate([images, images, masks], axis=-1)\n```",
      "replies": [
        {
          "id": 2062151,
          "postDate": "2022-12-11T19:12:39.093Z",
          "content": "<p><a href=\"https://www.kaggle.com/mariuszwisniewski\" target=\"_blank\">@mariuszwisniewski</a> I used integer encoded predicted segmentation mask as an additional input. That's not the best approach but it was the first thing (combined with 3D crop around each vertebra) that worked for me. The other approach is to just use binary masks.<br>\nWhy not just img + mask?  Was a bit lazy to recompute the first conv weights to 2 channels.</p>",
          "rawMarkdown": "@mariuszwisniewski I used integer encoded predicted segmentation mask as an additional input. That's not the best approach but it was the first thing (combined with 3D crop around each vertebra) that worked for me. The other approach is to just use binary masks.\nWhy not just img + mask?  Was a bit lazy to recompute the first conv weights to 2 channels.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2024863,
      "postDate": "2022-11-10T20:00:33.797Z",
      "content": "<p>Congratulations and thanks for sharing your code (RSNA CSN segmentor &amp; classifier ), GitHub and solution too.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing your code (RSNA CSN segmentor & classifier ), GitHub and solution too."
    },
    {
      "id": 2022199,
      "postDate": "2022-11-08T19:48:36.620Z",
      "content": "<p>Wow, awesome result with just 10 days of active submissions <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a>! I appreciate your diligence and efforts. Congratulations!!<br>\nWishing you the best!! <br>\nI appreciate the detailed approach note too, this is informative..</p>",
      "rawMarkdown": "Wow, awesome result with just 10 days of active submissions @selimsef! I appreciate your diligence and efforts. Congratulations!!\nWishing you the best!! \nI appreciate the detailed approach note too, this is informative.."
    }
  ],
  "comments": [
    {
      "id": 2060896,
      "author_name": "Mariusz Wiśniewski",
      "author_url": "",
      "post_date": "2022-12-10T14:36:34.997000",
      "content": "<p><a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a>, could you explain why the input of the classifiers is in the form of 2x images + masks? Below is the part of the code from <code>DatasetCrops.__getitem__()</code> I am asking about.</p>\n<pre><code>images = np.concatenate([images, images, masks], axis=-)\n</code></pre>",
      "votes": 0,
      "replies": [
        {
          "id": 2062151,
          "author_name": "Selim Seferbekov",
          "author_url": "",
          "post_date": "2022-12-11T19:12:39.093000",
          "content": "<p><a href=\"https://www.kaggle.com/mariuszwisniewski\" target=\"_blank\">@mariuszwisniewski</a> I used integer encoded predicted segmentation mask as an additional input. That's not the best approach but it was the first thing (combined with 3D crop around each vertebra) that worked for me. The other approach is to just use binary masks.<br>\nWhy not just img + mask?  Was a bit lazy to recompute the first conv weights to 2 channels.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2024863,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2022-11-10T20:00:33.797000",
      "content": "<p>Congratulations and thanks for sharing your code (RSNA CSN segmentor &amp; classifier ), GitHub and solution too.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2022199,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2022-11-08T19:48:36.620000",
      "content": "<p>Wow, awesome result with just 10 days of active submissions <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a>! I appreciate your diligence and efforts. Congratulations!!<br>\nWishing you the best!! <br>\nI appreciate the detailed approach note too, this is informative..</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2021867": "I started working on this challenge really late, made first sub 10 days before deadline, so could not run a lot of experiments and used codebase/tricks from  [GI tract segmentation challenge](https://www.kaggle.com/competitions/uw-madison-gi-tract-image-segmentation) and NASA comet detection challenge.\n\n###TLDR\nTwo stage approach, 3D segmentation + 3D classification for each vertebra crop\n\n###The main trick\nA lot of  failed attempts to use 3D networks for classification are related to overfitting . That's the case for pure 3D convolutions which usually are not needed and we can easily work with architectures that achieve state of the art performance on video/action classification datasets. \nSo the options are:\n- inflate some state of the art 2d nets, like effnets but start using 3Ds from 2nd or 3rd block . It works but still tends to overfit and  imagenet pretraining is not the best in this case\n- use 2D networks with LSTM/ConvLSTM/Transformer heads. It works and allows to use a lot of architectures, the main issue here is that 3D nature is considered only on the latest stage.\n- use pretrained networks from ig65m/kinetics datasest. That gave the best results for me. \n\n\n\n \nFrom https://paperswithcode.com/sota/action-classification-on-kinetics-400 it is clearly seen that the best convolutional architecture for video classifcation in 2022 is still [ir-CSN-152](https://arxiv.org/abs/1904.02811). Even though transformers achieve higher scores on Kinetics datasets they don't perform as good on small datasets like RSNA and CSNs are the best for small sized datasets and if one needs fast and accurate 3D segmentation and/or classification.\n\nI used mmaction2's implementation of CSNs.\n\n### Segmentation\n- UNet like multiclass segmentor\n- encoder: **ir-CSN-50**\n- decoder: standard unet decoder with nn upsampling but with [(2+1)d convolutions](https://arxiv.org/abs/1711.11248v3)\n- pure 3d convolutions in decoder lead to NaNs in amp training\n\n**Training**:\n\n- 2x subsampling(just slice ::2) by z axis\n- 2x linear downsampled images\n- memory mapping to reduce IO/CPU overhead\n- AdamW + wd, cosine LR annealing\n- Loss: focal-jaccard optimized loss for faster multiclass jaccard computation\n- 2D augs, ReplayCompose in Albumentations library, lightweight geometric augs + hflip\n\n### Classification\n\n- 4 folds of ir(ip)-CSN-152 with global max pooling\n- multilabel 8 classes, if less than 30% of vertebra is visible then make vertebra's label 0 and recompute overall label, otherwise keep the same as provided\n\n**Training**\n\n- 3 channel input (img, img, integer encoded segmentation mask\n- 2x subsampling(just slice ::2) by z axis\n- using 40 slices(80 in original data) around each vertebra\n- crops were resized to 256x256\n- BCE loss\n- target metric as validation\n- AdamW + wd, cosine LR annealing\n- Augmentations: 2D augs, replay compose in Albumentations library, flips, rotations, geometric. Needed to make more augs compared to segmentation pipeline\n\n### Code\n- inference kernel and weights: https://www.kaggle.com/code/selimsef/rsna-csn-segmentor-classifier\n- github code for training: https://github.com/selimsef/rsna_cervical_fracture/",
    "2060896": "@selimsef, could you explain why the input of the classifiers is in the form of 2x images + masks? Below is the part of the code from `DatasetCrops.__getitem__()` I am asking about.\n\n```py\nimages = np.concatenate([images, images, masks], axis=-1)\n```",
    "2024863": "Congratulations and thanks for sharing your code (RSNA CSN segmentor & classifier ), GitHub and solution too.",
    "2022199": "Wow, awesome result with just 10 days of active submissions @selimsef! I appreciate your diligence and efforts. Congratulations!!\nWishing you the best!! \nI appreciate the detailed approach note too, this is informative.."
  }
}