{
  "id": 365115,
  "title": "2nd place solution ： Segmentation + 2.5D CNN + GRU Attention",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/365115",
  "author_name": "Ryan R",
  "post_date": "2022-11-09T21:14:56.748000",
  "votes": 41,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Thanks to Kaggle and RSNA for such a great competition, we are very happy to have finished second. From this complex game, we tried to find the most concise and efficient solution, and gained a lot of knowledge. It was also one of the most hard-drive intensive competitions I've ever seen, and we wasted time loading data because we didn't have enough space to save the high-resolution pseudo-label voxel.</p>\n<p>To cut to the chase, our solution consists of two stages, and use 2.5D CNNs, which we learned from <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@Awsaf</a> in UWM <a href=\"https://www.kaggle.com/code/awsaf49/uwmgi-2-5d-train-pytorch\" target=\"_blank\">UWMGI: 2.5D Train [PyTorch] | Kaggle</a> </p>\n<p>stage1： 2.5D CNN + Unet for Segmentation </p>\n<p>stage2： CNN + BiGRU + Attention for Classification </p>\n<h2>Stage 1</h2>\n<p>First, we used the 87 studies of segmentation samples provided by the organizers . We recreated the mask labels according to the following method</p>\n<pre><code>0 ---&gt; background  \n1 ---&gt; C1  \n2 ---&gt; C2  \n...\n8 ---&gt; T1 - T12  \n</code></pre>\n<p>We used the more general 2.5D and with 3 channels of image data, i.e., the original image i and its sides: i-1, i+1. Thanks to Kaggle and RSNA for such a great competition, we are very happy to have finished second. From this complex game, we tried to find the most concise and efficient solution, and gained a lot of knowledge. It was also one of the most hard-drive intensive competition I've ever seen, and we wasted time loading data because we didn't have enough space to save the high-resolution pseudo-label voxel.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3799319%2F29389b33392e4b84d56d0d53be39c23e%2FSnipaste_2022-11-09_01-18-47.png?generation=1668028486937787&amp;alt=media\" alt=\"\"></p>\n<p>The data augmentation section here is as follows, similar to <a href=\"https://www.kaggle.com/code/awsaf49/uwmgi-2-5d-train-pytorch\" target=\"_blank\">UWMGI: 2.5D [Train] [PyTorch] | Kaggle</a>, without much change.</p>\n<p>We also tried heavier data augmentation, but it did not work better.</p>\n<pre><code>Resize(CFG.img_size, CFG.img_size, interpolation=cv2.INTER_NEAREST),\nHorizontalFlip(p=0.5),\nShiftScaleRotate(shift_limit=0.0625, scale_limit=0.05, rotate_limit=10, p=0.5),\nOneOf([\n    GridDistortion(num_steps=5, distort_limit=0.05, p=1.0),\n    ElasticTransform(alpha=1, sigma=50, alpha_affine=50, p=1.0)\n], p=0.25),\n</code></pre>\n<p>For segmentation model，we used segmentation_models_pytorch lib, backbone was efficientnet-b0, decoder was unet</p>\n<p>Optimizer=\"AdamW\" </p>\n<p>Scheduler=\"CosineAnnealingLR\" + \"GradualWarmupSchedulerV3\"</p>\n<h2>Crop Voxel</h2>\n<p>Once we trained the segmentation model, we generalized it to all 2019 studies, we did the same preprocessing as before for the input data, and after the model predicted the results, we manually looked at several predicted images and found that the accuracy was pretty good.</p>\n<p>We cropped out all 7 cervical vertebrae of each study separately. Each cervical vertebrae to a fracture label from train.csv. According to our EDA, most of the studies contain 200-300 slices, so the average of each vertebrae is about 30 slices. We chose 24 slices, which will be satisfied by most vertebras. For cervical vertebrae with more than 24 slice, we used a simple numpy function to get 24 slices evenly</p>\n<pre><code>sample_index = np.linspace(0, len(one_study_cid)-1, sample_num, dtype=int)\n</code></pre>\n<p>One of the challenges for us was the training images for this competition are 300GB, and if we were to save the cropped 3D high-resolution training images locally, it would exceed the capacity of the hard disk, so we are forced to choose to record the cropped, [x0:x1, y0:y1, z0:z1] and the corresponding slice's dcm file number for the training process in stage2 for reading and cropping.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3799319%2F4ab026464572fb40996597d4aaa424c2%2Fimage-20221110043219911.png?generation=1668028409059020&amp;alt=media\" alt=\"\"></p>\n<h2>Stage 2</h2>\n<p>After we crop, the data shape of our input was (bs, 24, img_size, img_size), 24 channels, representing 24 uniformly distributed slices, and also the seq_len of GRU.</p>\n<p>For data sampling, we ignored the wrong study 1.2.826.0.1.3680043.20574 and 1.2.826.0.1.3680043.29952</p>\n<p>Regarding data augmentation, we used similar methods to stage1 with a little new augmentation.</p>\n<p>For the model we used CNN + biGRU + Attention, where for the CNN backbone we used tf_efficientnetv2_s and resnest50d from the timm library. For some other details, we initialized the GRU, since it seems that the original GRU weights on Pytorch are not very good. We also added SpatialDropout , which also gives us a little improvement.</p>\n<h2>Things we didn't had time to do</h2>\n<ol>\n<li>use bbox csv in yolo</li>\n<li>Transformer for sequential model</li>\n<li>buy new hard-drive :)</li>\n</ol>\n<h2>Code</h2>\n<p><a href=\"https://github.com/ryanyuerong/RSNA2022RAWE\" target=\"_blank\">https://github.com/ryanyuerong/RSNA2022RAWE</a></p>",
  "messages": [
    {
      "id": 2023562,
      "postDate": "2022-11-09T21:14:56.750Z",
      "content": "<p>Thanks to Kaggle and RSNA for such a great competition, we are very happy to have finished second. From this complex game, we tried to find the most concise and efficient solution, and gained a lot of knowledge. It was also one of the most hard-drive intensive competitions I've ever seen, and we wasted time loading data because we didn't have enough space to save the high-resolution pseudo-label voxel.</p>\n<p>To cut to the chase, our solution consists of two stages, and use 2.5D CNNs, which we learned from <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@Awsaf</a> in UWM <a href=\"https://www.kaggle.com/code/awsaf49/uwmgi-2-5d-train-pytorch\" target=\"_blank\">UWMGI: 2.5D Train [PyTorch] | Kaggle</a> </p>\n<p>stage1： 2.5D CNN + Unet for Segmentation </p>\n<p>stage2： CNN + BiGRU + Attention for Classification </p>\n<h2>Stage 1</h2>\n<p>First, we used the 87 studies of segmentation samples provided by the organizers . We recreated the mask labels according to the following method</p>\n<pre><code>0 ---&gt; background  \n1 ---&gt; C1  \n2 ---&gt; C2  \n...\n8 ---&gt; T1 - T12  \n</code></pre>\n<p>We used the more general 2.5D and with 3 channels of image data, i.e., the original image i and its sides: i-1, i+1. Thanks to Kaggle and RSNA for such a great competition, we are very happy to have finished second. From this complex game, we tried to find the most concise and efficient solution, and gained a lot of knowledge. It was also one of the most hard-drive intensive competition I've ever seen, and we wasted time loading data because we didn't have enough space to save the high-resolution pseudo-label voxel.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3799319%2F29389b33392e4b84d56d0d53be39c23e%2FSnipaste_2022-11-09_01-18-47.png?generation=1668028486937787&amp;alt=media\" alt=\"\"></p>\n<p>The data augmentation section here is as follows, similar to <a href=\"https://www.kaggle.com/code/awsaf49/uwmgi-2-5d-train-pytorch\" target=\"_blank\">UWMGI: 2.5D [Train] [PyTorch] | Kaggle</a>, without much change.</p>\n<p>We also tried heavier data augmentation, but it did not work better.</p>\n<pre><code>Resize(CFG.img_size, CFG.img_size, interpolation=cv2.INTER_NEAREST),\nHorizontalFlip(p=0.5),\nShiftScaleRotate(shift_limit=0.0625, scale_limit=0.05, rotate_limit=10, p=0.5),\nOneOf([\n    GridDistortion(num_steps=5, distort_limit=0.05, p=1.0),\n    ElasticTransform(alpha=1, sigma=50, alpha_affine=50, p=1.0)\n], p=0.25),\n</code></pre>\n<p>For segmentation model，we used segmentation_models_pytorch lib, backbone was efficientnet-b0, decoder was unet</p>\n<p>Optimizer=\"AdamW\" </p>\n<p>Scheduler=\"CosineAnnealingLR\" + \"GradualWarmupSchedulerV3\"</p>\n<h2>Crop Voxel</h2>\n<p>Once we trained the segmentation model, we generalized it to all 2019 studies, we did the same preprocessing as before for the input data, and after the model predicted the results, we manually looked at several predicted images and found that the accuracy was pretty good.</p>\n<p>We cropped out all 7 cervical vertebrae of each study separately. Each cervical vertebrae to a fracture label from train.csv. According to our EDA, most of the studies contain 200-300 slices, so the average of each vertebrae is about 30 slices. We chose 24 slices, which will be satisfied by most vertebras. For cervical vertebrae with more than 24 slice, we used a simple numpy function to get 24 slices evenly</p>\n<pre><code>sample_index = np.linspace(0, len(one_study_cid)-1, sample_num, dtype=int)\n</code></pre>\n<p>One of the challenges for us was the training images for this competition are 300GB, and if we were to save the cropped 3D high-resolution training images locally, it would exceed the capacity of the hard disk, so we are forced to choose to record the cropped, [x0:x1, y0:y1, z0:z1] and the corresponding slice's dcm file number for the training process in stage2 for reading and cropping.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3799319%2F4ab026464572fb40996597d4aaa424c2%2Fimage-20221110043219911.png?generation=1668028409059020&amp;alt=media\" alt=\"\"></p>\n<h2>Stage 2</h2>\n<p>After we crop, the data shape of our input was (bs, 24, img_size, img_size), 24 channels, representing 24 uniformly distributed slices, and also the seq_len of GRU.</p>\n<p>For data sampling, we ignored the wrong study 1.2.826.0.1.3680043.20574 and 1.2.826.0.1.3680043.29952</p>\n<p>Regarding data augmentation, we used similar methods to stage1 with a little new augmentation.</p>\n<p>For the model we used CNN + biGRU + Attention, where for the CNN backbone we used tf_efficientnetv2_s and resnest50d from the timm library. For some other details, we initialized the GRU, since it seems that the original GRU weights on Pytorch are not very good. We also added SpatialDropout , which also gives us a little improvement.</p>\n<h2>Things we didn't had time to do</h2>\n<ol>\n<li>use bbox csv in yolo</li>\n<li>Transformer for sequential model</li>\n<li>buy new hard-drive :)</li>\n</ol>\n<h2>Code</h2>\n<p><a href=\"https://github.com/ryanyuerong/RSNA2022RAWE\" target=\"_blank\">https://github.com/ryanyuerong/RSNA2022RAWE</a></p>",
      "rawMarkdown": "Thanks to Kaggle and RSNA for such a great competition, we are very happy to have finished second. From this complex game, we tried to find the most concise and efficient solution, and gained a lot of knowledge. It was also one of the most hard-drive intensive competitions I've ever seen, and we wasted time loading data because we didn't have enough space to save the high-resolution pseudo-label voxel.\n\nTo cut to the chase, our solution consists of two stages, and use 2.5D CNNs, which we learned from [@Awsaf](https://www.kaggle.com/awsaf49) in UWM [UWMGI: 2.5D Train [PyTorch] | Kaggle](https://www.kaggle.com/code/awsaf49/uwmgi-2-5d-train-pytorch) \n\nstage1： 2.5D CNN + Unet for Segmentation \n\nstage2： CNN + BiGRU + Attention for Classification \n               \n            \n## Stage 1\n\nFirst, we used the 87 studies of segmentation samples provided by the organizers . We recreated the mask labels according to the following method\n```\n0 ---> background  \n1 ---> C1  \n2 ---> C2  \n...\n8 ---> T1 - T12  \n```\n\nWe used the more general 2.5D and with 3 channels of image data, i.e., the original image i and its sides: i-1, i+1. Thanks to Kaggle and RSNA for such a great competition, we are very happy to have finished second. From this complex game, we tried to find the most concise and efficient solution, and gained a lot of knowledge. It was also one of the most hard-drive intensive competition I've ever seen, and we wasted time loading data because we didn't have enough space to save the high-resolution pseudo-label voxel.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3799319%2F29389b33392e4b84d56d0d53be39c23e%2FSnipaste_2022-11-09_01-18-47.png?generation=1668028486937787&alt=media)\n\n\nThe data augmentation section here is as follows, similar to [UWMGI: 2.5D [Train\\] [PyTorch] | Kaggle](https://www.kaggle.com/code/awsaf49/uwmgi-2-5d-train-pytorch), without much change.\n\nWe also tried heavier data augmentation, but it did not work better.\n\n```\nResize(CFG.img_size, CFG.img_size, interpolation=cv2.INTER_NEAREST),\nHorizontalFlip(p=0.5),\nShiftScaleRotate(shift_limit=0.0625, scale_limit=0.05, rotate_limit=10, p=0.5),\nOneOf([\n    GridDistortion(num_steps=5, distort_limit=0.05, p=1.0),\n    ElasticTransform(alpha=1, sigma=50, alpha_affine=50, p=1.0)\n], p=0.25),\n```\n\n\n\nFor segmentation model，we used segmentation_models_pytorch lib, backbone was efficientnet-b0, decoder was unet\n\nOptimizer=\"AdamW\" \n\nScheduler=\"CosineAnnealingLR\" + \"GradualWarmupSchedulerV3\"\n\n\n\n\n\n\n                     \n                       \n## Crop Voxel\n\nOnce we trained the segmentation model, we generalized it to all 2019 studies, we did the same preprocessing as before for the input data, and after the model predicted the results, we manually looked at several predicted images and found that the accuracy was pretty good.\n\nWe cropped out all 7 cervical vertebrae of each study separately. Each cervical vertebrae to a fracture label from train.csv. According to our EDA, most of the studies contain 200-300 slices, so the average of each vertebrae is about 30 slices. We chose 24 slices, which will be satisfied by most vertebras. For cervical vertebrae with more than 24 slice, we used a simple numpy function to get 24 slices evenly\n\n\n```\nsample_index = np.linspace(0, len(one_study_cid)-1, sample_num, dtype=int)\n```\n\nOne of the challenges for us was the training images for this competition are 300GB, and if we were to save the cropped 3D high-resolution training images locally, it would exceed the capacity of the hard disk, so we are forced to choose to record the cropped, [x0:x1, y0:y1, z0:z1] and the corresponding slice's dcm file number for the training process in stage2 for reading and cropping.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3799319%2F4ab026464572fb40996597d4aaa424c2%2Fimage-20221110043219911.png?generation=1668028409059020&alt=media)\n\n\n\n\n\n                   \n                     \n## Stage 2\n\n\n\nAfter we crop, the data shape of our input was (bs, 24, img_size, img_size), 24 channels, representing 24 uniformly distributed slices, and also the seq_len of GRU.\n\nFor data sampling, we ignored the wrong study 1.2.826.0.1.3680043.20574 and 1.2.826.0.1.3680043.29952\n\nRegarding data augmentation, we used similar methods to stage1 with a little new augmentation.\n\nFor the model we used CNN + biGRU + Attention, where for the CNN backbone we used tf_efficientnetv2_s and resnest50d from the timm library. For some other details, we initialized the GRU, since it seems that the original GRU weights on Pytorch are not very good. We also added SpatialDropout , which also gives us a little improvement.\n\n\n\n## Things we didn't had time to do\n\n1. use bbox csv in yolo\n2. Transformer for sequential model\n3. buy new hard-drive :)\n\n\n## Code\nhttps://github.com/ryanyuerong/RSNA2022RAWE",
      "votes": 41
    },
    {
      "id": 2204519,
      "postDate": "2023-03-31T16:25:16.640Z",
      "content": "<p>Can someone please tell me what accuracy we get for stage 2?</p>",
      "rawMarkdown": "Can someone please tell me what accuracy we get for stage 2?\n"
    },
    {
      "id": 2101293,
      "postDate": "2023-01-15T19:34:21.280Z",
      "content": "<p>Hello. I have followed your notebooks. Very good solution. However, I got an error while running your inference notebook that says,</p>\n<p>RuntimeError: Error(s) in loading state_dict for RSNAClassifier:<br>\n    size mismatch for model.bn1.weight: copying a param with shape torch.Size([24]) from checkpoint, the shape in current model is torch.Size([64]).</p>\n<p>Why would I get a current model with torch.Size([64])? I am running the same notebook for your inference?<br>\n<a href=\"https://www.kaggle.com/ryanrong\" target=\"_blank\">@ryanrong</a> </p>",
      "rawMarkdown": "Hello. I have followed your notebooks. Very good solution. However, I got an error while running your inference notebook that says,\n\nRuntimeError: Error(s) in loading state_dict for RSNAClassifier:\n\tsize mismatch for model.bn1.weight: copying a param with shape torch.Size([24]) from checkpoint, the shape in current model is torch.Size([64]).\n\nWhy would I get a current model with torch.Size([64])? I am running the same notebook for your inference?\n@ryanrong ",
      "replies": [
        {
          "id": 2204523,
          "postDate": "2023-03-31T16:26:32.903Z",
          "content": "<p>Heyyy. Can you please tell me how much accuracy do you get in stage 2?</p>",
          "rawMarkdown": "Heyyy. Can you please tell me how much accuracy do you get in stage 2?"
        }
      ]
    },
    {
      "id": 2072893,
      "postDate": "2022-12-22T13:54:13.280Z",
      "content": "<p>Congratulations!<br>\nI found that your segmentation model is different from 1st solution. He change con2d to conv3d, is that necessary?<br>\nWhy would you use GRU and Attention as classifier? Could you please explain your motivation?<br>\nwhy does MLPAttentionNetwork called as attention?<br>\nThank you.</p>",
      "rawMarkdown": "Congratulations!\nI found that your segmentation model is different from 1st solution. He change con2d to conv3d, is that necessary?\nWhy would you use GRU and Attention as classifier? Could you please explain your motivation?\nwhy does MLPAttentionNetwork called as attention?\nThank you.",
      "replies": [
        {
          "id": 2213006,
          "postDate": "2023-04-07T08:41:40.467Z",
          "content": "<p>Hi, we considered the computation cost and chose 2d conv. But we definitely think 3d can result in a better result in this case. It is sadly that, due to  limited resources we haved, we didn't done this experiments.</p>",
          "rawMarkdown": "Hi, we considered the computation cost and chose 2d conv. But we definitely think 3d can result in a better result in this case. It is sadly that, due to  limited resources we haved, we didn't done this experiments."
        }
      ]
    },
    {
      "id": 2063671,
      "postDate": "2022-12-13T07:30:24.730Z",
      "content": "<p>Thanks for sharing your nice solution!<br>\nx0,x1,y0,y1 are not appearing in stage2.ipynb on github. Is stage2.ipynb a partial reproduction?</p>",
      "rawMarkdown": "\nThanks for sharing your nice solution!\nx0,x1,y0,y1 are not appearing in stage2.ipynb on github. Is stage2.ipynb a partial reproduction?",
      "replies": [
        {
          "id": 2066554,
          "postDate": "2022-12-15T20:04:21.837Z",
          "content": "<p>thanks for catching this, it was a old version of code. I have updated the code in GitHub</p>",
          "rawMarkdown": "thanks for catching this, it was a old version of code. I have updated the code in GitHub",
          "replies": [
            {
              "id": 2203984,
              "postDate": "2023-03-31T09:34:56.100Z",
              "content": "<p>Heyyy. I just wanted to ask how much accuracy do you get in stage 2?</p>",
              "rawMarkdown": "Heyyy. I just wanted to ask how much accuracy do you get in stage 2?\n"
            }
          ]
        }
      ]
    },
    {
      "id": 2024508,
      "postDate": "2022-11-10T15:21:21.963Z",
      "content": "<p>Awesome notebook <a href=\"https://www.kaggle.com/ryanrong\" target=\"_blank\">@ryanrong</a> </p>",
      "rawMarkdown": "Awesome notebook @ryanrong "
    },
    {
      "id": 2029407,
      "postDate": "2022-11-14T16:21:40.117Z",
      "content": "<p>Thanks for the share</p>",
      "rawMarkdown": "Thanks for the share"
    },
    {
      "id": 2027215,
      "postDate": "2022-11-12T17:10:17.323Z",
      "content": "<p>Thanks for the share ..</p>",
      "rawMarkdown": "Thanks for the share .."
    },
    {
      "id": 2024918,
      "postDate": "2022-11-10T21:34:58.977Z",
      "content": "<p>Thanks for this notebook…</p>",
      "rawMarkdown": "Thanks for this notebook..."
    }
  ],
  "comments": [
    {
      "id": 2204519,
      "author_name": "SHINIT SHETTY",
      "author_url": "",
      "post_date": "2023-03-31T16:25:16.640000",
      "content": "<p>Can someone please tell me what accuracy we get for stage 2?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2101293,
      "author_name": "Atika Rahman Paddo",
      "author_url": "",
      "post_date": "2023-01-15T19:34:21.280000",
      "content": "<p>Hello. I have followed your notebooks. Very good solution. However, I got an error while running your inference notebook that says,</p>\n<p>RuntimeError: Error(s) in loading state_dict for RSNAClassifier:<br>\n    size mismatch for model.bn1.weight: copying a param with shape torch.Size([24]) from checkpoint, the shape in current model is torch.Size([64]).</p>\n<p>Why would I get a current model with torch.Size([64])? I am running the same notebook for your inference?<br>\n<a href=\"https://www.kaggle.com/ryanrong\" target=\"_blank\">@ryanrong</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2204523,
          "author_name": "SHINIT SHETTY",
          "author_url": "",
          "post_date": "2023-03-31T16:26:32.903000",
          "content": "<p>Heyyy. Can you please tell me how much accuracy do you get in stage 2?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2072893,
      "author_name": "Chenjie",
      "author_url": "",
      "post_date": "2022-12-22T13:54:13.280000",
      "content": "<p>Congratulations!<br>\nI found that your segmentation model is different from 1st solution. He change con2d to conv3d, is that necessary?<br>\nWhy would you use GRU and Attention as classifier? Could you please explain your motivation?<br>\nwhy does MLPAttentionNetwork called as attention?<br>\nThank you.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2213006,
          "author_name": "Hao Chen",
          "author_url": "",
          "post_date": "2023-04-07T08:41:40.467000",
          "content": "<p>Hi, we considered the computation cost and chose 2d conv. But we definitely think 3d can result in a better result in this case. It is sadly that, due to  limited resources we haved, we didn't done this experiments.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2063671,
      "author_name": "patriot",
      "author_url": "",
      "post_date": "2022-12-13T07:30:24.730000",
      "content": "<p>Thanks for sharing your nice solution!<br>\nx0,x1,y0,y1 are not appearing in stage2.ipynb on github. Is stage2.ipynb a partial reproduction?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2066554,
          "author_name": "Ryan R",
          "author_url": "",
          "post_date": "2022-12-15T20:04:21.837000",
          "content": "<p>thanks for catching this, it was a old version of code. I have updated the code in GitHub</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2203984,
              "author_name": "SHINIT SHETTY",
              "author_url": "",
              "post_date": "2023-03-31T09:34:56.100000",
              "content": "<p>Heyyy. I just wanted to ask how much accuracy do you get in stage 2?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2024508,
      "author_name": "Allena Venkata Sai Abhishek",
      "author_url": "",
      "post_date": "2022-11-10T15:21:21.963000",
      "content": "<p>Awesome notebook <a href=\"https://www.kaggle.com/ryanrong\" target=\"_blank\">@ryanrong</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2029407,
      "author_name": "Mohamed Elsorady",
      "author_url": "",
      "post_date": "2022-11-14T16:21:40.117000",
      "content": "<p>Thanks for the share</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2027215,
      "author_name": "Pavithra Devi M",
      "author_url": "",
      "post_date": "2022-11-12T17:10:17.323000",
      "content": "<p>Thanks for the share ..</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2024918,
      "author_name": "Raj Saha",
      "author_url": "",
      "post_date": "2022-11-10T21:34:58.977000",
      "content": "<p>Thanks for this notebook…</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2023562": "Thanks to Kaggle and RSNA for such a great competition, we are very happy to have finished second. From this complex game, we tried to find the most concise and efficient solution, and gained a lot of knowledge. It was also one of the most hard-drive intensive competitions I've ever seen, and we wasted time loading data because we didn't have enough space to save the high-resolution pseudo-label voxel.\n\nTo cut to the chase, our solution consists of two stages, and use 2.5D CNNs, which we learned from [@Awsaf](https://www.kaggle.com/awsaf49) in UWM [UWMGI: 2.5D Train [PyTorch] | Kaggle](https://www.kaggle.com/code/awsaf49/uwmgi-2-5d-train-pytorch) \n\nstage1： 2.5D CNN + Unet for Segmentation \n\nstage2： CNN + BiGRU + Attention for Classification \n               \n            \n## Stage 1\n\nFirst, we used the 87 studies of segmentation samples provided by the organizers . We recreated the mask labels according to the following method\n```\n0 ---> background  \n1 ---> C1  \n2 ---> C2  \n...\n8 ---> T1 - T12  \n```\n\nWe used the more general 2.5D and with 3 channels of image data, i.e., the original image i and its sides: i-1, i+1. Thanks to Kaggle and RSNA for such a great competition, we are very happy to have finished second. From this complex game, we tried to find the most concise and efficient solution, and gained a lot of knowledge. It was also one of the most hard-drive intensive competition I've ever seen, and we wasted time loading data because we didn't have enough space to save the high-resolution pseudo-label voxel.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3799319%2F29389b33392e4b84d56d0d53be39c23e%2FSnipaste_2022-11-09_01-18-47.png?generation=1668028486937787&alt=media)\n\n\nThe data augmentation section here is as follows, similar to [UWMGI: 2.5D [Train\\] [PyTorch] | Kaggle](https://www.kaggle.com/code/awsaf49/uwmgi-2-5d-train-pytorch), without much change.\n\nWe also tried heavier data augmentation, but it did not work better.\n\n```\nResize(CFG.img_size, CFG.img_size, interpolation=cv2.INTER_NEAREST),\nHorizontalFlip(p=0.5),\nShiftScaleRotate(shift_limit=0.0625, scale_limit=0.05, rotate_limit=10, p=0.5),\nOneOf([\n    GridDistortion(num_steps=5, distort_limit=0.05, p=1.0),\n    ElasticTransform(alpha=1, sigma=50, alpha_affine=50, p=1.0)\n], p=0.25),\n```\n\n\n\nFor segmentation model，we used segmentation_models_pytorch lib, backbone was efficientnet-b0, decoder was unet\n\nOptimizer=\"AdamW\" \n\nScheduler=\"CosineAnnealingLR\" + \"GradualWarmupSchedulerV3\"\n\n\n\n\n\n\n                     \n                       \n## Crop Voxel\n\nOnce we trained the segmentation model, we generalized it to all 2019 studies, we did the same preprocessing as before for the input data, and after the model predicted the results, we manually looked at several predicted images and found that the accuracy was pretty good.\n\nWe cropped out all 7 cervical vertebrae of each study separately. Each cervical vertebrae to a fracture label from train.csv. According to our EDA, most of the studies contain 200-300 slices, so the average of each vertebrae is about 30 slices. We chose 24 slices, which will be satisfied by most vertebras. For cervical vertebrae with more than 24 slice, we used a simple numpy function to get 24 slices evenly\n\n\n```\nsample_index = np.linspace(0, len(one_study_cid)-1, sample_num, dtype=int)\n```\n\nOne of the challenges for us was the training images for this competition are 300GB, and if we were to save the cropped 3D high-resolution training images locally, it would exceed the capacity of the hard disk, so we are forced to choose to record the cropped, [x0:x1, y0:y1, z0:z1] and the corresponding slice's dcm file number for the training process in stage2 for reading and cropping.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3799319%2F4ab026464572fb40996597d4aaa424c2%2Fimage-20221110043219911.png?generation=1668028409059020&alt=media)\n\n\n\n\n\n                   \n                     \n## Stage 2\n\n\n\nAfter we crop, the data shape of our input was (bs, 24, img_size, img_size), 24 channels, representing 24 uniformly distributed slices, and also the seq_len of GRU.\n\nFor data sampling, we ignored the wrong study 1.2.826.0.1.3680043.20574 and 1.2.826.0.1.3680043.29952\n\nRegarding data augmentation, we used similar methods to stage1 with a little new augmentation.\n\nFor the model we used CNN + biGRU + Attention, where for the CNN backbone we used tf_efficientnetv2_s and resnest50d from the timm library. For some other details, we initialized the GRU, since it seems that the original GRU weights on Pytorch are not very good. We also added SpatialDropout , which also gives us a little improvement.\n\n\n\n## Things we didn't had time to do\n\n1. use bbox csv in yolo\n2. Transformer for sequential model\n3. buy new hard-drive :)\n\n\n## Code\nhttps://github.com/ryanyuerong/RSNA2022RAWE",
    "2204519": "Can someone please tell me what accuracy we get for stage 2?\n",
    "2101293": "Hello. I have followed your notebooks. Very good solution. However, I got an error while running your inference notebook that says,\n\nRuntimeError: Error(s) in loading state_dict for RSNAClassifier:\n\tsize mismatch for model.bn1.weight: copying a param with shape torch.Size([24]) from checkpoint, the shape in current model is torch.Size([64]).\n\nWhy would I get a current model with torch.Size([64])? I am running the same notebook for your inference?\n@ryanrong ",
    "2072893": "Congratulations!\nI found that your segmentation model is different from 1st solution. He change con2d to conv3d, is that necessary?\nWhy would you use GRU and Attention as classifier? Could you please explain your motivation?\nwhy does MLPAttentionNetwork called as attention?\nThank you.",
    "2063671": "\nThanks for sharing your nice solution!\nx0,x1,y0,y1 are not appearing in stage2.ipynb on github. Is stage2.ipynb a partial reproduction?",
    "2024508": "Awesome notebook @ryanrong ",
    "2029407": "Thanks for the share",
    "2027215": "Thanks for the share ..",
    "2024918": "Thanks for this notebook..."
  }
}