{
  "id": 362771,
  "title": "14th place solution",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362771",
  "author_name": "Jesse",
  "post_date": "2022-10-29T02:26:13.776000",
  "votes": 18,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I wasn't sure if this was worth writing up, but since we seemed to do a few things differently, and not everyone used the bounding box annotations, figured we could share some insights. </p>\n<p>Hardware - 3090ti - thanks UT for this… perhaps minor compared to other hardware out there, but was a big upgrade personally over my now aging 1080tis (also huge thanks to crypto bust for making this available 😊)</p>\n<p>#</p>\n<h1>1st stage -</h1>\n<p>3D nnUnet pretrained on totalsegmentator. I was initially hesitant about 3D, mainly because I’ve never trained a 3D model before, and my computer’s 32GB of RAM kept having issues with dataloading beyond a certain size. My teammate Yee showed some good success with 3D, and this finally inspired me to overcome this fear. I also ended up using a library called rising <a href=\"https://rising.readthedocs.io/en/stable/transforms.html\" target=\"_blank\">https://rising.readthedocs.io/en/stable/transforms.html</a>, which made augmentations on the GPU easier. I wrote a 3D version of cutout (3D black cube) for it, but beyond that just used the library as is with standard rotate and random crop.</p>\n<ul>\n<li>trained on 87 segmentation cases to predict C1-C7</li>\n<li>trained for ~12 hours (400 epochs)</li>\n<li>intake 1x192x192x192</li>\n<li>modified to output 7x192x192x192;</li>\n<li>DICE + BCE loss; DSC of 0.95</li>\n<li>this gave us the best segmentation results</li>\n</ul>\n<p>Example (input, ground truth segmentation mask, predicted segmentation mask) -<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1598278%2Fa32d84a2533566a7e06e5c57250acc88%2Fim_130.png?generation=1667006906874204&amp;alt=media\" alt=\"\"></p>\n<p>All 2019 cases were then pseudolabeled with the segmentations. The total volume of CSpine was cropped to just C1-C7. </p>\n<p>#</p>\n<h1>2nd stage -</h1>\n<p>Unet-CNN 2.5D (EffB7 noisy student backbone), model from SMP. Images resized to 448x448. Axial slices were stacked by 3, and one channel also for providing level information (divide channel index by 8, so C1 was 0.125, C2 was 0.25, etc.). So CNN input was 4x448x448. Segmentation output is 2 channels for 2 separate tasks (below). Auxiliary embedding dim of 512 was taken from the middle of the Unet for additional 3 tasks.</p>\n<p>Task 1 - One segmentation channel is to predict bounding boxes (as segmentation maps), simple BCE with pos_weight of 7. </p>\n<p>Task 2 - Second segmentation channel is to predict the intersection of bounding box and vertebral segmentation mask (IF it’s positive based on the ground truth label for the case), trained with DICE + BCE (pos_weight 2) - see picture below for example. </p>\n<p>Task 3 - Auxiliary embedding (512 layer) output to 1 dim linear head to predict fx vs no fx, based on whether a bounding box slice exists. Since bounding boxes were only annotated for about ~50% of positive cases, we had to exclude about 250 cases. All slices from negative cases for fx (~1000 cases), plus positive cases without bounding box on the slice, could still be used as negative for fx. We intended to train a few bounding box models to pseudo label the remaining 250 cases but ended up losing time in trying other things. BCE loss, pos_weight 2.</p>\n<p>Task 4 - Auxiliary embedding (512 layer) output to 7 dim linear head to predict the ratio of vertebral bodies in the slice. We thresholded this at 1000, not sure this makes a difference. So if there are 2000 C5 pixels and 3000 C6 pixels and 500 C7 pixels, ground truth is [0,0,0,0,1,1,0.5] BCE loss, pos_weight 7.</p>\n<p>Task 5 - Auxiliary embedding (512 layer) output to 7 dim linear head to predict level fx vs no fx, which was determined by bounding box presence, level presence, and also weakly thresholded by pixel values. BCE loss, pos_weight 7.</p>\n<p>Examples (input 3 channel from 3 slices, single channel level segmentation mask, ground truth 2 channel segmentation mask, predicted 2 channel segmentation mask - red is bbox, white is intersection of bbox and segmented vert if positive) -</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1598278%2F14cb4d8a6aca61d605af8e45526ce14e%2Fim_1.jpg?generation=1667006208078177&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1598278%2F7e7e5e1a04cb25d9056155ccc6a9bb74%2Fim_19.jpg?generation=1667006224901265&amp;alt=media\" alt=\"\"></p>\n<p>For some reason the above model trained very quickly, despite aggressive regularization and augmentation, and performed worse (on 3rd stage) if trained for more than 1 epoch (~3 hrs). I think this needed better optimization, but a lesson I learned is that multitask is harder to optimize in general. The ideal ratio of loss weighting to balance the different tasks would need more hyperparameter optimization than a simpler solution? Had some slight success when favored segmentation losses over classification losses. Happy to get feedback/intuition on this in general if anyone has ideas.</p>\n<p>#</p>\n<h1>3rd stage -</h1>\n<p>RNN+Attention on the 512 dim embedding layer from the 2nd stage. Pretty standard stuff here.</p>\n<p>#</p>\n<p>We had high hopes on doing lots of ensembling of the above. Particularly because in addition to the RNN+Attention from the auxiliary layer, the segmentation channel outputs can be averaged and potentially be used for a 3rd stage 3D model, which some early tests showed modest success (just on fx channel though?). About 2 weeks from the end, I realized that the kaggle kernel will not handle even 2 folds of a simple pipeline. Oops. Lesson learned. Lots of optimizing of dataloading later, more than 2 folds of the above wasn’t possible. Out of 31 submissions, 6 were testing to get things working, and 19 were errored out (narrowed it down to VRAM more than CPU/RAM). So EffB7ns was too big for this comp perhaps. Last minute incomplete testing of smaller models didn’t learn level information as well though. </p>\n<p>Things that also kind of worked </p>\n<ul>\n<li>2.5D CNN without the 1st stage 3D model, but predicting 8/9 channels (7 channels for each level, loss only calculated for the 87 segmentation cases and weighting of ~20, 1-2 channels for fx detection, weighting of ~2)</li>\n<li>3D multitask model</li>\n</ul>\n<p>Things that didn't work</p>\n<ul>\n<li>Pure 3D model - we wasted so much time here haha</li>\n</ul>\n<p>#</p>\n<p>Had a lot of fun and learned a lot as usual. Congrats to winners, and all the teams that participated!</p>",
  "messages": [
    {
      "id": 2008387,
      "postDate": "2022-10-29T02:26:13.777Z",
      "content": "<p>I wasn't sure if this was worth writing up, but since we seemed to do a few things differently, and not everyone used the bounding box annotations, figured we could share some insights. </p>\n<p>Hardware - 3090ti - thanks UT for this… perhaps minor compared to other hardware out there, but was a big upgrade personally over my now aging 1080tis (also huge thanks to crypto bust for making this available 😊)</p>\n<p>#</p>\n<h1>1st stage -</h1>\n<p>3D nnUnet pretrained on totalsegmentator. I was initially hesitant about 3D, mainly because I’ve never trained a 3D model before, and my computer’s 32GB of RAM kept having issues with dataloading beyond a certain size. My teammate Yee showed some good success with 3D, and this finally inspired me to overcome this fear. I also ended up using a library called rising <a href=\"https://rising.readthedocs.io/en/stable/transforms.html\" target=\"_blank\">https://rising.readthedocs.io/en/stable/transforms.html</a>, which made augmentations on the GPU easier. I wrote a 3D version of cutout (3D black cube) for it, but beyond that just used the library as is with standard rotate and random crop.</p>\n<ul>\n<li>trained on 87 segmentation cases to predict C1-C7</li>\n<li>trained for ~12 hours (400 epochs)</li>\n<li>intake 1x192x192x192</li>\n<li>modified to output 7x192x192x192;</li>\n<li>DICE + BCE loss; DSC of 0.95</li>\n<li>this gave us the best segmentation results</li>\n</ul>\n<p>Example (input, ground truth segmentation mask, predicted segmentation mask) -<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1598278%2Fa32d84a2533566a7e06e5c57250acc88%2Fim_130.png?generation=1667006906874204&amp;alt=media\" alt=\"\"></p>\n<p>All 2019 cases were then pseudolabeled with the segmentations. The total volume of CSpine was cropped to just C1-C7. </p>\n<p>#</p>\n<h1>2nd stage -</h1>\n<p>Unet-CNN 2.5D (EffB7 noisy student backbone), model from SMP. Images resized to 448x448. Axial slices were stacked by 3, and one channel also for providing level information (divide channel index by 8, so C1 was 0.125, C2 was 0.25, etc.). So CNN input was 4x448x448. Segmentation output is 2 channels for 2 separate tasks (below). Auxiliary embedding dim of 512 was taken from the middle of the Unet for additional 3 tasks.</p>\n<p>Task 1 - One segmentation channel is to predict bounding boxes (as segmentation maps), simple BCE with pos_weight of 7. </p>\n<p>Task 2 - Second segmentation channel is to predict the intersection of bounding box and vertebral segmentation mask (IF it’s positive based on the ground truth label for the case), trained with DICE + BCE (pos_weight 2) - see picture below for example. </p>\n<p>Task 3 - Auxiliary embedding (512 layer) output to 1 dim linear head to predict fx vs no fx, based on whether a bounding box slice exists. Since bounding boxes were only annotated for about ~50% of positive cases, we had to exclude about 250 cases. All slices from negative cases for fx (~1000 cases), plus positive cases without bounding box on the slice, could still be used as negative for fx. We intended to train a few bounding box models to pseudo label the remaining 250 cases but ended up losing time in trying other things. BCE loss, pos_weight 2.</p>\n<p>Task 4 - Auxiliary embedding (512 layer) output to 7 dim linear head to predict the ratio of vertebral bodies in the slice. We thresholded this at 1000, not sure this makes a difference. So if there are 2000 C5 pixels and 3000 C6 pixels and 500 C7 pixels, ground truth is [0,0,0,0,1,1,0.5] BCE loss, pos_weight 7.</p>\n<p>Task 5 - Auxiliary embedding (512 layer) output to 7 dim linear head to predict level fx vs no fx, which was determined by bounding box presence, level presence, and also weakly thresholded by pixel values. BCE loss, pos_weight 7.</p>\n<p>Examples (input 3 channel from 3 slices, single channel level segmentation mask, ground truth 2 channel segmentation mask, predicted 2 channel segmentation mask - red is bbox, white is intersection of bbox and segmented vert if positive) -</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1598278%2F14cb4d8a6aca61d605af8e45526ce14e%2Fim_1.jpg?generation=1667006208078177&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1598278%2F7e7e5e1a04cb25d9056155ccc6a9bb74%2Fim_19.jpg?generation=1667006224901265&amp;alt=media\" alt=\"\"></p>\n<p>For some reason the above model trained very quickly, despite aggressive regularization and augmentation, and performed worse (on 3rd stage) if trained for more than 1 epoch (~3 hrs). I think this needed better optimization, but a lesson I learned is that multitask is harder to optimize in general. The ideal ratio of loss weighting to balance the different tasks would need more hyperparameter optimization than a simpler solution? Had some slight success when favored segmentation losses over classification losses. Happy to get feedback/intuition on this in general if anyone has ideas.</p>\n<p>#</p>\n<h1>3rd stage -</h1>\n<p>RNN+Attention on the 512 dim embedding layer from the 2nd stage. Pretty standard stuff here.</p>\n<p>#</p>\n<p>We had high hopes on doing lots of ensembling of the above. Particularly because in addition to the RNN+Attention from the auxiliary layer, the segmentation channel outputs can be averaged and potentially be used for a 3rd stage 3D model, which some early tests showed modest success (just on fx channel though?). About 2 weeks from the end, I realized that the kaggle kernel will not handle even 2 folds of a simple pipeline. Oops. Lesson learned. Lots of optimizing of dataloading later, more than 2 folds of the above wasn’t possible. Out of 31 submissions, 6 were testing to get things working, and 19 were errored out (narrowed it down to VRAM more than CPU/RAM). So EffB7ns was too big for this comp perhaps. Last minute incomplete testing of smaller models didn’t learn level information as well though. </p>\n<p>Things that also kind of worked </p>\n<ul>\n<li>2.5D CNN without the 1st stage 3D model, but predicting 8/9 channels (7 channels for each level, loss only calculated for the 87 segmentation cases and weighting of ~20, 1-2 channels for fx detection, weighting of ~2)</li>\n<li>3D multitask model</li>\n</ul>\n<p>Things that didn't work</p>\n<ul>\n<li>Pure 3D model - we wasted so much time here haha</li>\n</ul>\n<p>#</p>\n<p>Had a lot of fun and learned a lot as usual. Congrats to winners, and all the teams that participated!</p>",
      "rawMarkdown": "I wasn't sure if this was worth writing up, but since we seemed to do a few things differently, and not everyone used the bounding box annotations, figured we could share some insights. \n\nHardware - 3090ti - thanks UT for this... perhaps minor compared to other hardware out there, but was a big upgrade personally over my now aging 1080tis (also huge thanks to crypto bust for making this available 😊)\n\n#\n# 1st stage - \n3D nnUnet pretrained on totalsegmentator. I was initially hesitant about 3D, mainly because I’ve never trained a 3D model before, and my computer’s 32GB of RAM kept having issues with dataloading beyond a certain size. My teammate Yee showed some good success with 3D, and this finally inspired me to overcome this fear. I also ended up using a library called rising https://rising.readthedocs.io/en/stable/transforms.html, which made augmentations on the GPU easier. I wrote a 3D version of cutout (3D black cube) for it, but beyond that just used the library as is with standard rotate and random crop.\n- trained on 87 segmentation cases to predict C1-C7\n- trained for ~12 hours (400 epochs)\n- intake 1x192x192x192\n- modified to output 7x192x192x192;\n- DICE + BCE loss; DSC of 0.95\n- this gave us the best segmentation results\n\nExample (input, ground truth segmentation mask, predicted segmentation mask) -\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1598278%2Fa32d84a2533566a7e06e5c57250acc88%2Fim_130.png?generation=1667006906874204&alt=media)\n\nAll 2019 cases were then pseudolabeled with the segmentations. The total volume of CSpine was cropped to just C1-C7. \n\n#\n# 2nd stage - \nUnet-CNN 2.5D (EffB7 noisy student backbone), model from SMP. Images resized to 448x448. Axial slices were stacked by 3, and one channel also for providing level information (divide channel index by 8, so C1 was 0.125, C2 was 0.25, etc.). So CNN input was 4x448x448. Segmentation output is 2 channels for 2 separate tasks (below). Auxiliary embedding dim of 512 was taken from the middle of the Unet for additional 3 tasks.\n\nTask 1 - One segmentation channel is to predict bounding boxes (as segmentation maps), simple BCE with pos_weight of 7. \n\nTask 2 - Second segmentation channel is to predict the intersection of bounding box and vertebral segmentation mask (IF it’s positive based on the ground truth label for the case), trained with DICE + BCE (pos_weight 2) - see picture below for example. \n\nTask 3 - Auxiliary embedding (512 layer) output to 1 dim linear head to predict fx vs no fx, based on whether a bounding box slice exists. Since bounding boxes were only annotated for about ~50% of positive cases, we had to exclude about 250 cases. All slices from negative cases for fx (~1000 cases), plus positive cases without bounding box on the slice, could still be used as negative for fx. We intended to train a few bounding box models to pseudo label the remaining 250 cases but ended up losing time in trying other things. BCE loss, pos_weight 2.\n\nTask 4 - Auxiliary embedding (512 layer) output to 7 dim linear head to predict the ratio of vertebral bodies in the slice. We thresholded this at 1000, not sure this makes a difference. So if there are 2000 C5 pixels and 3000 C6 pixels and 500 C7 pixels, ground truth is [0,0,0,0,1,1,0.5] BCE loss, pos_weight 7.\n\nTask 5 - Auxiliary embedding (512 layer) output to 7 dim linear head to predict level fx vs no fx, which was determined by bounding box presence, level presence, and also weakly thresholded by pixel values. BCE loss, pos_weight 7.\n\nExamples (input 3 channel from 3 slices, single channel level segmentation mask, ground truth 2 channel segmentation mask, predicted 2 channel segmentation mask - red is bbox, white is intersection of bbox and segmented vert if positive) -\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1598278%2F14cb4d8a6aca61d605af8e45526ce14e%2Fim_1.jpg?generation=1667006208078177&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1598278%2F7e7e5e1a04cb25d9056155ccc6a9bb74%2Fim_19.jpg?generation=1667006224901265&alt=media)\n\nFor some reason the above model trained very quickly, despite aggressive regularization and augmentation, and performed worse (on 3rd stage) if trained for more than 1 epoch (~3 hrs). I think this needed better optimization, but a lesson I learned is that multitask is harder to optimize in general. The ideal ratio of loss weighting to balance the different tasks would need more hyperparameter optimization than a simpler solution? Had some slight success when favored segmentation losses over classification losses. Happy to get feedback/intuition on this in general if anyone has ideas.\n\n\n#\n# 3rd stage - \nRNN+Attention on the 512 dim embedding layer from the 2nd stage. Pretty standard stuff here.\n\n#\n\nWe had high hopes on doing lots of ensembling of the above. Particularly because in addition to the RNN+Attention from the auxiliary layer, the segmentation channel outputs can be averaged and potentially be used for a 3rd stage 3D model, which some early tests showed modest success (just on fx channel though?). About 2 weeks from the end, I realized that the kaggle kernel will not handle even 2 folds of a simple pipeline. Oops. Lesson learned. Lots of optimizing of dataloading later, more than 2 folds of the above wasn’t possible. Out of 31 submissions, 6 were testing to get things working, and 19 were errored out (narrowed it down to VRAM more than CPU/RAM). So EffB7ns was too big for this comp perhaps. Last minute incomplete testing of smaller models didn’t learn level information as well though. \n\nThings that also kind of worked \n- 2.5D CNN without the 1st stage 3D model, but predicting 8/9 channels (7 channels for each level, loss only calculated for the 87 segmentation cases and weighting of ~20, 1-2 channels for fx detection, weighting of ~2)\n- 3D multitask model\n\nThings that didn't work\n- Pure 3D model - we wasted so much time here haha\n\n#\n\nHad a lot of fun and learned a lot as usual. Congrats to winners, and all the teams that participated!\n",
      "votes": 18
    },
    {
      "id": 2008393,
      "postDate": "2022-10-29T02:42:17.540Z",
      "content": "<p>Nice pipeline <a href=\"https://www.kaggle.com/jcsagar\" target=\"_blank\">@jcsagar</a>. Your Task2 second segmentation channel is something I did not even consider. Was it particularly useful to the overall result?</p>",
      "rawMarkdown": "Nice pipeline @jcsagar. Your Task2 second segmentation channel is something I did not even consider. Was it particularly useful to the overall result?",
      "votes": 1,
      "replies": [
        {
          "id": 2008395,
          "postDate": "2022-10-29T02:47:08.673Z",
          "content": "<p>Edit: Sorry misunderstood initially. </p>\n<p>I think it made a difference. But not sure exactly how much. Local cv testing without any segmentation mask at all was ~0.45 vs ~0.3. Second segmentation channel alone is probably less of a difference… Never purely tested bbox channel vs both channel at the end, but earlier tests were ~0.05 difference.</p>",
          "rawMarkdown": "Edit: Sorry misunderstood initially. \n\nI think it made a difference. But not sure exactly how much. Local cv testing without any segmentation mask at all was ~0.45 vs ~0.3. Second segmentation channel alone is probably less of a difference... Never purely tested bbox channel vs both channel at the end, but earlier tests were ~0.05 difference.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2008393,
      "author_name": "David Roberts",
      "author_url": "",
      "post_date": "2022-10-29T02:42:17.540000",
      "content": "<p>Nice pipeline <a href=\"https://www.kaggle.com/jcsagar\" target=\"_blank\">@jcsagar</a>. Your Task2 second segmentation channel is something I did not even consider. Was it particularly useful to the overall result?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2008395,
          "author_name": "Jesse",
          "author_url": "",
          "post_date": "2022-10-29T02:47:08.673000",
          "content": "<p>Edit: Sorry misunderstood initially. </p>\n<p>I think it made a difference. But not sure exactly how much. Local cv testing without any segmentation mask at all was ~0.45 vs ~0.3. Second segmentation channel alone is probably less of a difference… Never purely tested bbox channel vs both channel at the end, but earlier tests were ~0.05 difference.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2008387": "I wasn't sure if this was worth writing up, but since we seemed to do a few things differently, and not everyone used the bounding box annotations, figured we could share some insights. \n\nHardware - 3090ti - thanks UT for this... perhaps minor compared to other hardware out there, but was a big upgrade personally over my now aging 1080tis (also huge thanks to crypto bust for making this available 😊)\n\n#\n# 1st stage - \n3D nnUnet pretrained on totalsegmentator. I was initially hesitant about 3D, mainly because I’ve never trained a 3D model before, and my computer’s 32GB of RAM kept having issues with dataloading beyond a certain size. My teammate Yee showed some good success with 3D, and this finally inspired me to overcome this fear. I also ended up using a library called rising https://rising.readthedocs.io/en/stable/transforms.html, which made augmentations on the GPU easier. I wrote a 3D version of cutout (3D black cube) for it, but beyond that just used the library as is with standard rotate and random crop.\n- trained on 87 segmentation cases to predict C1-C7\n- trained for ~12 hours (400 epochs)\n- intake 1x192x192x192\n- modified to output 7x192x192x192;\n- DICE + BCE loss; DSC of 0.95\n- this gave us the best segmentation results\n\nExample (input, ground truth segmentation mask, predicted segmentation mask) -\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1598278%2Fa32d84a2533566a7e06e5c57250acc88%2Fim_130.png?generation=1667006906874204&alt=media)\n\nAll 2019 cases were then pseudolabeled with the segmentations. The total volume of CSpine was cropped to just C1-C7. \n\n#\n# 2nd stage - \nUnet-CNN 2.5D (EffB7 noisy student backbone), model from SMP. Images resized to 448x448. Axial slices were stacked by 3, and one channel also for providing level information (divide channel index by 8, so C1 was 0.125, C2 was 0.25, etc.). So CNN input was 4x448x448. Segmentation output is 2 channels for 2 separate tasks (below). Auxiliary embedding dim of 512 was taken from the middle of the Unet for additional 3 tasks.\n\nTask 1 - One segmentation channel is to predict bounding boxes (as segmentation maps), simple BCE with pos_weight of 7. \n\nTask 2 - Second segmentation channel is to predict the intersection of bounding box and vertebral segmentation mask (IF it’s positive based on the ground truth label for the case), trained with DICE + BCE (pos_weight 2) - see picture below for example. \n\nTask 3 - Auxiliary embedding (512 layer) output to 1 dim linear head to predict fx vs no fx, based on whether a bounding box slice exists. Since bounding boxes were only annotated for about ~50% of positive cases, we had to exclude about 250 cases. All slices from negative cases for fx (~1000 cases), plus positive cases without bounding box on the slice, could still be used as negative for fx. We intended to train a few bounding box models to pseudo label the remaining 250 cases but ended up losing time in trying other things. BCE loss, pos_weight 2.\n\nTask 4 - Auxiliary embedding (512 layer) output to 7 dim linear head to predict the ratio of vertebral bodies in the slice. We thresholded this at 1000, not sure this makes a difference. So if there are 2000 C5 pixels and 3000 C6 pixels and 500 C7 pixels, ground truth is [0,0,0,0,1,1,0.5] BCE loss, pos_weight 7.\n\nTask 5 - Auxiliary embedding (512 layer) output to 7 dim linear head to predict level fx vs no fx, which was determined by bounding box presence, level presence, and also weakly thresholded by pixel values. BCE loss, pos_weight 7.\n\nExamples (input 3 channel from 3 slices, single channel level segmentation mask, ground truth 2 channel segmentation mask, predicted 2 channel segmentation mask - red is bbox, white is intersection of bbox and segmented vert if positive) -\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1598278%2F14cb4d8a6aca61d605af8e45526ce14e%2Fim_1.jpg?generation=1667006208078177&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1598278%2F7e7e5e1a04cb25d9056155ccc6a9bb74%2Fim_19.jpg?generation=1667006224901265&alt=media)\n\nFor some reason the above model trained very quickly, despite aggressive regularization and augmentation, and performed worse (on 3rd stage) if trained for more than 1 epoch (~3 hrs). I think this needed better optimization, but a lesson I learned is that multitask is harder to optimize in general. The ideal ratio of loss weighting to balance the different tasks would need more hyperparameter optimization than a simpler solution? Had some slight success when favored segmentation losses over classification losses. Happy to get feedback/intuition on this in general if anyone has ideas.\n\n\n#\n# 3rd stage - \nRNN+Attention on the 512 dim embedding layer from the 2nd stage. Pretty standard stuff here.\n\n#\n\nWe had high hopes on doing lots of ensembling of the above. Particularly because in addition to the RNN+Attention from the auxiliary layer, the segmentation channel outputs can be averaged and potentially be used for a 3rd stage 3D model, which some early tests showed modest success (just on fx channel though?). About 2 weeks from the end, I realized that the kaggle kernel will not handle even 2 folds of a simple pipeline. Oops. Lesson learned. Lots of optimizing of dataloading later, more than 2 folds of the above wasn’t possible. Out of 31 submissions, 6 were testing to get things working, and 19 were errored out (narrowed it down to VRAM more than CPU/RAM). So EffB7ns was too big for this comp perhaps. Last minute incomplete testing of smaller models didn’t learn level information as well though. \n\nThings that also kind of worked \n- 2.5D CNN without the 1st stage 3D model, but predicting 8/9 channels (7 channels for each level, loss only calculated for the 87 segmentation cases and weighting of ~20, 1-2 channels for fx detection, weighting of ~2)\n- 3D multitask model\n\nThings that didn't work\n- Pure 3D model - we wasted so much time here haha\n\n#\n\nHad a lot of fun and learned a lot as usual. Congrats to winners, and all the teams that participated!\n",
    "2008393": "Nice pipeline @jcsagar. Your Task2 second segmentation channel is something I did not even consider. Was it particularly useful to the overall result?"
  }
}