{
  "id": 447448,
  "title": "16th place solution ",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/447448",
  "author_name": "Ivan Panshin",
  "post_date": "2023-10-16T00:01:26.818000",
  "votes": 25,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thanks to the organizers for such a great competition. Unfortunately, I couldn’t take part last year in a similar one due to the lack of hardware. However, this year is different, and I’m really happy with the results that we managed to achieve. </p>\n<h2>Problem</h2>\n<p>In this competition we were tasked with predicting the intensity of injuries for different abdominal organs. The available data consists of big CT images (in DICOM format) with partial supplemental segmentation annotations (in NIFTI format).</p>\n<p>In ML terms, all comes down to 3D segmentation / classification models, lack of data / annotations, and heavily-penalizing metric.</p>\n<h2>Summary</h2>\n<ul>\n<li><strong>U-Net Bi-Conv-LSTM</strong> segmentation for organs (kidney, liver, spleen, bowel) and separate model for Extravasation based on boxes from here (thanks a lot!)</li>\n<li><strong>Resnet 3D CSN</strong> for 3D crops separated into 2 stages: (kidney, liver, spleen) and (bowel)</li>\n</ul>\n<h2>3D Semantic Segmentation</h2>\n<p>The semantic segmentation part to identify organs is quite straightforward (compared to the later classification) and could be effectively performed by <strong>U-Net Bi-Conv-LSTM</strong> with Effnet_v2_b0 backbone and <strong>CE-Dice-Focal</strong> loss. </p>\n<p>To make things efficient, we train semantic segmentation on <strong>96x256x256 crops</strong> and predict the whole image using crops of size 96x256x256 with <strong>overlaps of 48x256x256</strong> (later on overlaps were removed to save time and space in inference).</p>\n<p>To elevate overfitting (it’s not that critical, especially compared to classification), we added <strong>geometric augmentations</strong> like ShiftScaleRotate, RandomBrightnessContrast, Flips, GridDistortion, ElasticTransform. </p>\n<p>On average, about 10% of total volume was dedicated to kidney, liver, spleen, and about 20% - to bowel.</p>\n<p>The <strong>macro dice score per image</strong> is around <strong>0.96</strong>. </p>\n<p><strong>Training time</strong> - about <strong>12 hours</strong>.</p>\n<h2>3D classification</h2>\n<p>Based on extracted crops from segmentation masks, we train 2 models: one for kidney, liver, spleen, one for bowel. </p>\n<p>The <strong>CSN models</strong> from mmaction proved to be very fast and accurate. In order to figure out how to deal with temporal dimension, several possibilities were explored, but in the end basic interpolation (<strong>3D resize</strong>) was used to convert crops to <strong>96x256x256 resolution</strong>.</p>\n<p>To battle overfitting (which is really severe even with CSN), <strong>intensive geometric augmentations</strong> were used, including ShiftScaleRotate, RandomBrightnessContrast, and 4 different types of Blurs.</p>\n<p>The mean competition loss across all folds is <strong>0.401</strong> for kidney, liver, spleen and 0<strong>.156</strong> for bowel.</p>\n<p><strong>Training time</strong> - about <strong>4 hours</strong> per kidney, liver, spleen fold; and <strong>8-10 hours</strong> per bowel fold. </p>\n<h2>3D classification for Extravasation</h2>\n<p>In order to make predictions for Extravasation, a segmentation model was utilized. The motivation is simple: if semantic segmentation model predict anything, there is Extravasation, and it should be reflected in the probabilities. </p>\n<p>To make the <strong>transition from semantic segmentation to classification</strong>, the following trick was used:</p>\n<ul>\n<li>Turn 3D mask to 1D </li>\n<li>Sort probabilities </li>\n<li>Take top_n probabilities</li>\n<li>Find mean values of them. That’s the probability for positive Extravasation.</li>\n</ul>\n<p>In pseudo-code:<br>\n<code>cls_pred = np.mean(np.sort(np.ravel(sigmoid(mask)))[::-1][:top_n])</code></p>\n<p>The mean log loss across all folds is <strong>0.543</strong> for extravasation, and <strong>0.501</strong> for any_injury. </p>\n<p><strong>Training time</strong> - about <strong>4 hours</strong> per fold. </p>\n<h2>Validation</h2>\n<p><strong>StratifiedGroupKFold</strong> (stratification based on classification labels, grouping based on patients) with 4 folds. </p>\n<p>Mean log loss across all folds and all groups (kidney, liver, spleen, bowel, extravasation, any_injury), (which is the <strong>competition metric</strong>) is <strong>0.400</strong>. </p>\n<h2>Additional tricks</h2>\n<ul>\n<li>No post-processing.</li>\n<li>SWA on final checkpoints.</li>\n<li>EMA during training.</li>\n<li>Temporal shifting in classification to battle overfitting even more.</li>\n<li>Gradient checkpointing to have bigger batches (important for classification).</li>\n<li>memmap (uint8) using numpy to speed-up data reading and crop extraction. </li>\n<li>2 final subs: one minimizing competition loss, one maximizing AUC</li>\n</ul>\n<h2>Things that didn’t work</h2>\n<ul>\n<li>Samplers</li>\n<li>Heavier models (2+1D or Uniformer)</li>\n</ul>\n<h2>Final notes</h2>\n<p>During the final 2 days of the competition, we managed to improve the models for kidney, liver, and spleen from <strong>0.4</strong> to roughly <strong>0.38</strong>, which brought the overall loss from <strong>0.4</strong> to <strong>0.39</strong>, but made some errors in the submission process, which made them useless. The trick is simple - increase batch size. Usually we train with the batch of 14, but could increase it to 24 (with the help of A100 cards). </p>\n<p>The total <strong>training time</strong> (including all 4 folds for each stage) is around <strong>80 hours</strong> using a single RTX A6000 Ada.</p>\n<p>The total <strong>submission time</strong> is around <strong>8-9 hours</strong> using a single Tesla P100. </p>\n<p>We believe this solution could be pushed much further. However, we made the first sub (that isn’t sample submission or just a bunch of static predictions) 2 days before the competition ended, so that also played some role.</p>\n<p>P.S. Man that sucked to mess up the models for 0.39 :) </p>",
  "messages": [
    {
      "id": 2483713,
      "postDate": "2023-10-16T00:01:26.820Z",
      "content": "<p>Thanks to the organizers for such a great competition. Unfortunately, I couldn’t take part last year in a similar one due to the lack of hardware. However, this year is different, and I’m really happy with the results that we managed to achieve. </p>\n<h2>Problem</h2>\n<p>In this competition we were tasked with predicting the intensity of injuries for different abdominal organs. The available data consists of big CT images (in DICOM format) with partial supplemental segmentation annotations (in NIFTI format).</p>\n<p>In ML terms, all comes down to 3D segmentation / classification models, lack of data / annotations, and heavily-penalizing metric.</p>\n<h2>Summary</h2>\n<ul>\n<li><strong>U-Net Bi-Conv-LSTM</strong> segmentation for organs (kidney, liver, spleen, bowel) and separate model for Extravasation based on boxes from here (thanks a lot!)</li>\n<li><strong>Resnet 3D CSN</strong> for 3D crops separated into 2 stages: (kidney, liver, spleen) and (bowel)</li>\n</ul>\n<h2>3D Semantic Segmentation</h2>\n<p>The semantic segmentation part to identify organs is quite straightforward (compared to the later classification) and could be effectively performed by <strong>U-Net Bi-Conv-LSTM</strong> with Effnet_v2_b0 backbone and <strong>CE-Dice-Focal</strong> loss. </p>\n<p>To make things efficient, we train semantic segmentation on <strong>96x256x256 crops</strong> and predict the whole image using crops of size 96x256x256 with <strong>overlaps of 48x256x256</strong> (later on overlaps were removed to save time and space in inference).</p>\n<p>To elevate overfitting (it’s not that critical, especially compared to classification), we added <strong>geometric augmentations</strong> like ShiftScaleRotate, RandomBrightnessContrast, Flips, GridDistortion, ElasticTransform. </p>\n<p>On average, about 10% of total volume was dedicated to kidney, liver, spleen, and about 20% - to bowel.</p>\n<p>The <strong>macro dice score per image</strong> is around <strong>0.96</strong>. </p>\n<p><strong>Training time</strong> - about <strong>12 hours</strong>.</p>\n<h2>3D classification</h2>\n<p>Based on extracted crops from segmentation masks, we train 2 models: one for kidney, liver, spleen, one for bowel. </p>\n<p>The <strong>CSN models</strong> from mmaction proved to be very fast and accurate. In order to figure out how to deal with temporal dimension, several possibilities were explored, but in the end basic interpolation (<strong>3D resize</strong>) was used to convert crops to <strong>96x256x256 resolution</strong>.</p>\n<p>To battle overfitting (which is really severe even with CSN), <strong>intensive geometric augmentations</strong> were used, including ShiftScaleRotate, RandomBrightnessContrast, and 4 different types of Blurs.</p>\n<p>The mean competition loss across all folds is <strong>0.401</strong> for kidney, liver, spleen and 0<strong>.156</strong> for bowel.</p>\n<p><strong>Training time</strong> - about <strong>4 hours</strong> per kidney, liver, spleen fold; and <strong>8-10 hours</strong> per bowel fold. </p>\n<h2>3D classification for Extravasation</h2>\n<p>In order to make predictions for Extravasation, a segmentation model was utilized. The motivation is simple: if semantic segmentation model predict anything, there is Extravasation, and it should be reflected in the probabilities. </p>\n<p>To make the <strong>transition from semantic segmentation to classification</strong>, the following trick was used:</p>\n<ul>\n<li>Turn 3D mask to 1D </li>\n<li>Sort probabilities </li>\n<li>Take top_n probabilities</li>\n<li>Find mean values of them. That’s the probability for positive Extravasation.</li>\n</ul>\n<p>In pseudo-code:<br>\n<code>cls_pred = np.mean(np.sort(np.ravel(sigmoid(mask)))[::-1][:top_n])</code></p>\n<p>The mean log loss across all folds is <strong>0.543</strong> for extravasation, and <strong>0.501</strong> for any_injury. </p>\n<p><strong>Training time</strong> - about <strong>4 hours</strong> per fold. </p>\n<h2>Validation</h2>\n<p><strong>StratifiedGroupKFold</strong> (stratification based on classification labels, grouping based on patients) with 4 folds. </p>\n<p>Mean log loss across all folds and all groups (kidney, liver, spleen, bowel, extravasation, any_injury), (which is the <strong>competition metric</strong>) is <strong>0.400</strong>. </p>\n<h2>Additional tricks</h2>\n<ul>\n<li>No post-processing.</li>\n<li>SWA on final checkpoints.</li>\n<li>EMA during training.</li>\n<li>Temporal shifting in classification to battle overfitting even more.</li>\n<li>Gradient checkpointing to have bigger batches (important for classification).</li>\n<li>memmap (uint8) using numpy to speed-up data reading and crop extraction. </li>\n<li>2 final subs: one minimizing competition loss, one maximizing AUC</li>\n</ul>\n<h2>Things that didn’t work</h2>\n<ul>\n<li>Samplers</li>\n<li>Heavier models (2+1D or Uniformer)</li>\n</ul>\n<h2>Final notes</h2>\n<p>During the final 2 days of the competition, we managed to improve the models for kidney, liver, and spleen from <strong>0.4</strong> to roughly <strong>0.38</strong>, which brought the overall loss from <strong>0.4</strong> to <strong>0.39</strong>, but made some errors in the submission process, which made them useless. The trick is simple - increase batch size. Usually we train with the batch of 14, but could increase it to 24 (with the help of A100 cards). </p>\n<p>The total <strong>training time</strong> (including all 4 folds for each stage) is around <strong>80 hours</strong> using a single RTX A6000 Ada.</p>\n<p>The total <strong>submission time</strong> is around <strong>8-9 hours</strong> using a single Tesla P100. </p>\n<p>We believe this solution could be pushed much further. However, we made the first sub (that isn’t sample submission or just a bunch of static predictions) 2 days before the competition ended, so that also played some role.</p>\n<p>P.S. Man that sucked to mess up the models for 0.39 :) </p>",
      "rawMarkdown": "Thanks to the organizers for such a great competition. Unfortunately, I couldn’t take part last year in a similar one due to the lack of hardware. However, this year is different, and I’m really happy with the results that we managed to achieve. \n\n## Problem\nIn this competition we were tasked with predicting the intensity of injuries for different abdominal organs. The available data consists of big CT images (in DICOM format) with partial supplemental segmentation annotations (in NIFTI format).\n\nIn ML terms, all comes down to 3D segmentation / classification models, lack of data / annotations, and heavily-penalizing metric.\n\n## Summary\n- **U-Net Bi-Conv-LSTM** segmentation for organs (kidney, liver, spleen, bowel) and separate model for Extravasation based on boxes from here (thanks a lot!)\n- **Resnet 3D CSN** for 3D crops separated into 2 stages: (kidney, liver, spleen) and (bowel)\n\n\n## 3D Semantic Segmentation\n\nThe semantic segmentation part to identify organs is quite straightforward (compared to the later classification) and could be effectively performed by **U-Net Bi-Conv-LSTM** with Effnet_v2_b0 backbone and **CE-Dice-Focal** loss. \n\nTo make things efficient, we train semantic segmentation on **96x256x256 crops** and predict the whole image using crops of size 96x256x256 with **overlaps of 48x256x256** (later on overlaps were removed to save time and space in inference).\n\nTo elevate overfitting (it’s not that critical, especially compared to classification), we added **geometric augmentations** like ShiftScaleRotate, RandomBrightnessContrast, Flips, GridDistortion, ElasticTransform. \n\nOn average, about 10% of total volume was dedicated to kidney, liver, spleen, and about 20% - to bowel.\n\nThe **macro dice score per image** is around **0.96**. \n\n**Training time** - about **12 hours**.\n\n## 3D classification\n\nBased on extracted crops from segmentation masks, we train 2 models: one for kidney, liver, spleen, one for bowel. \n\nThe **CSN models** from mmaction proved to be very fast and accurate. In order to figure out how to deal with temporal dimension, several possibilities were explored, but in the end basic interpolation (**3D resize**) was used to convert crops to **96x256x256 resolution**.\n\nTo battle overfitting (which is really severe even with CSN), **intensive geometric augmentations** were used, including ShiftScaleRotate, RandomBrightnessContrast, and 4 different types of Blurs.\n\nThe mean competition loss across all folds is **0.401** for kidney, liver, spleen and 0**.156** for bowel.\n\n**Training time** - about **4 hours** per kidney, liver, spleen fold; and **8-10 hours** per bowel fold. \n\n## 3D classification for Extravasation\n\nIn order to make predictions for Extravasation, a segmentation model was utilized. The motivation is simple: if semantic segmentation model predict anything, there is Extravasation, and it should be reflected in the probabilities. \n\nTo make the **transition from semantic segmentation to classification**, the following trick was used:\n- Turn 3D mask to 1D \n- Sort probabilities \n- Take top_n probabilities\n- Find mean values of them. That’s the probability for positive Extravasation.\n\nIn pseudo-code:\n`cls_pred = np.mean(np.sort(np.ravel(sigmoid(mask)))[::-1][:top_n])`\n\nThe mean log loss across all folds is **0.543** for extravasation, and **0.501** for any_injury. \n\n**Training time** - about **4 hours** per fold. \n\n## Validation\n\n**StratifiedGroupKFold** (stratification based on classification labels, grouping based on patients) with 4 folds. \n\nMean log loss across all folds and all groups (kidney, liver, spleen, bowel, extravasation, any_injury), (which is the **competition metric**) is **0.400**. \n\n## Additional tricks\n\n- No post-processing.\n- SWA on final checkpoints.\n- EMA during training.\n- Temporal shifting in classification to battle overfitting even more.\n- Gradient checkpointing to have bigger batches (important for classification).\n- memmap (uint8) using numpy to speed-up data reading and crop extraction. \n- 2 final subs: one minimizing competition loss, one maximizing AUC\n\n\n## Things that didn’t work\n- Samplers\n- Heavier models (2+1D or Uniformer)\n\n## Final notes \nDuring the final 2 days of the competition, we managed to improve the models for kidney, liver, and spleen from **0.4** to roughly **0.38**, which brought the overall loss from **0.4** to **0.39**, but made some errors in the submission process, which made them useless. The trick is simple - increase batch size. Usually we train with the batch of 14, but could increase it to 24 (with the help of A100 cards). \n\nThe total **training time** (including all 4 folds for each stage) is around **80 hours** using a single RTX A6000 Ada.\n\nThe total **submission time** is around **8-9 hours** using a single Tesla P100. \n\nWe believe this solution could be pushed much further. However, we made the first sub (that isn’t sample submission or just a bunch of static predictions) 2 days before the competition ended, so that also played some role.\n\nP.S. Man that sucked to mess up the models for 0.39 :) \n",
      "votes": 25
    },
    {
      "id": 2484400,
      "postDate": "2023-10-16T13:13:09.547Z",
      "content": "<p>Congratulation for the high rank - especially regarding such a late first sub, and thank you for the clear explanations and well organized thoughts. </p>\n<p>The only thing I didn't get is the 3D classification for Extravasation. How do you train it to obtain the 3D mask that you process to get the probability? </p>",
      "rawMarkdown": "Congratulation for the high rank - especially regarding such a late first sub, and thank you for the clear explanations and well organized thoughts. \n\nThe only thing I didn't get is the 3D classification for Extravasation. How do you train it to obtain the 3D mask that you process to get the probability? \n",
      "votes": 1,
      "replies": [
        {
          "id": 2484436,
          "postDate": "2023-10-16T13:43:46.080Z",
          "content": "<p>Yeah, sorry for not being clear. </p>\n<p>I take bboxes from <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/441402\" target=\"_blank\">here</a>, treat bboxes as semantic segmentation masks, and train <strong>U-Net Bi-Conv-LSTM</strong> on it. </p>\n<p>Because the annotations are obviously not pixel-wise perfect, the Dice score is not perfect either (about 0.2-0.3 per image), but in practice it looks more or less like this: <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2755695%2Fcbe24ea20acdd591b9f831443853da31%2F2023-10-16%2016.42.48.jpg?generation=1697463778578613&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Yeah, sorry for not being clear. \n\nI take bboxes from [here](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/441402), treat bboxes as semantic segmentation masks, and train **U-Net Bi-Conv-LSTM** on it. \n\nBecause the annotations are obviously not pixel-wise perfect, the Dice score is not perfect either (about 0.2-0.3 per image), but in practice it looks more or less like this: ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2755695%2Fcbe24ea20acdd591b9f831443853da31%2F2023-10-16%2016.42.48.jpg?generation=1697463778578613&alt=media)"
        }
      ]
    },
    {
      "id": 2484273,
      "postDate": "2023-10-16T11:01:20.037Z",
      "content": "<p>Congratulations on topping the LB at 16th Rank. Thanks for sharing the details of your solution. </p>",
      "rawMarkdown": "Congratulations on topping the LB at 16th Rank. Thanks for sharing the details of your solution. ",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2484400,
      "author_name": "Kiurtis",
      "author_url": "",
      "post_date": "2023-10-16T13:13:09.547000",
      "content": "<p>Congratulation for the high rank - especially regarding such a late first sub, and thank you for the clear explanations and well organized thoughts. </p>\n<p>The only thing I didn't get is the 3D classification for Extravasation. How do you train it to obtain the 3D mask that you process to get the probability? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2484436,
          "author_name": "Ivan Panshin",
          "author_url": "",
          "post_date": "2023-10-16T13:43:46.080000",
          "content": "<p>Yeah, sorry for not being clear. </p>\n<p>I take bboxes from <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/441402\" target=\"_blank\">here</a>, treat bboxes as semantic segmentation masks, and train <strong>U-Net Bi-Conv-LSTM</strong> on it. </p>\n<p>Because the annotations are obviously not pixel-wise perfect, the Dice score is not perfect either (about 0.2-0.3 per image), but in practice it looks more or less like this: <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2755695%2Fcbe24ea20acdd591b9f831443853da31%2F2023-10-16%2016.42.48.jpg?generation=1697463778578613&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2484273,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-10-16T11:01:20.037000",
      "content": "<p>Congratulations on topping the LB at 16th Rank. Thanks for sharing the details of your solution. </p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2483713": "Thanks to the organizers for such a great competition. Unfortunately, I couldn’t take part last year in a similar one due to the lack of hardware. However, this year is different, and I’m really happy with the results that we managed to achieve. \n\n## Problem\nIn this competition we were tasked with predicting the intensity of injuries for different abdominal organs. The available data consists of big CT images (in DICOM format) with partial supplemental segmentation annotations (in NIFTI format).\n\nIn ML terms, all comes down to 3D segmentation / classification models, lack of data / annotations, and heavily-penalizing metric.\n\n## Summary\n- **U-Net Bi-Conv-LSTM** segmentation for organs (kidney, liver, spleen, bowel) and separate model for Extravasation based on boxes from here (thanks a lot!)\n- **Resnet 3D CSN** for 3D crops separated into 2 stages: (kidney, liver, spleen) and (bowel)\n\n\n## 3D Semantic Segmentation\n\nThe semantic segmentation part to identify organs is quite straightforward (compared to the later classification) and could be effectively performed by **U-Net Bi-Conv-LSTM** with Effnet_v2_b0 backbone and **CE-Dice-Focal** loss. \n\nTo make things efficient, we train semantic segmentation on **96x256x256 crops** and predict the whole image using crops of size 96x256x256 with **overlaps of 48x256x256** (later on overlaps were removed to save time and space in inference).\n\nTo elevate overfitting (it’s not that critical, especially compared to classification), we added **geometric augmentations** like ShiftScaleRotate, RandomBrightnessContrast, Flips, GridDistortion, ElasticTransform. \n\nOn average, about 10% of total volume was dedicated to kidney, liver, spleen, and about 20% - to bowel.\n\nThe **macro dice score per image** is around **0.96**. \n\n**Training time** - about **12 hours**.\n\n## 3D classification\n\nBased on extracted crops from segmentation masks, we train 2 models: one for kidney, liver, spleen, one for bowel. \n\nThe **CSN models** from mmaction proved to be very fast and accurate. In order to figure out how to deal with temporal dimension, several possibilities were explored, but in the end basic interpolation (**3D resize**) was used to convert crops to **96x256x256 resolution**.\n\nTo battle overfitting (which is really severe even with CSN), **intensive geometric augmentations** were used, including ShiftScaleRotate, RandomBrightnessContrast, and 4 different types of Blurs.\n\nThe mean competition loss across all folds is **0.401** for kidney, liver, spleen and 0**.156** for bowel.\n\n**Training time** - about **4 hours** per kidney, liver, spleen fold; and **8-10 hours** per bowel fold. \n\n## 3D classification for Extravasation\n\nIn order to make predictions for Extravasation, a segmentation model was utilized. The motivation is simple: if semantic segmentation model predict anything, there is Extravasation, and it should be reflected in the probabilities. \n\nTo make the **transition from semantic segmentation to classification**, the following trick was used:\n- Turn 3D mask to 1D \n- Sort probabilities \n- Take top_n probabilities\n- Find mean values of them. That’s the probability for positive Extravasation.\n\nIn pseudo-code:\n`cls_pred = np.mean(np.sort(np.ravel(sigmoid(mask)))[::-1][:top_n])`\n\nThe mean log loss across all folds is **0.543** for extravasation, and **0.501** for any_injury. \n\n**Training time** - about **4 hours** per fold. \n\n## Validation\n\n**StratifiedGroupKFold** (stratification based on classification labels, grouping based on patients) with 4 folds. \n\nMean log loss across all folds and all groups (kidney, liver, spleen, bowel, extravasation, any_injury), (which is the **competition metric**) is **0.400**. \n\n## Additional tricks\n\n- No post-processing.\n- SWA on final checkpoints.\n- EMA during training.\n- Temporal shifting in classification to battle overfitting even more.\n- Gradient checkpointing to have bigger batches (important for classification).\n- memmap (uint8) using numpy to speed-up data reading and crop extraction. \n- 2 final subs: one minimizing competition loss, one maximizing AUC\n\n\n## Things that didn’t work\n- Samplers\n- Heavier models (2+1D or Uniformer)\n\n## Final notes \nDuring the final 2 days of the competition, we managed to improve the models for kidney, liver, and spleen from **0.4** to roughly **0.38**, which brought the overall loss from **0.4** to **0.39**, but made some errors in the submission process, which made them useless. The trick is simple - increase batch size. Usually we train with the batch of 14, but could increase it to 24 (with the help of A100 cards). \n\nThe total **training time** (including all 4 folds for each stage) is around **80 hours** using a single RTX A6000 Ada.\n\nThe total **submission time** is around **8-9 hours** using a single Tesla P100. \n\nWe believe this solution could be pushed much further. However, we made the first sub (that isn’t sample submission or just a bunch of static predictions) 2 days before the competition ended, so that also played some role.\n\nP.S. Man that sucked to mess up the models for 0.39 :) \n",
    "2484400": "Congratulation for the high rank - especially regarding such a late first sub, and thank you for the clear explanations and well organized thoughts. \n\nThe only thing I didn't get is the 3D classification for Extravasation. How do you train it to obtain the 3D mask that you process to get the probability? \n",
    "2484273": "Congratulations on topping the LB at 16th Rank. Thanks for sharing the details of your solution. "
  }
}