{
  "id": 447449,
  "title": "1st Place Solution: Team Oxygen",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/447449",
  "author_name": "Nischay Dhankhar",
  "post_date": "2023-10-16T00:02:45.234000",
  "votes": 133,
  "comment_count": 49,
  "views": 0,
  "content": "<p>Firstly, Thank you RSNA for hosting another interesting competition &amp; my teammates <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> - formation of <strong>Team oxygen</strong> ? :)) It was amazing to be #1 on public leaderboard for almost a month. I am sharing a quick overview of our solution, we will release the entire solution soon. It was really fun competing for #1 with <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>  <br>\n<strong>Edit: Full solution published.</strong>  <br>\nHere is the inference code you may refer: <a href=\"https://www.kaggle.com/nischaydnk/rsna-super-mega-lb-ensemble\" target=\"_blank\">link</a> <br>\nOur GitHub repo w/ all preprocessing + training code: <a href=\"https://github.com/Nischaydnk/RSNA-2023-1st-place-solution\" target=\"_blank\">link</a><br>\nDemo Inference notebook: <a href=\"https://www.kaggle.com/code/haqishen/rsna-2023-1st-place-best-model-infer-cleaned\" target=\"_blank\">link</a></p>\n<h4><strong>Split used:</strong> 4 Fold GroupKFold ( Patient Level)</h4>\n<h2><strong>Our solution is divided into three parts:</strong></h2>\n<p><strong>Part 1:</strong> 3D segmentation for generating masks / crops [Stage 1]<br>\n<strong>Part 2:</strong> 2D CNN + RNN based approach for Kidney, Liver, Spleen &amp; Bowel [Stage 2]<br>\n<strong>Part 3:</strong> 2D CNN + RNN based approach for Bowel + Extravasation [Stage 2]</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fee28ed8eef7827d8f2cc69601875e5c2%2FScreenshot%202023-10-22%20at%2011.54.15%20AM.png?generation=1697955883461262&amp;alt=media\" alt=\"\"></p>\n<h2><strong>Data Preprocessing:</strong></h2>\n<p>Here comes the key part of our solution, we will describe it later in more depth. <strong>Note:</strong> <em>All models were trained on image size 384 x 384. We use datasets preprocessing from <a href=\"https://www.kaggle.com/TheoVeol\" target=\"_blank\">@TheoVeol</a> and our data which we made rescale dicoms and applying soft-tissue windowing.</em></p>\n<p>We take a patient/study, we run a 3d segmentation model on it, it outputs masks for each slice, we make a study-level crop here based on boundaries of organs - liver, spleen, kidney &amp; liver. </p>\n<p>Next, we make volumes from the patient, each volume extracted with equi-distant 96 slices for a study which is then reshaped to (32, 3, image_size, image_size) in a 2.5D manner for training CNN based models.</p>\n<p>3 channels are formed by using the adjacent slices.</p>\n<p>All our model takes in input in shape (2, 32, 3, height, width) and outputs it as (2, 32, n_classes) as the targets are also kept in shape (2, 32, n_classes).</p>\n<p>To make the targets, we need 2 things, patient-level target of each organ and how much the organ is visible compared to its maximum visibility, this data is available after normalizing segmentation model masks in 0-1 based on number of positive pixels</p>\n<p>Then we multiply targets * patient-level target for each middle slice of the sequence and that is our label</p>\n<p>For example if a patient has label 0 for liver-injury and the liver visibility is as follows in the slice sequence</p>\n<p>[0., 0., 0., 0.01, 0.05, 0.1, 0.23, 0.5, 0.7, 0.95, 0.99, 1., 0.95, 0.8, 0.4 …. 0. ,0., 0.]</p>\n<p>We multiply it with label which is currently 0 results in an all zeros list as output, but if target label for liver-injury was 1, then we use the list mentioned above as our soft labels.</p>\n<h2><strong>Stage2: 2.5D Approach ( 2D CNN + RNN):</strong></h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fe8df4581839b1fa7dcadf68fe2a715a1%2FScreenshot%202023-10-22%20at%205.31.23%20AM.png?generation=1697935695484067&amp;alt=media\" alt=\"\"></p>\n<p>In stage 2, we trained our models using the volumes either based on our windowing or theo's preprocessing approach and the masks/crops generated from 3D segmentation approach. Each model is trained for multiple tasks (segmentation + classification). For all 32 sequences, we predicted slice level masks and sigmoid predictions. Further, simple maximum aggregation is applied on sigmoid predictions to fetch study level prediction used in submissions. </p>\n<p>For training our models, some common settings were:</p>\n<ul>\n<li><strong>Learning rate:</strong> (1e-4 to 4e-4) range</li>\n<li><strong>Optimizer:</strong> AdamW</li>\n<li><strong>Scheduler:</strong> Cosine Annealing w/ Warmup </li>\n<li><strong>Loss:</strong> BCE Loss for Classification, Dice Loss for segmentation</li>\n</ul>\n<h3><strong>Auxiliary Segmentation Loss:</strong></h3>\n<p>One of the key things which made our training much more stable and helped in improving scores was using auxiliary losses based on segmentation. </p>\n<p>Encoder was kept same for both classification &amp; segmentation decoders,  we used two types of segmentation head:</p>\n<ul>\n<li><strong><em>Unet based decoder</em></strong> for generating masks</li>\n<li><strong><em>2D-CNN</em></strong> based head </li>\n</ul>\n<pre><code>nn.Sequential(\n            nn.Conv2d(nb_ft, 128, =3, =1),\n            nn.BatchNorm2d(128),\n            nn.ReLU(=),\n            nn.Conv2d(128, 128, =3, =1),\n            nn.BatchNorm2d(128),\n            nn.ReLU(=),\n            nn.Conv2d(128, 4, =1, =0),\n        )\n</code></pre>\n<pre><code>        self = self(true_encoder)\n        self = self(true_encoder)\n</code></pre>\n<p>We used the feature maps generated mainly from last and 2nd last blocks of the backbones &amp; apply dice loss on the predicted masks &amp; true masks. This trick gave us around +0.01 to +0.03 boost in our models. We used similar technique in Covid 19 detection competition held few years back, you can also refer my solution for more detailed use of auxiliary loss &amp; code snippets. <br>\n<a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/266571\" target=\"_blank\">link of discussion</a></p>\n<p>Here is an example code for applying aux loss:</p>\n<pre><code></code></pre>\n<h3><strong>Architectures used in Final ensemble:</strong></h3>\n<ul>\n<li>Coat Lite Medium w/ GRU - <a href=\"https://github.com/mlpc-ucsd/CoaT\" target=\"_blank\">original source code</a></li>\n<li>Coat Lite Small w/ GRU!</li>\n<li>Efficientnet v2s w/ GRU [Timm]</li>\n</ul>\n<h3><strong>Augmentations:</strong></h3>\n<p>We couldn't come up with several augmentations to use, but these were the ones which we used in our training.</p>\n<pre><code>        .Perspective(p=.),\n        .HorizontalFlip(p=.),\n        .VerticalFlip(p=.),\n        .Rotate(p=., limit=(-, )),\n</code></pre>\n<h2><strong>Post Processing / Ensemble:</strong></h2>\n<p>Final ensemble for all organs model includes <strong>multiple Coat medium and V2s based models</strong> trained on either 4 Folds or Full data. </p>\n<p>For extravasation, We mainly used Coat Small and v2s in ensemble. <br>\n<strong>No major postprocessing</strong> was applied except for tuning scaling factors based on CV scores.<br>\nTo get the predictions, we aggregated the model outputs at slice level and simply took the maximum value for each patient.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fd6fa2cc524588b85b82906cccb6552bf%2FScreenshot%202023-10-22%20at%206.02.32%20AM.png?generation=1697936329146043&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fb69766140607127e717966085c927bce%2FScreenshot%202023-10-22%20at%206.07.13%20AM.png?generation=1697936371125187&amp;alt=media\" alt=\"\"></p>\n<h4><strong>Ensemble:</strong></h4>\n<p>Within folds of each models, we are doing slice level ensemble.<br>\nFor different architectures &amp; cross data models (theo/ours), we did ensemble after the max aggregation. </p>\n<h4><strong>Best Ensemble OOF CV</strong>: 0.31x</h4>\n<h4><strong>Best single model 4 fold OOF CV</strong>: 0.326 [Coat lite Medium]</h4>\n<p>Organ level OOF for single model looks like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F9143afd07ba3069f2b2259b1d8fe80eb%2FScreenshot%202023-10-16%20at%204.05.07%20AM.png?generation=1697413214073394&amp;alt=media\" alt=\"\"></p>\n<p>Thank you. </p>\n<p>EDIT 1: 3D segmentation code: <a href=\"https://www.kaggle.com/code/haqishen/rsna-2023-1st-place-solution-train-3d-seg/notebook\" target=\"_blank\">notebook link</a></p>",
  "messages": [
    {
      "id": 2483715,
      "postDate": "2023-10-16T00:02:45.233Z",
      "content": "<p>Firstly, Thank you RSNA for hosting another interesting competition &amp; my teammates <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> - formation of <strong>Team oxygen</strong> ? :)) It was amazing to be #1 on public leaderboard for almost a month. I am sharing a quick overview of our solution, we will release the entire solution soon. It was really fun competing for #1 with <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>  <br>\n<strong>Edit: Full solution published.</strong>  <br>\nHere is the inference code you may refer: <a href=\"https://www.kaggle.com/nischaydnk/rsna-super-mega-lb-ensemble\" target=\"_blank\">link</a> <br>\nOur GitHub repo w/ all preprocessing + training code: <a href=\"https://github.com/Nischaydnk/RSNA-2023-1st-place-solution\" target=\"_blank\">link</a><br>\nDemo Inference notebook: <a href=\"https://www.kaggle.com/code/haqishen/rsna-2023-1st-place-best-model-infer-cleaned\" target=\"_blank\">link</a></p>\n<h4><strong>Split used:</strong> 4 Fold GroupKFold ( Patient Level)</h4>\n<h2><strong>Our solution is divided into three parts:</strong></h2>\n<p><strong>Part 1:</strong> 3D segmentation for generating masks / crops [Stage 1]<br>\n<strong>Part 2:</strong> 2D CNN + RNN based approach for Kidney, Liver, Spleen &amp; Bowel [Stage 2]<br>\n<strong>Part 3:</strong> 2D CNN + RNN based approach for Bowel + Extravasation [Stage 2]</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fee28ed8eef7827d8f2cc69601875e5c2%2FScreenshot%202023-10-22%20at%2011.54.15%20AM.png?generation=1697955883461262&amp;alt=media\" alt=\"\"></p>\n<h2><strong>Data Preprocessing:</strong></h2>\n<p>Here comes the key part of our solution, we will describe it later in more depth. <strong>Note:</strong> <em>All models were trained on image size 384 x 384. We use datasets preprocessing from <a href=\"https://www.kaggle.com/TheoVeol\" target=\"_blank\">@TheoVeol</a> and our data which we made rescale dicoms and applying soft-tissue windowing.</em></p>\n<p>We take a patient/study, we run a 3d segmentation model on it, it outputs masks for each slice, we make a study-level crop here based on boundaries of organs - liver, spleen, kidney &amp; liver. </p>\n<p>Next, we make volumes from the patient, each volume extracted with equi-distant 96 slices for a study which is then reshaped to (32, 3, image_size, image_size) in a 2.5D manner for training CNN based models.</p>\n<p>3 channels are formed by using the adjacent slices.</p>\n<p>All our model takes in input in shape (2, 32, 3, height, width) and outputs it as (2, 32, n_classes) as the targets are also kept in shape (2, 32, n_classes).</p>\n<p>To make the targets, we need 2 things, patient-level target of each organ and how much the organ is visible compared to its maximum visibility, this data is available after normalizing segmentation model masks in 0-1 based on number of positive pixels</p>\n<p>Then we multiply targets * patient-level target for each middle slice of the sequence and that is our label</p>\n<p>For example if a patient has label 0 for liver-injury and the liver visibility is as follows in the slice sequence</p>\n<p>[0., 0., 0., 0.01, 0.05, 0.1, 0.23, 0.5, 0.7, 0.95, 0.99, 1., 0.95, 0.8, 0.4 …. 0. ,0., 0.]</p>\n<p>We multiply it with label which is currently 0 results in an all zeros list as output, but if target label for liver-injury was 1, then we use the list mentioned above as our soft labels.</p>\n<h2><strong>Stage2: 2.5D Approach ( 2D CNN + RNN):</strong></h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fe8df4581839b1fa7dcadf68fe2a715a1%2FScreenshot%202023-10-22%20at%205.31.23%20AM.png?generation=1697935695484067&amp;alt=media\" alt=\"\"></p>\n<p>In stage 2, we trained our models using the volumes either based on our windowing or theo's preprocessing approach and the masks/crops generated from 3D segmentation approach. Each model is trained for multiple tasks (segmentation + classification). For all 32 sequences, we predicted slice level masks and sigmoid predictions. Further, simple maximum aggregation is applied on sigmoid predictions to fetch study level prediction used in submissions. </p>\n<p>For training our models, some common settings were:</p>\n<ul>\n<li><strong>Learning rate:</strong> (1e-4 to 4e-4) range</li>\n<li><strong>Optimizer:</strong> AdamW</li>\n<li><strong>Scheduler:</strong> Cosine Annealing w/ Warmup </li>\n<li><strong>Loss:</strong> BCE Loss for Classification, Dice Loss for segmentation</li>\n</ul>\n<h3><strong>Auxiliary Segmentation Loss:</strong></h3>\n<p>One of the key things which made our training much more stable and helped in improving scores was using auxiliary losses based on segmentation. </p>\n<p>Encoder was kept same for both classification &amp; segmentation decoders,  we used two types of segmentation head:</p>\n<ul>\n<li><strong><em>Unet based decoder</em></strong> for generating masks</li>\n<li><strong><em>2D-CNN</em></strong> based head </li>\n</ul>\n<pre><code>nn.Sequential(\n            nn.Conv2d(nb_ft, 128, =3, =1),\n            nn.BatchNorm2d(128),\n            nn.ReLU(=),\n            nn.Conv2d(128, 128, =3, =1),\n            nn.BatchNorm2d(128),\n            nn.ReLU(=),\n            nn.Conv2d(128, 4, =1, =0),\n        )\n</code></pre>\n<pre><code>        self = self(true_encoder)\n        self = self(true_encoder)\n</code></pre>\n<p>We used the feature maps generated mainly from last and 2nd last blocks of the backbones &amp; apply dice loss on the predicted masks &amp; true masks. This trick gave us around +0.01 to +0.03 boost in our models. We used similar technique in Covid 19 detection competition held few years back, you can also refer my solution for more detailed use of auxiliary loss &amp; code snippets. <br>\n<a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/266571\" target=\"_blank\">link of discussion</a></p>\n<p>Here is an example code for applying aux loss:</p>\n<pre><code></code></pre>\n<h3><strong>Architectures used in Final ensemble:</strong></h3>\n<ul>\n<li>Coat Lite Medium w/ GRU - <a href=\"https://github.com/mlpc-ucsd/CoaT\" target=\"_blank\">original source code</a></li>\n<li>Coat Lite Small w/ GRU!</li>\n<li>Efficientnet v2s w/ GRU [Timm]</li>\n</ul>\n<h3><strong>Augmentations:</strong></h3>\n<p>We couldn't come up with several augmentations to use, but these were the ones which we used in our training.</p>\n<pre><code>        .Perspective(p=.),\n        .HorizontalFlip(p=.),\n        .VerticalFlip(p=.),\n        .Rotate(p=., limit=(-, )),\n</code></pre>\n<h2><strong>Post Processing / Ensemble:</strong></h2>\n<p>Final ensemble for all organs model includes <strong>multiple Coat medium and V2s based models</strong> trained on either 4 Folds or Full data. </p>\n<p>For extravasation, We mainly used Coat Small and v2s in ensemble. <br>\n<strong>No major postprocessing</strong> was applied except for tuning scaling factors based on CV scores.<br>\nTo get the predictions, we aggregated the model outputs at slice level and simply took the maximum value for each patient.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fd6fa2cc524588b85b82906cccb6552bf%2FScreenshot%202023-10-22%20at%206.02.32%20AM.png?generation=1697936329146043&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fb69766140607127e717966085c927bce%2FScreenshot%202023-10-22%20at%206.07.13%20AM.png?generation=1697936371125187&amp;alt=media\" alt=\"\"></p>\n<h4><strong>Ensemble:</strong></h4>\n<p>Within folds of each models, we are doing slice level ensemble.<br>\nFor different architectures &amp; cross data models (theo/ours), we did ensemble after the max aggregation. </p>\n<h4><strong>Best Ensemble OOF CV</strong>: 0.31x</h4>\n<h4><strong>Best single model 4 fold OOF CV</strong>: 0.326 [Coat lite Medium]</h4>\n<p>Organ level OOF for single model looks like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F9143afd07ba3069f2b2259b1d8fe80eb%2FScreenshot%202023-10-16%20at%204.05.07%20AM.png?generation=1697413214073394&amp;alt=media\" alt=\"\"></p>\n<p>Thank you. </p>\n<p>EDIT 1: 3D segmentation code: <a href=\"https://www.kaggle.com/code/haqishen/rsna-2023-1st-place-solution-train-3d-seg/notebook\" target=\"_blank\">notebook link</a></p>",
      "rawMarkdown": "Firstly, Thank you RSNA for hosting another interesting competition & my teammates @haqishen @harshitsheoran - formation of **Team oxygen** ? :)) It was amazing to be #1 on public leaderboard for almost a month. I am sharing a quick overview of our solution, we will release the entire solution soon. It was really fun competing for #1 with @theoviel  \n**Edit: Full solution published.**  \nHere is the inference code you may refer: [link](https://www.kaggle.com/nischaydnk/rsna-super-mega-lb-ensemble) \nOur GitHub repo w/ all preprocessing + training code: [link](https://github.com/Nischaydnk/RSNA-2023-1st-place-solution)\nDemo Inference notebook: [link](https://www.kaggle.com/code/haqishen/rsna-2023-1st-place-best-model-infer-cleaned)\n\n#### **Split used:** 4 Fold GroupKFold ( Patient Level)\n\n## **Our solution is divided into three parts:**\n**Part 1:** 3D segmentation for generating masks / crops [Stage 1]\n**Part 2:** 2D CNN + RNN based approach for Kidney, Liver, Spleen & Bowel [Stage 2]\n**Part 3:** 2D CNN + RNN based approach for Bowel + Extravasation [Stage 2]\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fee28ed8eef7827d8f2cc69601875e5c2%2FScreenshot%202023-10-22%20at%2011.54.15%20AM.png?generation=1697955883461262&alt=media)\n\n## **Data Preprocessing:**\n\nHere comes the key part of our solution, we will describe it later in more depth. **Note:** *All models were trained on image size 384 x 384. We use datasets preprocessing from @TheoVeol and our data which we made rescale dicoms and applying soft-tissue windowing.*\n\nWe take a patient/study, we run a 3d segmentation model on it, it outputs masks for each slice, we make a study-level crop here based on boundaries of organs - liver, spleen, kidney & liver. \n\nNext, we make volumes from the patient, each volume extracted with equi-distant 96 slices for a study which is then reshaped to (32, 3, image_size, image_size) in a 2.5D manner for training CNN based models.\n\n3 channels are formed by using the adjacent slices.\n\nAll our model takes in input in shape (2, 32, 3, height, width) and outputs it as (2, 32, n_classes) as the targets are also kept in shape (2, 32, n_classes).\n\nTo make the targets, we need 2 things, patient-level target of each organ and how much the organ is visible compared to its maximum visibility, this data is available after normalizing segmentation model masks in 0-1 based on number of positive pixels\n\nThen we multiply targets * patient-level target for each middle slice of the sequence and that is our label\n\nFor example if a patient has label 0 for liver-injury and the liver visibility is as follows in the slice sequence\n\n[0., 0., 0., 0.01, 0.05, 0.1, 0.23, 0.5, 0.7, 0.95, 0.99, 1., 0.95, 0.8, 0.4 …. 0. ,0., 0.]\n\nWe multiply it with label which is currently 0 results in an all zeros list as output, but if target label for liver-injury was 1, then we use the list mentioned above as our soft labels.\n\n\n## **Stage2: 2.5D Approach ( 2D CNN + RNN):**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fe8df4581839b1fa7dcadf68fe2a715a1%2FScreenshot%202023-10-22%20at%205.31.23%20AM.png?generation=1697935695484067&alt=media)\n\nIn stage 2, we trained our models using the volumes either based on our windowing or theo's preprocessing approach and the masks/crops generated from 3D segmentation approach. Each model is trained for multiple tasks (segmentation + classification). For all 32 sequences, we predicted slice level masks and sigmoid predictions. Further, simple maximum aggregation is applied on sigmoid predictions to fetch study level prediction used in submissions. \n\nFor training our models, some common settings were:\n- **Learning rate:** (1e-4 to 4e-4) range\n- **Optimizer:** AdamW\n- **Scheduler:** Cosine Annealing w/ Warmup \n- **Loss:** BCE Loss for Classification, Dice Loss for segmentation\n\n\n\n###**Auxiliary Segmentation Loss:** \nOne of the key things which made our training much more stable and helped in improving scores was using auxiliary losses based on segmentation. \n\nEncoder was kept same for both classification & segmentation decoders,  we used two types of segmentation head:\n- ***Unet based decoder*** for generating masks\n- ***2D-CNN*** based head \n```\nnn.Sequential(\n            nn.Conv2d(nb_ft, 128, kernel_size=3, padding=1),\n            nn.BatchNorm2d(128),\n            nn.ReLU(inplace=True),\n            nn.Conv2d(128, 128, kernel_size=3, padding=1),\n            nn.BatchNorm2d(128),\n            nn.ReLU(inplace=True),\n            nn.Conv2d(128, 4, kernel_size=1, padding=0),\n        )\n```\n```\n        self.mask_head_3 = self.get_mask_head(true_encoder.feature_info[-2]['num_chs'])\n        self.mask_head_4 = self.get_mask_head(true_encoder.feature_info[-1]['num_chs'])\n```\nWe used the feature maps generated mainly from last and 2nd last blocks of the backbones & apply dice loss on the predicted masks & true masks. This trick gave us around +0.01 to +0.03 boost in our models. We used similar technique in Covid 19 detection competition held few years back, you can also refer my solution for more detailed use of auxiliary loss & code snippets. \n[link of discussion](https://www.kaggle.com/c/siim-covid19-detection/discussion/266571)\n\nHere is an example code for applying aux loss:\n```\nclass CustomLoss(nn.Module):\n    def __init__(self):\n        super(CustomLoss, self).__init__()\n        #self.bce = nn.BCEWithLogitsLoss(pos_weight=torch.as_tensor([2.318]).cuda()).cuda()\n        self.bce = nn.BCEWithLogitsLoss()\n        self.dice = smp.losses.DiceLoss(smp.losses.MULTILABEL_MODE, from_logits=True)\n\n    def forward(self, outputs, targets, masks_outputs, masks_outputs2, masks_targets):\n        loss1 = self.bce(outputs, targets.float())\n\n        masks_outputs = masks_outputs.float()\n        masks_outputs2 = masks_outputs2.float()\n\n        masks_targets = masks_targets.float().flatten(0, 1)\n\n        loss2 = self.dice(masks_outputs, masks_targets) + self.dice(masks_outputs2, masks_targets)\n\n\n        loss = loss1 + (loss2 * 0.125) \n\n        return loss\n```\n\n\n### **Architectures used in Final ensemble:**\n- Coat Lite Medium w/ GRU - [original source code](https://github.com/mlpc-ucsd/CoaT)\n- Coat Lite Small w/ GRU!\n- Efficientnet v2s w/ GRU [Timm]\n\n\n### **Augmentations:**\n\nWe couldn't come up with several augmentations to use, but these were the ones which we used in our training.\n\n```\n        A.Perspective(p=0.5),\n        A.HorizontalFlip(p=0.5),\n        A.VerticalFlip(p=0.5),\n        A.Rotate(p=0.5, limit=(-25, 25)),\n```\n\n## **Post Processing / Ensemble:**\n\nFinal ensemble for all organs model includes **multiple Coat medium and V2s based models** trained on either 4 Folds or Full data. \n\nFor extravasation, We mainly used Coat Small and v2s in ensemble. \n**No major postprocessing** was applied except for tuning scaling factors based on CV scores.\nTo get the predictions, we aggregated the model outputs at slice level and simply took the maximum value for each patient.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fd6fa2cc524588b85b82906cccb6552bf%2FScreenshot%202023-10-22%20at%206.02.32%20AM.png?generation=1697936329146043&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fb69766140607127e717966085c927bce%2FScreenshot%202023-10-22%20at%206.07.13%20AM.png?generation=1697936371125187&alt=media)\n\n#### **Ensemble:**\n\nWithin folds of each models, we are doing slice level ensemble.\nFor different architectures & cross data models (theo/ours), we did ensemble after the max aggregation. \n\n#### **Best Ensemble OOF CV**: 0.31x\n#### **Best single model 4 fold OOF CV**: 0.326 [Coat lite Medium]\n\nOrgan level OOF for single model looks like this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F9143afd07ba3069f2b2259b1d8fe80eb%2FScreenshot%202023-10-16%20at%204.05.07%20AM.png?generation=1697413214073394&alt=media)\n\nThank you. \n\n\nEDIT 1: 3D segmentation code: [notebook link](https://www.kaggle.com/code/haqishen/rsna-2023-1st-place-solution-train-3d-seg/notebook)\n",
      "votes": 133
    },
    {
      "id": 2484321,
      "postDate": "2023-10-16T11:57:05.177Z",
      "content": "<p>Congratulation and thanks for sharing those insights ! Waiting for the full solution. </p>",
      "rawMarkdown": "Congratulation and thanks for sharing those insights ! Waiting for the full solution. ",
      "votes": 5
    },
    {
      "id": 2484264,
      "postDate": "2023-10-16T10:56:02.173Z",
      "content": "<p>Congratulations on winning the top position. Thanks for sharing the solution. </p>",
      "rawMarkdown": "Congratulations on winning the top position. Thanks for sharing the solution. ",
      "votes": 3
    },
    {
      "id": 2483894,
      "postDate": "2023-10-16T04:43:40.883Z",
      "content": "<p>😀😀😀<br>\nCongrats and Big Thanks to my wonderful teammates!! <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> </p>",
      "rawMarkdown": "😀😀😀\nCongrats and Big Thanks to my wonderful teammates!! @nischaydnk @harshitsheoran ",
      "votes": 4
    },
    {
      "id": 2513825,
      "postDate": "2023-11-05T18:52:36.063Z",
      "content": "<p>Congratulations on securing the top position .Looking forward to see full code</p>",
      "rawMarkdown": "Congratulations on securing the top position .Looking forward to see full code\n\n\n",
      "votes": 1
    },
    {
      "id": 2486869,
      "postDate": "2023-10-18T07:56:00.003Z",
      "content": "<p>Tq for your work </p>",
      "rawMarkdown": "Tq for your work ",
      "votes": 1
    },
    {
      "id": 2486059,
      "postDate": "2023-10-17T16:49:21.100Z",
      "content": "<p>Congratulations, team! Nice job :) Looking forward to seeing the final notebook. </p>",
      "rawMarkdown": "Congratulations, team! Nice job :) Looking forward to seeing the final notebook. ",
      "votes": 1
    },
    {
      "id": 2485202,
      "postDate": "2023-10-17T02:52:12.537Z",
      "content": "<p>Congrats!!! thinks for sharing！</p>",
      "rawMarkdown": "Congrats!!! thinks for sharing！",
      "votes": 1
    },
    {
      "id": 2485163,
      "postDate": "2023-10-17T01:55:20.390Z",
      "content": "<p>Congratulations！I would appreciate it if you share the code。</p>",
      "rawMarkdown": "Congratulations！I would appreciate it if you share the code。",
      "votes": 1
    },
    {
      "id": 2483735,
      "postDate": "2023-10-16T00:21:04.300Z",
      "content": "<p>Congratulations on the win! Your solution is very elegant :)</p>\n<p>For the auxiliary segmentation objective, did you use the masks generated by your 3D segmentation models as the ground truth labels? If so, does that mean you also added a decoder module to your classification models to predict segmentation masks?</p>",
      "rawMarkdown": "Congratulations on the win! Your solution is very elegant :)\n\nFor the auxiliary segmentation objective, did you use the masks generated by your 3D segmentation models as the ground truth labels? If so, does that mean you also added a decoder module to your classification models to predict segmentation masks?",
      "votes": 1,
      "replies": [
        {
          "id": 2483747,
          "postDate": "2023-10-16T00:46:12.980Z",
          "content": "<p>Thank you, yes we used 3D segmentation models output as ground truth, while training classification models, along with linear classifier head, we used 1 / 2 separate mask heads/decoders. So, our loss function looks something like this.</p>\n<pre><code></code></pre>",
          "rawMarkdown": "Thank you, yes we used 3D segmentation models output as ground truth, while training classification models, along with linear classifier head, we used 1 / 2 separate mask heads/decoders. So, our loss function looks something like this.\n```\nclass CustomLoss(nn.Module):\n    def __init__(self):\n        super(CustomLoss, self).__init__()\n        #self.bce = nn.BCEWithLogitsLoss(pos_weight=torch.as_tensor([2.318]).cuda()).cuda()\n        self.bce = nn.BCEWithLogitsLoss()\n        self.dice = smp.losses.DiceLoss(smp.losses.MULTILABEL_MODE, from_logits=True)\n        \n    def forward(self, outputs, targets, masks_outputs, masks_outputs2, masks_targets):\n        loss1 = self.bce(outputs, targets.float())\n        \n        masks_outputs = masks_outputs.float()\n        masks_outputs2 = masks_outputs2.float()\n        \n        masks_targets = masks_targets.float().flatten(0, 1)\n        \n        loss2 = self.dice(masks_outputs, masks_targets) + self.dice(masks_outputs2, masks_targets)\n        \n        \n        loss = loss1 + (loss2 * 0.125) \n        \n        return loss\n``` ",
          "votes": 3,
          "replies": [
            {
              "id": 2483753,
              "postDate": "2023-10-16T00:49:37.097Z",
              "content": "<p>Great, thank you!</p>",
              "rawMarkdown": "Great, thank you!",
              "votes": 1
            }
          ]
        },
        {
          "id": 2483748,
          "postDate": "2023-10-16T00:46:34.113Z",
          "content": "<p>Yes, we did use masks prediction by 3D segmentation model as the ground truth label, we used 2 approaches on how to predict them in the model, one is like you say, Unet decoder blocks, and the other is direct conv head from the encoder</p>",
          "rawMarkdown": "Yes, we did use masks prediction by 3D segmentation model as the ground truth label, we used 2 approaches on how to predict them in the model, one is like you say, Unet decoder blocks, and the other is direct conv head from the encoder",
          "votes": 2,
          "replies": [
            {
              "id": 2483754,
              "postDate": "2023-10-16T00:50:39.120Z",
              "content": "<p>That makes sense, thanks. Out of curiosity, did you find that using two decoders/mask heads improved performance over using just one?</p>",
              "rawMarkdown": "That makes sense, thanks. Out of curiosity, did you find that using two decoders/mask heads improved performance over using just one?",
              "votes": 1
            },
            {
              "id": 2483786,
              "postDate": "2023-10-16T01:51:14.517Z",
              "content": "<p>Not much difference in performance, but it brought some diversity between the models, improving the ensemble scores. </p>",
              "rawMarkdown": "Not much difference in performance, but it brought some diversity between the models, improving the ensemble scores. ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2484862,
      "postDate": "2023-10-16T18:13:00.713Z",
      "content": "<p>Nice work from all three of you, as usual. Thanks for sharing with us!</p>",
      "rawMarkdown": "Nice work from all three of you, as usual. Thanks for sharing with us!",
      "votes": 2
    },
    {
      "id": 2484510,
      "postDate": "2023-10-16T14:33:35.810Z",
      "content": "<p>congrats and thanks, looking forward to the training code</p>",
      "rawMarkdown": "congrats and thanks, looking forward to the training code",
      "votes": 2
    },
    {
      "id": 2484000,
      "postDate": "2023-10-16T06:57:13.857Z",
      "content": "<p><a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> congratulations,</p>\n<p>After many years of hard work, your routine is now .. eat sleep get 1st rank .. post it in  LinkedIn and repeat..</p>\n<p>Is there any secret sauce for noobies? 😁</p>",
      "rawMarkdown": "@nischaydnk congratulations,\n\nAfter many years of hard work, your routine is now .. eat sleep get 1st rank .. post it in  LinkedIn and repeat..\n\nIs there any secret sauce for noobies? 😁",
      "votes": 2,
      "replies": [
        {
          "id": 2491971,
          "postDate": "2023-10-22T07:08:37.343Z",
          "content": "<p>Haha Thank you. I guess you already mentioned secret sauce. \"hard work, routine\" :D </p>",
          "rawMarkdown": "Haha Thank you. I guess you already mentioned secret sauce. \"hard work, routine\" :D ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2483891,
      "postDate": "2023-10-16T04:41:48.047Z",
      "content": "<p>Hi! Congratulations on your 1st. Your solutions always amaze me. </p>\n<p>And I think this is a really amazing idea and the heart of the competition!</p>\n<p>\"For example if a patient has label 0 for liver-injury and the liver visibility is as follows in the slice sequence</p>\n<p>[0., 0., 0., 0.01, 0.05, 0.1, 0.23, 0.5, 0.7, 0.95, 0.99, 1., 0.95, 0.8, 0.4 …. 0. ,0., 0.]\"</p>\n<p>(However, for example, if the upper part of the organ is damaged and the lower part is normal, this may be an incorrect soft-label. What do you think?)</p>",
      "rawMarkdown": "Hi! Congratulations on your 1st. Your solutions always amaze me. \n\nAnd I think this is a really amazing idea and the heart of the competition!\n\n\"For example if a patient has label 0 for liver-injury and the liver visibility is as follows in the slice sequence\n\n[0., 0., 0., 0.01, 0.05, 0.1, 0.23, 0.5, 0.7, 0.95, 0.99, 1., 0.95, 0.8, 0.4 …. 0. ,0., 0.]\"\n\n(However, for example, if the upper part of the organ is damaged and the lower part is normal, this may be an incorrect soft-label. What do you think?)",
      "votes": 2
    },
    {
      "id": 2483728,
      "postDate": "2023-10-16T00:14:48.240Z",
      "content": "<p>Wow a really clean way of creating image level labels, congrats!</p>",
      "rawMarkdown": "Wow a really clean way of creating image level labels, congrats!",
      "votes": 2
    },
    {
      "id": 2483726,
      "postDate": "2023-10-16T00:12:45.107Z",
      "content": "<p>thx for sharing. To honest, Host should show 3 or 4 decimals, to see how close the competition is .</p>",
      "rawMarkdown": "thx for sharing. To honest, Host should show 3 or 4 decimals, to see how close the competition is .",
      "votes": 2,
      "replies": [
        {
          "id": 2483729,
          "postDate": "2023-10-16T00:14:56.073Z",
          "content": "<p>Thank you. We are curious about that too :))</p>",
          "rawMarkdown": "Thank you. We are curious about that too :))",
          "votes": 1
        }
      ]
    },
    {
      "id": 3298170,
      "postDate": "2025-10-04T18:12:16.883Z",
      "content": "<p>Congrats!!! thinks for sharing！</p>",
      "rawMarkdown": "Congrats!!! thinks for sharing！"
    },
    {
      "id": 3051257,
      "postDate": "2024-11-21T05:03:16.573Z",
      "content": "<p>I would like to be allowed to ask a question from the future, more than a year after the competition has ended.</p>\n<p>Where does the <code>f‘{PATHS.TOTAL_SEGMENTOR_FOLDER}/meta.csv’</code> in <code>make_segmentation_data1.py</code> in the repository come from?</p>",
      "rawMarkdown": "I would like to be allowed to ask a question from the future, more than a year after the competition has ended.\n\nWhere does the `f‘{PATHS.TOTAL_SEGMENTOR_FOLDER}/meta.csv’` in `make_segmentation_data1.py` in the repository come from?",
      "replies": [
        {
          "id": 3298121,
          "postDate": "2025-10-04T14:53:11.587Z",
          "content": "<p>The file exists in Totalsegmentator_dataset, which is used to train the segmentation model as described in the original dataset description.</p>\n<blockquote>\n  <p>segmentations/ Model generated pixel-level annotations of the relevant organs and some major bones for a subset of the scans in the training set. This data is provided in the nifti file format. The filenames are series IDs. You can find a description of the source model (total segmentator) here and the data used to train that model here.</p>\n</blockquote>",
          "rawMarkdown": "The file exists in Totalsegmentator_dataset, which is used to train the segmentation model as described in the original dataset description.\n\n> segmentations/ Model generated pixel-level annotations of the relevant organs and some major bones for a subset of the scans in the training set. This data is provided in the nifti file format. The filenames are series IDs. You can find a description of the source model (total segmentator) here and the data used to train that model here."
        }
      ]
    },
    {
      "id": 3016803,
      "postDate": "2024-10-14T07:15:05.883Z",
      "content": "<p>Congratulations and thank you for the amazing work.<br>\nBut I have some questions to ask.<br>\nWhen I did the preprocessing part, it showed that there is no file called, meta.csv. Did I miss anything, or is there any preparation for the datasets I should do first?</p>",
      "rawMarkdown": "Congratulations and thank you for the amazing work.\nBut I have some questions to ask.\nWhen I did the preprocessing part, it showed that there is no file called, meta.csv. Did I miss anything, or is there any preparation for the datasets I should do first?"
    },
    {
      "id": 2539245,
      "postDate": "2023-11-26T21:45:36.960Z",
      "content": "<p>You guys have used a contrails model. What is that and why are you using that? </p>",
      "rawMarkdown": "You guys have used a contrails model. What is that and why are you using that? ",
      "replies": [
        {
          "id": 2539383,
          "postDate": "2023-11-27T03:10:36.023Z",
          "content": "<p>We do not use contrails model, we used the code of the model that comes from another competition which is regarding contrails segmentation</p>",
          "rawMarkdown": "We do not use contrails model, we used the code of the model that comes from another competition which is regarding contrails segmentation"
        }
      ]
    },
    {
      "id": 2529491,
      "postDate": "2023-11-18T09:50:44.840Z",
      "content": "<p>Congratulations and thanks for sharing your excellent work. I ran into some problems as I was replicating your solution.</p>\n<p>The library <strong>src.coat</strong> and <strong>src.layers</strong> in <strong>Models/coatmed384fullseed_model.py</strong> are missing. </p>\n<p><code>from src.coat import CoaT,coat_lite_mini,coat_lite_small,coat_lite_medium</code><br>\n<code>from src.layers import *</code></p>\n<p>I guess this library is from repo <strong>mlpc-ucsd/CoaT</strong>. I have downloaded the repo and set PATHS.CONTRAIL_MODEL_BASE. But I checked its source code and it turned out the CoaT repo doesn't have anything as src.coat and src.layers. It does have src.model.coat, which may be a substitute to src.coat. But I can't find anything similar to src.layers. Is there something wrong with the code?</p>",
      "rawMarkdown": "Congratulations and thanks for sharing your excellent work. I ran into some problems as I was replicating your solution.\n\nThe library **src.coat** and **src.layers** in **Models/coatmed384fullseed_model.py** are missing. \n\n`from src.coat import CoaT,coat_lite_mini,coat_lite_small,coat_lite_medium`\n`from src.layers import *`\n\nI guess this library is from repo **mlpc-ucsd/CoaT**. I have downloaded the repo and set PATHS.CONTRAIL_MODEL_BASE. But I checked its source code and it turned out the CoaT repo doesn't have anything as src.coat and src.layers. It does have src.model.coat, which may be a substitute to src.coat. But I can't find anything similar to src.layers. Is there something wrong with the code?",
      "replies": [
        {
          "id": 2529744,
          "postDate": "2023-11-18T14:01:01.833Z",
          "content": "<p>We are sorry, we somehow forgot to mention this, <a href=\"https://www.kaggle.com/datasets/iafoss/contrails-model-def1\" target=\"_blank\">https://www.kaggle.com/datasets/iafoss/contrails-model-def1</a>, you might need to rename the folder \"src_inference1\" to \"src\"</p>",
          "rawMarkdown": "We are sorry, we somehow forgot to mention this, https://www.kaggle.com/datasets/iafoss/contrails-model-def1, you might need to rename the folder \"src_inference1\" to \"src\"",
          "replies": [
            {
              "id": 2529807,
              "postDate": "2023-11-18T15:02:47.333Z",
              "content": "<p>Thank you so much.</p>",
              "rawMarkdown": "Thank you so much."
            }
          ]
        }
      ]
    },
    {
      "id": 2504104,
      "postDate": "2023-10-29T17:04:47.603Z",
      "content": "<p>Hey, congratulations on the win!<br>\nI had a few question:</p>\n<ol>\n<li>You mentioned using <a href=\"https://www.kaggle.com/Theviel\" target=\"_blank\">@Theviel</a>'s dataset processing code. I saw on their profile and the one that I found converts dcm to png. Did you guys also do it?</li>\n<li>&gt;each volume extracted with equi-distant 96 slices.<br>\nWhat does this mean? Let's say a study has 1000 slices, you'd pick a slice after every 96 slices? Also why did you do this?</li>\n</ol>\n<p>Sorry for the basic questions, I am a noob trying to replicate your solution for understanding purposes.</p>",
      "rawMarkdown": "Hey, congratulations on the win!\nI had a few question:\n1.  You mentioned using @Theviel's dataset processing code. I saw on their profile and the one that I found converts dcm to png. Did you guys also do it?\n2. >each volume extracted with equi-distant 96 slices.\nWhat does this mean? Let's say a study has 1000 slices, you'd pick a slice after every 96 slices? Also why did you do this?\n\nSorry for the basic questions, I am a noob trying to replicate your solution for understanding purposes.\n",
      "replies": [
        {
          "id": 2504341,
          "postDate": "2023-10-29T19:24:58.480Z",
          "content": "<p>Answer 1</p>\n<p>Yes, Early in the competition we created our own dataset, we used it for a long time, near the end, we tried theo's public dataset, which results in similar scores to our own but the ensemble boost was relevant</p>\n<p>Answer 2</p>\n<p>For example, If a study has 1000 slices, then we make volumes by selecting (0-96), (96-192), (192-288) and so on… and then each volume will be given to the model to make predictions on, generally taking more slices than 96 might result in better performance but it also uses a lot of GPU memory</p>",
          "rawMarkdown": "Answer 1\n\nYes, Early in the competition we created our own dataset, we used it for a long time, near the end, we tried theo's public dataset, which results in similar scores to our own but the ensemble boost was relevant\n\nAnswer 2\n\nFor example, If a study has 1000 slices, then we make volumes by selecting (0-96), (96-192), (192-288) and so on... and then each volume will be given to the model to make predictions on, generally taking more slices than 96 might result in better performance but it also uses a lot of GPU memory",
          "replies": [
            {
              "id": 2504483,
              "postDate": "2023-10-29T22:06:28.807Z",
              "content": "<p>Sorry for asking the same kind of question again, but just to clarify, PNGs produced a better result than dicoms?</p>\n<p>Also, did you use the nii files provided by the authors to do something?</p>",
              "rawMarkdown": "Sorry for asking the same kind of question again, but just to clarify, PNGs produced a better result than dicoms?\n\nAlso, did you use the nii files provided by the authors to do something?\n"
            },
            {
              "id": 2504728,
              "postDate": "2023-10-30T05:56:34.660Z",
              "content": "<p>No, it was not the PNGs that produced the results, what produces results is how dicom, which is the raw data, is handled and preprocessed</p>\n<p>Yes, we used .nii files for 3D segmentation model</p>",
              "rawMarkdown": "No, it was not the PNGs that produced the results, what produces results is how dicom, which is the raw data, is handled and preprocessed\n\nYes, we used .nii files for 3D segmentation model"
            },
            {
              "id": 2596802,
              "postDate": "2024-01-11T10:41:30.167Z",
              "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> <br>\nThanks for the good explanation, but I still don't understand it very well, so I have a question. If a patient has 1000 slices, only some parts will have extravasation, but if you cut it with 96 equidistant slices, there will be at least 10 volumes and have more parts without extravasation, so I wonder if you labeled all volumes (10 volumes in this case) as extravasation and trained the model.</p>",
              "rawMarkdown": "@harshitsheoran \nThanks for the good explanation, but I still don't understand it very well, so I have a question. If a patient has 1000 slices, only some parts will have extravasation, but if you cut it with 96 equidistant slices, there will be at least 10 volumes and have more parts without extravasation, so I wonder if you labeled all volumes (10 volumes in this case) as extravasation and trained the model."
            },
            {
              "id": 2597140,
              "postDate": "2024-01-11T14:39:08.697Z",
              "content": "<p>Image-level labels for extravasation and bowel are provided in the training data.</p>",
              "rawMarkdown": "Image-level labels for extravasation and bowel are provided in the training data."
            }
          ]
        }
      ]
    },
    {
      "id": 2502381,
      "postDate": "2023-10-28T06:57:04.443Z",
      "content": "<p>Congratulations! You have done an excellent job.<br>\nHowever, I still have 2 questions.</p>\n<p>Question 1:<br>\nWhen you train the extravasation model, how do you obtain the segmentation head label? I'm asking because the dataset only provides image-level information for extravasation.<br>\nIs it set like this: If the image indicates a true extravasation injury, is the segmentation training label set to all 1s? If not, is it set to all 0s?</p>\n<p>Question 2:<br>\nIn your GitHub code, the criterion uses <code>nn.BCEWithLogitsLoss()</code> directly without any weights. However, during inference, you added the weight as a post-process. Have you tried using a weighted <code>BCEWithLogitsLoss</code>?</p>",
      "rawMarkdown": "Congratulations! You have done an excellent job.\nHowever, I still have 2 questions.\n\nQuestion 1:\nWhen you train the extravasation model, how do you obtain the segmentation head label? I'm asking because the dataset only provides image-level information for extravasation.\nIs it set like this: If the image indicates a true extravasation injury, is the segmentation training label set to all 1s? If not, is it set to all 0s?\n\nQuestion 2:\nIn your GitHub code, the criterion uses `nn.BCEWithLogitsLoss()` directly without any weights. However, during inference, you added the weight as a post-process. Have you tried using a weighted `BCEWithLogitsLoss`?\n",
      "replies": [
        {
          "id": 2502638,
          "postDate": "2023-10-28T10:46:59.800Z",
          "content": "<p>Answer 1<br>\nSegmentation head segments liver-kidney-spleen, and not extravasation, it is like an auxiliary guidance which does not directly helps extravasation</p>\n<p>Answer 2<br>\nYes, we did try weighted BCE loss, in our experiments, it did not help model converge, one reason for this could be our sampling, in our GitHub code, you can see that we are sampling injured and non-injured patients equally</p>",
          "rawMarkdown": "Answer 1\nSegmentation head segments liver-kidney-spleen, and not extravasation, it is like an auxiliary guidance which does not directly helps extravasation\n\nAnswer 2\nYes, we did try weighted BCE loss, in our experiments, it did not help model converge, one reason for this could be our sampling, in our GitHub code, you can see that we are sampling injured and non-injured patients equally",
          "replies": [
            {
              "id": 2504779,
              "postDate": "2023-10-30T06:34:37.153Z",
              "content": "<p>Thank you very much！</p>",
              "rawMarkdown": "Thank you very much！"
            }
          ]
        }
      ]
    },
    {
      "id": 3012809,
      "postDate": "2024-10-09T11:47:00.020Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2529486,
      "postDate": "2023-11-18T09:47:22.960Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2516769,
      "postDate": "2023-11-07T23:45:53.390Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2516761,
      "postDate": "2023-11-07T23:35:35.120Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2485824,
      "postDate": "2023-10-17T13:32:45.713Z",
      "content": "<p>Congratulations and thanks for sharing! For beginners like me, could you explain why using 32 equidistant slices and 3 adjacent slices is better than stacking 96 slices in the same channel? You also employed this technique in the cervical challenge.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing! For beginners like me, could you explain why using 32 equidistant slices and 3 adjacent slices is better than stacking 96 slices in the same channel? You also employed this technique in the cervical challenge.",
      "isDeleted": true,
      "replies": [
        {
          "id": 2491969,
          "postDate": "2023-10-22T07:07:30.660Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/nathanbanaste\" target=\"_blank\">@nathanbanaste</a> It was more up to the experimentation. One major downside of using a single channel and 96 slices would be a large increase in computation time as we are passing 3x num images to the encoder. Performance-wise, we tried channels 1, 3, 5, 6, out of these 3 &amp; 5 worked the best for us. </p>",
          "rawMarkdown": "Thanks @nathanbanaste It was more up to the experimentation. One major downside of using a single channel and 96 slices would be a large increase in computation time as we are passing 3x num images to the encoder. Performance-wise, we tried channels 1, 3, 5, 6, out of these 3 & 5 worked the best for us. \n ",
          "replies": [
            {
              "id": 2492184,
              "postDate": "2023-10-22T10:12:28.460Z",
              "content": "<p>What i meant was that you used 32,3,image_size,image_size for input shape.<br>\nSo you say that this input shape  32,3,img_size,img_size is less computation than 96,img_size,img_size (with same batch_size)?</p>",
              "rawMarkdown": "What i meant was that you used 32,3,image_size,image_size for input shape.\nSo you say that this input shape  32,3,img_size,img_size is less computation than 96,img_size,img_size (with same batch_size)?",
              "isDeleted": true
            },
            {
              "id": 2492265,
              "postDate": "2023-10-22T10:52:07.070Z",
              "content": "<p>Yes, (2, 32, 3, 384, 384) is much less compute than (2, 96, 1, 384, 384), shape being (batch_size, n_slices, n_channels, img_size, img_size)</p>",
              "rawMarkdown": "Yes, (2, 32, 3, 384, 384) is much less compute than (2, 96, 1, 384, 384), shape being (batch_size, n_slices, n_channels, img_size, img_size)"
            }
          ]
        }
      ]
    },
    {
      "id": 2485236,
      "postDate": "2023-10-17T03:41:07.937Z",
      "content": "<p>Thanks for providing us your solution</p>",
      "rawMarkdown": "Thanks for providing us your solution",
      "votes": 3
    }
  ],
  "comments": [
    {
      "id": 2484321,
      "author_name": "Kiurtis",
      "author_url": "",
      "post_date": "2023-10-16T11:57:05.177000",
      "content": "<p>Congratulation and thanks for sharing those insights ! Waiting for the full solution. </p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2484264,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-10-16T10:56:02.173000",
      "content": "<p>Congratulations on winning the top position. Thanks for sharing the solution. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2483894,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2023-10-16T04:43:40.883000",
      "content": "<p>😀😀😀<br>\nCongrats and Big Thanks to my wonderful teammates!! <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2513825,
      "author_name": "EE22M215 VIJAY RAJ",
      "author_url": "",
      "post_date": "2023-11-05T18:52:36.063000",
      "content": "<p>Congratulations on securing the top position .Looking forward to see full code</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2486869,
      "author_name": "Grey Benhin ",
      "author_url": "",
      "post_date": "2023-10-18T07:56:00.003000",
      "content": "<p>Tq for your work </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2486059,
      "author_name": "Reynita_777",
      "author_url": "",
      "post_date": "2023-10-17T16:49:21.100000",
      "content": "<p>Congratulations, team! Nice job :) Looking forward to seeing the final notebook. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2485202,
      "author_name": "koihoo Wang",
      "author_url": "",
      "post_date": "2023-10-17T02:52:12.537000",
      "content": "<p>Congrats!!! thinks for sharing！</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2485163,
      "author_name": "wenzhia",
      "author_url": "",
      "post_date": "2023-10-17T01:55:20.390000",
      "content": "<p>Congratulations！I would appreciate it if you share the code。</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2483735,
      "author_name": "Romain Hardy",
      "author_url": "",
      "post_date": "2023-10-16T00:21:04.300000",
      "content": "<p>Congratulations on the win! Your solution is very elegant :)</p>\n<p>For the auxiliary segmentation objective, did you use the masks generated by your 3D segmentation models as the ground truth labels? If so, does that mean you also added a decoder module to your classification models to predict segmentation masks?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2483747,
          "author_name": "Nischay Dhankhar",
          "author_url": "",
          "post_date": "2023-10-16T00:46:12.980000",
          "content": "<p>Thank you, yes we used 3D segmentation models output as ground truth, while training classification models, along with linear classifier head, we used 1 / 2 separate mask heads/decoders. So, our loss function looks something like this.</p>\n<pre><code></code></pre>",
          "votes": 3,
          "replies": [
            {
              "id": 2483753,
              "author_name": "Romain Hardy",
              "author_url": "",
              "post_date": "2023-10-16T00:49:37.097000",
              "content": "<p>Great, thank you!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2483748,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2023-10-16T00:46:34.113000",
          "content": "<p>Yes, we did use masks prediction by 3D segmentation model as the ground truth label, we used 2 approaches on how to predict them in the model, one is like you say, Unet decoder blocks, and the other is direct conv head from the encoder</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2483754,
              "author_name": "Romain Hardy",
              "author_url": "",
              "post_date": "2023-10-16T00:50:39.120000",
              "content": "<p>That makes sense, thanks. Out of curiosity, did you find that using two decoders/mask heads improved performance over using just one?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2483786,
              "author_name": "Nischay Dhankhar",
              "author_url": "",
              "post_date": "2023-10-16T01:51:14.517000",
              "content": "<p>Not much difference in performance, but it brought some diversity between the models, improving the ensemble scores. </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2484862,
      "author_name": "David Roberts",
      "author_url": "",
      "post_date": "2023-10-16T18:13:00.713000",
      "content": "<p>Nice work from all three of you, as usual. Thanks for sharing with us!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2484510,
      "author_name": "ls",
      "author_url": "",
      "post_date": "2023-10-16T14:33:35.810000",
      "content": "<p>congrats and thanks, looking forward to the training code</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2484000,
      "author_name": "MLV Prasad",
      "author_url": "",
      "post_date": "2023-10-16T06:57:13.857000",
      "content": "<p><a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> congratulations,</p>\n<p>After many years of hard work, your routine is now .. eat sleep get 1st rank .. post it in  LinkedIn and repeat..</p>\n<p>Is there any secret sauce for noobies? 😁</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2491971,
          "author_name": "Nischay Dhankhar",
          "author_url": "",
          "post_date": "2023-10-22T07:08:37.343000",
          "content": "<p>Haha Thank you. I guess you already mentioned secret sauce. \"hard work, routine\" :D </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2483891,
      "author_name": "devchopin",
      "author_url": "",
      "post_date": "2023-10-16T04:41:48.047000",
      "content": "<p>Hi! Congratulations on your 1st. Your solutions always amaze me. </p>\n<p>And I think this is a really amazing idea and the heart of the competition!</p>\n<p>\"For example if a patient has label 0 for liver-injury and the liver visibility is as follows in the slice sequence</p>\n<p>[0., 0., 0., 0.01, 0.05, 0.1, 0.23, 0.5, 0.7, 0.95, 0.99, 1., 0.95, 0.8, 0.4 …. 0. ,0., 0.]\"</p>\n<p>(However, for example, if the upper part of the organ is damaged and the lower part is normal, this may be an incorrect soft-label. What do you think?)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2483728,
      "author_name": "Feng Qilong",
      "author_url": "",
      "post_date": "2023-10-16T00:14:48.240000",
      "content": "<p>Wow a really clean way of creating image level labels, congrats!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2483726,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2023-10-16T00:12:45.107000",
      "content": "<p>thx for sharing. To honest, Host should show 3 or 4 decimals, to see how close the competition is .</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2483729,
          "author_name": "Nischay Dhankhar",
          "author_url": "",
          "post_date": "2023-10-16T00:14:56.073000",
          "content": "<p>Thank you. We are curious about that too :))</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3298170,
      "author_name": "Khushi Yadav",
      "author_url": "",
      "post_date": "2025-10-04T18:12:16.883000",
      "content": "<p>Congrats!!! thinks for sharing！</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3051257,
      "author_name": "Ataracsia",
      "author_url": "",
      "post_date": "2024-11-21T05:03:16.573000",
      "content": "<p>I would like to be allowed to ask a question from the future, more than a year after the competition has ended.</p>\n<p>Where does the <code>f‘{PATHS.TOTAL_SEGMENTOR_FOLDER}/meta.csv’</code> in <code>make_segmentation_data1.py</code> in the repository come from?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3298121,
          "author_name": "Lunran",
          "author_url": "",
          "post_date": "2025-10-04T14:53:11.587000",
          "content": "<p>The file exists in Totalsegmentator_dataset, which is used to train the segmentation model as described in the original dataset description.</p>\n<blockquote>\n  <p>segmentations/ Model generated pixel-level annotations of the relevant organs and some major bones for a subset of the scans in the training set. This data is provided in the nifti file format. The filenames are series IDs. You can find a description of the source model (total segmentator) here and the data used to train that model here.</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3016803,
      "author_name": "Po-Cheng Cheng",
      "author_url": "",
      "post_date": "2024-10-14T07:15:05.883000",
      "content": "<p>Congratulations and thank you for the amazing work.<br>\nBut I have some questions to ask.<br>\nWhen I did the preprocessing part, it showed that there is no file called, meta.csv. Did I miss anything, or is there any preparation for the datasets I should do first?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2539245,
      "author_name": "Syed Muhammad Sameer",
      "author_url": "",
      "post_date": "2023-11-26T21:45:36.960000",
      "content": "<p>You guys have used a contrails model. What is that and why are you using that? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2539383,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2023-11-27T03:10:36.023000",
          "content": "<p>We do not use contrails model, we used the code of the model that comes from another competition which is regarding contrails segmentation</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2529491,
      "author_name": "Richard Luo",
      "author_url": "",
      "post_date": "2023-11-18T09:50:44.840000",
      "content": "<p>Congratulations and thanks for sharing your excellent work. I ran into some problems as I was replicating your solution.</p>\n<p>The library <strong>src.coat</strong> and <strong>src.layers</strong> in <strong>Models/coatmed384fullseed_model.py</strong> are missing. </p>\n<p><code>from src.coat import CoaT,coat_lite_mini,coat_lite_small,coat_lite_medium</code><br>\n<code>from src.layers import *</code></p>\n<p>I guess this library is from repo <strong>mlpc-ucsd/CoaT</strong>. I have downloaded the repo and set PATHS.CONTRAIL_MODEL_BASE. But I checked its source code and it turned out the CoaT repo doesn't have anything as src.coat and src.layers. It does have src.model.coat, which may be a substitute to src.coat. But I can't find anything similar to src.layers. Is there something wrong with the code?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2529744,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2023-11-18T14:01:01.833000",
          "content": "<p>We are sorry, we somehow forgot to mention this, <a href=\"https://www.kaggle.com/datasets/iafoss/contrails-model-def1\" target=\"_blank\">https://www.kaggle.com/datasets/iafoss/contrails-model-def1</a>, you might need to rename the folder \"src_inference1\" to \"src\"</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2529807,
              "author_name": "Richard Luo",
              "author_url": "",
              "post_date": "2023-11-18T15:02:47.333000",
              "content": "<p>Thank you so much.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2504104,
      "author_name": "Syed Muhammad Sameer",
      "author_url": "",
      "post_date": "2023-10-29T17:04:47.603000",
      "content": "<p>Hey, congratulations on the win!<br>\nI had a few question:</p>\n<ol>\n<li>You mentioned using <a href=\"https://www.kaggle.com/Theviel\" target=\"_blank\">@Theviel</a>'s dataset processing code. I saw on their profile and the one that I found converts dcm to png. Did you guys also do it?</li>\n<li>&gt;each volume extracted with equi-distant 96 slices.<br>\nWhat does this mean? Let's say a study has 1000 slices, you'd pick a slice after every 96 slices? Also why did you do this?</li>\n</ol>\n<p>Sorry for the basic questions, I am a noob trying to replicate your solution for understanding purposes.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2504341,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2023-10-29T19:24:58.480000",
          "content": "<p>Answer 1</p>\n<p>Yes, Early in the competition we created our own dataset, we used it for a long time, near the end, we tried theo's public dataset, which results in similar scores to our own but the ensemble boost was relevant</p>\n<p>Answer 2</p>\n<p>For example, If a study has 1000 slices, then we make volumes by selecting (0-96), (96-192), (192-288) and so on… and then each volume will be given to the model to make predictions on, generally taking more slices than 96 might result in better performance but it also uses a lot of GPU memory</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2504483,
              "author_name": "Syed Muhammad Sameer",
              "author_url": "",
              "post_date": "2023-10-29T22:06:28.807000",
              "content": "<p>Sorry for asking the same kind of question again, but just to clarify, PNGs produced a better result than dicoms?</p>\n<p>Also, did you use the nii files provided by the authors to do something?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2504728,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2023-10-30T05:56:34.660000",
              "content": "<p>No, it was not the PNGs that produced the results, what produces results is how dicom, which is the raw data, is handled and preprocessed</p>\n<p>Yes, we used .nii files for 3D segmentation model</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2596802,
              "author_name": "kuchoco97",
              "author_url": "",
              "post_date": "2024-01-11T10:41:30.167000",
              "content": "<p><a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> <br>\nThanks for the good explanation, but I still don't understand it very well, so I have a question. If a patient has 1000 slices, only some parts will have extravasation, but if you cut it with 96 equidistant slices, there will be at least 10 volumes and have more parts without extravasation, so I wonder if you labeled all volumes (10 volumes in this case) as extravasation and trained the model.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2597140,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2024-01-11T14:39:08.697000",
              "content": "<p>Image-level labels for extravasation and bowel are provided in the training data.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2502381,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2023-10-28T06:57:04.443000",
      "content": "<p>Congratulations! You have done an excellent job.<br>\nHowever, I still have 2 questions.</p>\n<p>Question 1:<br>\nWhen you train the extravasation model, how do you obtain the segmentation head label? I'm asking because the dataset only provides image-level information for extravasation.<br>\nIs it set like this: If the image indicates a true extravasation injury, is the segmentation training label set to all 1s? If not, is it set to all 0s?</p>\n<p>Question 2:<br>\nIn your GitHub code, the criterion uses <code>nn.BCEWithLogitsLoss()</code> directly without any weights. However, during inference, you added the weight as a post-process. Have you tried using a weighted <code>BCEWithLogitsLoss</code>?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2502638,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2023-10-28T10:46:59.800000",
          "content": "<p>Answer 1<br>\nSegmentation head segments liver-kidney-spleen, and not extravasation, it is like an auxiliary guidance which does not directly helps extravasation</p>\n<p>Answer 2<br>\nYes, we did try weighted BCE loss, in our experiments, it did not help model converge, one reason for this could be our sampling, in our GitHub code, you can see that we are sampling injured and non-injured patients equally</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2504779,
              "author_name": "Alex",
              "author_url": "",
              "post_date": "2023-10-30T06:34:37.153000",
              "content": "<p>Thank you very much！</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3012809,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-09T11:47:00.020000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2529486,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-11-18T09:47:22.960000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2516769,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-11-07T23:45:53.390000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2516761,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-11-07T23:35:35.120000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2485824,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-17T13:32:45.713000",
      "content": "<p>Congratulations and thanks for sharing! For beginners like me, could you explain why using 32 equidistant slices and 3 adjacent slices is better than stacking 96 slices in the same channel? You also employed this technique in the cervical challenge.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2491969,
          "author_name": "Nischay Dhankhar",
          "author_url": "",
          "post_date": "2023-10-22T07:07:30.660000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/nathanbanaste\" target=\"_blank\">@nathanbanaste</a> It was more up to the experimentation. One major downside of using a single channel and 96 slices would be a large increase in computation time as we are passing 3x num images to the encoder. Performance-wise, we tried channels 1, 3, 5, 6, out of these 3 &amp; 5 worked the best for us. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2492184,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-10-22T10:12:28.460000",
              "content": "<p>What i meant was that you used 32,3,image_size,image_size for input shape.<br>\nSo you say that this input shape  32,3,img_size,img_size is less computation than 96,img_size,img_size (with same batch_size)?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2492265,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2023-10-22T10:52:07.070000",
              "content": "<p>Yes, (2, 32, 3, 384, 384) is much less compute than (2, 96, 1, 384, 384), shape being (batch_size, n_slices, n_channels, img_size, img_size)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2485236,
      "author_name": "Indra Sonowal",
      "author_url": "",
      "post_date": "2023-10-17T03:41:07.937000",
      "content": "<p>Thanks for providing us your solution</p>",
      "votes": 3,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2483715": "Firstly, Thank you RSNA for hosting another interesting competition & my teammates @haqishen @harshitsheoran - formation of **Team oxygen** ? :)) It was amazing to be #1 on public leaderboard for almost a month. I am sharing a quick overview of our solution, we will release the entire solution soon. It was really fun competing for #1 with @theoviel  \n**Edit: Full solution published.**  \nHere is the inference code you may refer: [link](https://www.kaggle.com/nischaydnk/rsna-super-mega-lb-ensemble) \nOur GitHub repo w/ all preprocessing + training code: [link](https://github.com/Nischaydnk/RSNA-2023-1st-place-solution)\nDemo Inference notebook: [link](https://www.kaggle.com/code/haqishen/rsna-2023-1st-place-best-model-infer-cleaned)\n\n#### **Split used:** 4 Fold GroupKFold ( Patient Level)\n\n## **Our solution is divided into three parts:**\n**Part 1:** 3D segmentation for generating masks / crops [Stage 1]\n**Part 2:** 2D CNN + RNN based approach for Kidney, Liver, Spleen & Bowel [Stage 2]\n**Part 3:** 2D CNN + RNN based approach for Bowel + Extravasation [Stage 2]\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fee28ed8eef7827d8f2cc69601875e5c2%2FScreenshot%202023-10-22%20at%2011.54.15%20AM.png?generation=1697955883461262&alt=media)\n\n## **Data Preprocessing:**\n\nHere comes the key part of our solution, we will describe it later in more depth. **Note:** *All models were trained on image size 384 x 384. We use datasets preprocessing from @TheoVeol and our data which we made rescale dicoms and applying soft-tissue windowing.*\n\nWe take a patient/study, we run a 3d segmentation model on it, it outputs masks for each slice, we make a study-level crop here based on boundaries of organs - liver, spleen, kidney & liver. \n\nNext, we make volumes from the patient, each volume extracted with equi-distant 96 slices for a study which is then reshaped to (32, 3, image_size, image_size) in a 2.5D manner for training CNN based models.\n\n3 channels are formed by using the adjacent slices.\n\nAll our model takes in input in shape (2, 32, 3, height, width) and outputs it as (2, 32, n_classes) as the targets are also kept in shape (2, 32, n_classes).\n\nTo make the targets, we need 2 things, patient-level target of each organ and how much the organ is visible compared to its maximum visibility, this data is available after normalizing segmentation model masks in 0-1 based on number of positive pixels\n\nThen we multiply targets * patient-level target for each middle slice of the sequence and that is our label\n\nFor example if a patient has label 0 for liver-injury and the liver visibility is as follows in the slice sequence\n\n[0., 0., 0., 0.01, 0.05, 0.1, 0.23, 0.5, 0.7, 0.95, 0.99, 1., 0.95, 0.8, 0.4 …. 0. ,0., 0.]\n\nWe multiply it with label which is currently 0 results in an all zeros list as output, but if target label for liver-injury was 1, then we use the list mentioned above as our soft labels.\n\n\n## **Stage2: 2.5D Approach ( 2D CNN + RNN):**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fe8df4581839b1fa7dcadf68fe2a715a1%2FScreenshot%202023-10-22%20at%205.31.23%20AM.png?generation=1697935695484067&alt=media)\n\nIn stage 2, we trained our models using the volumes either based on our windowing or theo's preprocessing approach and the masks/crops generated from 3D segmentation approach. Each model is trained for multiple tasks (segmentation + classification). For all 32 sequences, we predicted slice level masks and sigmoid predictions. Further, simple maximum aggregation is applied on sigmoid predictions to fetch study level prediction used in submissions. \n\nFor training our models, some common settings were:\n- **Learning rate:** (1e-4 to 4e-4) range\n- **Optimizer:** AdamW\n- **Scheduler:** Cosine Annealing w/ Warmup \n- **Loss:** BCE Loss for Classification, Dice Loss for segmentation\n\n\n\n###**Auxiliary Segmentation Loss:** \nOne of the key things which made our training much more stable and helped in improving scores was using auxiliary losses based on segmentation. \n\nEncoder was kept same for both classification & segmentation decoders,  we used two types of segmentation head:\n- ***Unet based decoder*** for generating masks\n- ***2D-CNN*** based head \n```\nnn.Sequential(\n            nn.Conv2d(nb_ft, 128, kernel_size=3, padding=1),\n            nn.BatchNorm2d(128),\n            nn.ReLU(inplace=True),\n            nn.Conv2d(128, 128, kernel_size=3, padding=1),\n            nn.BatchNorm2d(128),\n            nn.ReLU(inplace=True),\n            nn.Conv2d(128, 4, kernel_size=1, padding=0),\n        )\n```\n```\n        self.mask_head_3 = self.get_mask_head(true_encoder.feature_info[-2]['num_chs'])\n        self.mask_head_4 = self.get_mask_head(true_encoder.feature_info[-1]['num_chs'])\n```\nWe used the feature maps generated mainly from last and 2nd last blocks of the backbones & apply dice loss on the predicted masks & true masks. This trick gave us around +0.01 to +0.03 boost in our models. We used similar technique in Covid 19 detection competition held few years back, you can also refer my solution for more detailed use of auxiliary loss & code snippets. \n[link of discussion](https://www.kaggle.com/c/siim-covid19-detection/discussion/266571)\n\nHere is an example code for applying aux loss:\n```\nclass CustomLoss(nn.Module):\n    def __init__(self):\n        super(CustomLoss, self).__init__()\n        #self.bce = nn.BCEWithLogitsLoss(pos_weight=torch.as_tensor([2.318]).cuda()).cuda()\n        self.bce = nn.BCEWithLogitsLoss()\n        self.dice = smp.losses.DiceLoss(smp.losses.MULTILABEL_MODE, from_logits=True)\n\n    def forward(self, outputs, targets, masks_outputs, masks_outputs2, masks_targets):\n        loss1 = self.bce(outputs, targets.float())\n\n        masks_outputs = masks_outputs.float()\n        masks_outputs2 = masks_outputs2.float()\n\n        masks_targets = masks_targets.float().flatten(0, 1)\n\n        loss2 = self.dice(masks_outputs, masks_targets) + self.dice(masks_outputs2, masks_targets)\n\n\n        loss = loss1 + (loss2 * 0.125) \n\n        return loss\n```\n\n\n### **Architectures used in Final ensemble:**\n- Coat Lite Medium w/ GRU - [original source code](https://github.com/mlpc-ucsd/CoaT)\n- Coat Lite Small w/ GRU!\n- Efficientnet v2s w/ GRU [Timm]\n\n\n### **Augmentations:**\n\nWe couldn't come up with several augmentations to use, but these were the ones which we used in our training.\n\n```\n        A.Perspective(p=0.5),\n        A.HorizontalFlip(p=0.5),\n        A.VerticalFlip(p=0.5),\n        A.Rotate(p=0.5, limit=(-25, 25)),\n```\n\n## **Post Processing / Ensemble:**\n\nFinal ensemble for all organs model includes **multiple Coat medium and V2s based models** trained on either 4 Folds or Full data. \n\nFor extravasation, We mainly used Coat Small and v2s in ensemble. \n**No major postprocessing** was applied except for tuning scaling factors based on CV scores.\nTo get the predictions, we aggregated the model outputs at slice level and simply took the maximum value for each patient.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fd6fa2cc524588b85b82906cccb6552bf%2FScreenshot%202023-10-22%20at%206.02.32%20AM.png?generation=1697936329146043&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2Fb69766140607127e717966085c927bce%2FScreenshot%202023-10-22%20at%206.07.13%20AM.png?generation=1697936371125187&alt=media)\n\n#### **Ensemble:**\n\nWithin folds of each models, we are doing slice level ensemble.\nFor different architectures & cross data models (theo/ours), we did ensemble after the max aggregation. \n\n#### **Best Ensemble OOF CV**: 0.31x\n#### **Best single model 4 fold OOF CV**: 0.326 [Coat lite Medium]\n\nOrgan level OOF for single model looks like this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4712534%2F9143afd07ba3069f2b2259b1d8fe80eb%2FScreenshot%202023-10-16%20at%204.05.07%20AM.png?generation=1697413214073394&alt=media)\n\nThank you. \n\n\nEDIT 1: 3D segmentation code: [notebook link](https://www.kaggle.com/code/haqishen/rsna-2023-1st-place-solution-train-3d-seg/notebook)\n",
    "2484321": "Congratulation and thanks for sharing those insights ! Waiting for the full solution. ",
    "2484264": "Congratulations on winning the top position. Thanks for sharing the solution. ",
    "2483894": "😀😀😀\nCongrats and Big Thanks to my wonderful teammates!! @nischaydnk @harshitsheoran ",
    "2513825": "Congratulations on securing the top position .Looking forward to see full code\n\n\n",
    "2486869": "Tq for your work ",
    "2486059": "Congratulations, team! Nice job :) Looking forward to seeing the final notebook. ",
    "2485202": "Congrats!!! thinks for sharing！",
    "2485163": "Congratulations！I would appreciate it if you share the code。",
    "2483735": "Congratulations on the win! Your solution is very elegant :)\n\nFor the auxiliary segmentation objective, did you use the masks generated by your 3D segmentation models as the ground truth labels? If so, does that mean you also added a decoder module to your classification models to predict segmentation masks?",
    "2484862": "Nice work from all three of you, as usual. Thanks for sharing with us!",
    "2484510": "congrats and thanks, looking forward to the training code",
    "2484000": "@nischaydnk congratulations,\n\nAfter many years of hard work, your routine is now .. eat sleep get 1st rank .. post it in  LinkedIn and repeat..\n\nIs there any secret sauce for noobies? 😁",
    "2483891": "Hi! Congratulations on your 1st. Your solutions always amaze me. \n\nAnd I think this is a really amazing idea and the heart of the competition!\n\n\"For example if a patient has label 0 for liver-injury and the liver visibility is as follows in the slice sequence\n\n[0., 0., 0., 0.01, 0.05, 0.1, 0.23, 0.5, 0.7, 0.95, 0.99, 1., 0.95, 0.8, 0.4 …. 0. ,0., 0.]\"\n\n(However, for example, if the upper part of the organ is damaged and the lower part is normal, this may be an incorrect soft-label. What do you think?)",
    "2483728": "Wow a really clean way of creating image level labels, congrats!",
    "2483726": "thx for sharing. To honest, Host should show 3 or 4 decimals, to see how close the competition is .",
    "3298170": "Congrats!!! thinks for sharing！",
    "3051257": "I would like to be allowed to ask a question from the future, more than a year after the competition has ended.\n\nWhere does the `f‘{PATHS.TOTAL_SEGMENTOR_FOLDER}/meta.csv’` in `make_segmentation_data1.py` in the repository come from?",
    "3016803": "Congratulations and thank you for the amazing work.\nBut I have some questions to ask.\nWhen I did the preprocessing part, it showed that there is no file called, meta.csv. Did I miss anything, or is there any preparation for the datasets I should do first?",
    "2539245": "You guys have used a contrails model. What is that and why are you using that? ",
    "2529491": "Congratulations and thanks for sharing your excellent work. I ran into some problems as I was replicating your solution.\n\nThe library **src.coat** and **src.layers** in **Models/coatmed384fullseed_model.py** are missing. \n\n`from src.coat import CoaT,coat_lite_mini,coat_lite_small,coat_lite_medium`\n`from src.layers import *`\n\nI guess this library is from repo **mlpc-ucsd/CoaT**. I have downloaded the repo and set PATHS.CONTRAIL_MODEL_BASE. But I checked its source code and it turned out the CoaT repo doesn't have anything as src.coat and src.layers. It does have src.model.coat, which may be a substitute to src.coat. But I can't find anything similar to src.layers. Is there something wrong with the code?",
    "2504104": "Hey, congratulations on the win!\nI had a few question:\n1.  You mentioned using @Theviel's dataset processing code. I saw on their profile and the one that I found converts dcm to png. Did you guys also do it?\n2. >each volume extracted with equi-distant 96 slices.\nWhat does this mean? Let's say a study has 1000 slices, you'd pick a slice after every 96 slices? Also why did you do this?\n\nSorry for the basic questions, I am a noob trying to replicate your solution for understanding purposes.\n",
    "2502381": "Congratulations! You have done an excellent job.\nHowever, I still have 2 questions.\n\nQuestion 1:\nWhen you train the extravasation model, how do you obtain the segmentation head label? I'm asking because the dataset only provides image-level information for extravasation.\nIs it set like this: If the image indicates a true extravasation injury, is the segmentation training label set to all 1s? If not, is it set to all 0s?\n\nQuestion 2:\nIn your GitHub code, the criterion uses `nn.BCEWithLogitsLoss()` directly without any weights. However, during inference, you added the weight as a post-process. Have you tried using a weighted `BCEWithLogitsLoss`?\n",
    "3012809": "",
    "2529486": "",
    "2516769": "",
    "2516761": "",
    "2485824": "Congratulations and thanks for sharing! For beginners like me, could you explain why using 32 equidistant slices and 3 adjacent slices is better than stacking 96 slices in the same channel? You also employed this technique in the cervical challenge.",
    "2485236": "Thanks for providing us your solution"
  }
}