{
  "id": 447539,
  "title": "12th place solution",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/447539",
  "author_name": "Ahmed El Fazouani",
  "post_date": "2023-10-16T10:43:35.118000",
  "votes": 23,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Thanks kaggle and the hosts for this interesting competition, <br>\nBig thanks to kagglers out there for their great ideas and engaging descussions.<br>\nThanks a lot as well to my great teammate <a href=\"https://www.kaggle.com/siwooyong\" target=\"_blank\">@siwooyong</a> </p>\n<h1>Summary</h1>\n<p>Our solution is an ensemble of one stage approach without segmentation and two stage approach with segmentation.</p>\n<h2>One stage approach (public LB = 0.4Thanks kaggle and the hosts for this interesting competition,</h2>\n<p>Big thanks to kagglers out there for their great ideas and engaging descussions.<br>\nThanks a lot as well to my great teammate <a href=\"https://www.kaggle.com/siwooyong\" target=\"_blank\">@siwooyong</a> </p>\n<h1>Summary</h1>\n<p>Our solution is an ensemble of one stage approach without segmentation and two stage approach with segmentation.</p>\n<h2>One stage approach (public LB = 0.5, private LB = 0.45)</h2>\n<h3>Data Pre-Processing</h3>\n<p>If the dimension for the image is greater than (512, 512), we cropped the area with a higher density of pixels to get a (512, 512) image, then the input is resized to (96, 256, 256) for each serie following the same preprocessing steps that were used by hengck23 in his great <a href=\"https://www.kaggle.com/code/hengck23/lb0-55-2-5d-3d-sample-model\" target=\"_blank\">notebook</a></p>\n<h3>Model : resnest50d + GRU Attention</h3>\n<p>We tried to predict each target independently from the others so we have 13 outputs</p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        self.seq_len = seq_len\n        self.model_arch = model_arch\n        self.model = timm.create_model(model_arch, in_chans=, pretrained=pretrained)\n\n\n        cnn_feature = self.model.fc.in_features\n        self.model.global_pool = nn.Identity()\n        self.model.fc = nn.Identity()\n        self.pooling = nn.AdaptiveAvgPool2d()\n\n\n        self.spatialdropout = SpatialDropout(CFG.dropout)\n        self.gru = nn.GRU(cnn_feature, hidden_dim, num_layers=, batch_first=, bidirectional=)\n        self.mlp_attention_layer = MLPAttentionNetwork( * hidden_dim)\n        self.logits = nn.Sequential(\n            nn.Linear( * hidden_dim, ),\n        )\n     ():\n        bs = x.size()\n        x = x.reshape(bs*self.seq_len//, , x.size(), x.size())\n        features = self.model(x)\n        features = self.pooling(features).view(bs*self.seq_len//, -)\n        features = self.spatialdropout(features) \n        \n        features = features.reshape(bs, self.seq_len//, -) \n        features, _ = self.gru(features)            \n        atten_out = self.mlp_attention_layer(features) \n        pred = self.logits(atten_out)\n        pred = pred.view(bs, -)\n         pred\n</code></pre>\n<h3>Augmentation</h3>\n<ul>\n<li>Mixup</li>\n<li>Random crop + resize</li>\n<li>Random shift, scale, rotate</li>\n<li>shuffle randomly the indexes of the sequence, but respecting the same order and keeping the dependency between each three consecutive images:</li>\n</ul>\n<pre><code>    inds = np.random.choice(np.arange(, -), , replace = )\n    inds.sort()\n    inds = np.stack([inds-, inds, inds+]).T.flatten()\n    image = image[inds]\n</code></pre>\n<p>Loss : BCEWithLogitsLoss<br>\nscheduler : CosineAnnealingLR<br>\noptimizer : AdamW<br>\nlearning rate : 5e-5</p>\n<h3>Postprocessing</h3>\n<p>We simply multiplied the output by the weights of the competition metric :</p>\n<pre><code>preds.loc[:, [, , , ]] *= \npreds.loc[:, [, , ]] *= \npreds.loc[:, []] *= \n</code></pre>\n<h2>Two stage approach (public LB = 0.45, private LB = 0.43)</h2>\n<h3>stage1 : Segmentation</h3>\n<p>Model : regnety002 + unet</p>\n<p>Even with only 160 of 200 data (1th fold) used as training data, the model has already shown good performance.</p>\n<pre><code> (nn.Module):\n     ():\n        (SegModel, self).__init__()\n\n        self.n_classes = (\n            [\n                ,\n                ,\n                ,\n                ,\n                ,\n                \n            ])\n        in_chans = \n        self.encoder = timm.create_model(\n            ,\n            pretrained=,\n            features_only=,\n            in_chans=in_chans,\n        )\n        encoder_channels = (\n            [in_chans]\n            + [\n                self.encoder.feature_info[i][]\n                 i  ((self.encoder.feature_info))\n            ]\n        )\n        self.decoder = UnetDecoder(\n            encoder_channels=encoder_channels,\n            decoder_channels=(, , , , ),\n            n_blocks=,\n            use_batchnorm=,\n            center=,\n            attention_type=,\n        )\n\n        self.segmentation_head = SegmentationHead(\n            in_channels=,\n            out_channels=self.n_classes,\n            activation=,\n            kernel_size=,\n        )\n\n        self.bce_seg = nn.BCEWithLogitsLoss()\n\n     ():\n        enc_out = self.encoder(x_in)\n\n        decoder_out = self.decoder(*[x_in] + enc_out)\n        x_seg = self.segmentation_head(decoder_out)\n\n         nn.Sigmoid()(x_seg)\n</code></pre>\n<h3>stage2 : 2.5DCNN</h3>\n<h4>Data Pre-Processing:</h4>\n<p>We used the segmentation logits obtained from stage1 to crop livers, spleen, and kidney, and then resized each to (96, 224, 224). <br>\n(We use 10-size padding when we crop the organs with segmentation logits)<br>\nIn addition, full ct data not cropped is resized to (128, 224, 224) and a total of four inputs are put into the model (full_video, crop_liver, crop_spleen, crop_kidney)</p>\n<h4>Model : regnety002 + transformer</h4>\n<p>We initially used a custom any_injury_loss function, but found that it did not improve the performance.  For the model input channel, we experimented with different values, including 2, 3, 4, and 8. <br>\nWe found that a channel size of 2 performed the best, we also initially tried using a shared CNN and transformer model for all organs,  but found that separate CNN and transformer models for each organ performed better.  we also experimented with increasing the size of the CNN (using ConvNeXt and EfficientNet models), but this resulted in a decrease in performance.  Therefore, we used the RegNet002 model, which is a smaller CNN model.</p>\n<pre><code> (nn.Module):\n     ():\n        (FeatureExtractor, self).__init__()\n\n        self.hidden = hidden\n        self.num_channel = num_channel\n\n        self.cnn = timm.create_model(model_name = ,\n                                     pretrained = ,\n                                     num_classes = ,\n                                     in_chans = num_channel)\n\n        self.fc = nn.Linear(hidden, hidden//)\n\n     ():\n        batch_size, num_frame, h, w = x.shape\n        x = x.reshape(batch_size, num_frame//self.num_channel, self.num_channel, h, w)\n        x = x.reshape(-, self.num_channel, h, w)\n        x = self.cnn(x)\n        x = x.reshape(batch_size, num_frame//self.num_channel, self.hidden)\n\n        x = self.fc(x)\n         x\n\n (nn.Module):\n     ():\n        (ContextProcessor, self).__init__()\n        self.transformer = RobertaPreLayerNormModel(\n            RobertaPreLayerNormConfig(\n                hidden_size = hidden//,\n                num_hidden_layers = ,\n                num_attention_heads = ,\n                intermediate_size = hidden*,\n                hidden_act = ,\n                )\n            )\n\n         self.transformer.embeddings.word_embeddings\n\n        self.dense = nn.Linear(hidden, hidden)\n        self.activation = nn.ReLU()\n\n\n     ():\n        x = self.transformer(inputs_embeds = x).last_hidden_state\n\n        apool = torch.mean(x, dim = )\n        mpool, _ = torch.(x, dim = )\n        x = torch.cat([mpool, apool], dim = -)\n\n        x = self.dense(x)\n        x = self.activation(x)\n         x\n\n (nn.Module):\n     ():\n        (Custom3DCNN, self).__init__()\n\n        self.full_extractor = FeatureExtractor(hidden=hidden, num_channel=num_channel)\n        self.kidney_extractor = FeatureExtractor(hidden=hidden, num_channel=num_channel)\n        self.liver_extractor = FeatureExtractor(hidden=hidden, num_channel=num_channel)\n        self.spleen_extractor = FeatureExtractor(hidden=hidden, num_channel=num_channel)\n\n        self.full_processor = ContextProcessor(hidden=hidden)\n        self.kidney_processor = ContextProcessor(hidden=hidden)\n        self.liver_processor = ContextProcessor(hidden=hidden)\n        self.spleen_processor = ContextProcessor(hidden=hidden)\n\n        self.bowel = nn.Linear(hidden, )\n        self.extravasation = nn.Linear(hidden, )\n        self.kidney = nn.Linear(hidden, )\n        self.liver = nn.Linear(hidden, )\n        self.spleen = nn.Linear(hidden, )\n\n        self.softmax = nn.Softmax(dim = -)\n\n     ():\n        full_output = self.full_extractor(full_input)\n        kidney_output = self.kidney_extractor(crop_kidney)\n        liver_output = self.liver_extractor(crop_liver)\n        spleen_output = self.spleen_extractor(crop_spleen)\n\n        full_output2 = self.full_processor(torch.cat([full_output, kidney_output, liver_output, spleen_output], dim = ))\n        kidney_output2 = self.kidney_processor(torch.cat([full_output, kidney_output], dim = ))\n        liver_output2 = self.liver_processor(torch.cat([full_output, liver_output], dim = ))\n        spleen_output2 = self.spleen_processor(torch.cat([full_output, spleen_output], dim = ))\n\n        bowel = self.bowel(full_output2)\n        extravasation = self.extravasation(full_output2)\n        kidney = self.kidney(kidney_output2)\n        liver = self.liver(liver_output2)\n        spleen = self.spleen(spleen_output2)\n\n\n        any_injury = torch.stack([\n            self.softmax(bowel)[:, ],\n            self.softmax(extravasation)[:, ],\n            self.softmax(kidney)[:, ],\n            self.softmax(liver)[:, ],\n            self.softmax(spleen)[:, ]\n        ], dim = -)\n        any_injury =  - any_injury\n        any_injury, _ = any_injury.()\n         bowel, extravasation, kidney, liver, spleen, any_injury\n</code></pre>\n<h4>Augmentation</h4>\n<pre><code> (nn.Module):\n     ():\n        (CustomAug, self).__init__()\n        self.prob = prob\n\n        self.do_random_rotate = v2.RandomRotation(\n            degrees = (-, ),\n            interpolation = torchvision.transforms.InterpolationMode.BILINEAR,\n            expand = ,\n            center = ,\n            fill = \n        )\n        self.do_random_scale = v2.ScaleJitter(\n            target_size = [s, s],\n            scale_range = (, ),\n            interpolation = torchvision.transforms.InterpolationMode.BILINEAR,\n            antialias = )\n\n        self.do_random_crop = v2.RandomCrop(\n            size = [s, s],\n            \n            pad_if_needed = ,\n            fill = ,\n            padding_mode = \n        )\n\n        self.do_horizontal_flip = v2.RandomHorizontalFlip(self.prob)\n        self.do_vertical_flip = v2.RandomVerticalFlip(self.prob)\n     ():\n         np.random.rand() &lt; self.prob:\n            x = self.do_random_rotate(x)\n\n         np.random.rand() &lt; self.prob:\n            x = self.do_random_scale(x)\n            x = self.do_random_crop(x)\n\n        x = self.do_horizontal_flip(x)\n        x = self.do_vertical_flip(x)\n         x\n</code></pre>\n<p>Loss : nn.CrossEntropyLoss(no class weight)<br>\nscheduler : cosine_schedule_with_warmup<br>\noptimizer : AdamW<br>\nlearning rate :2e-4</p>\n<h4>Postprocessing</h4>\n<p>We multiplied by the value that maximizes the validation score for each pred_df obtained for each fold.</p>\n<pre><code>weights = [\n    [, , , , , , , ],\n    [, , , , , , , ],\n    [, , , , , , , ],\n    [, , , , , , , ],\n    [, , , , , , , ]\n]\n\ny_pred = pred_df.copy().groupby().mean().reset_index()\n\nw1, w2, w3, w4, w5, w6, w7, w8 = weights[i]\n\ny_pred[] *= w1\ny_pred[] *= w2\ny_pred[] *= w3\ny_pred[] *= w4\ny_pred[] *= w5\ny_pred[] *= w6\ny_pred[] *= w7\ny_pred[] *= w8\n\ny_pred = y_pred ** \n</code></pre>\n<h3>Reference</h3>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362607\" target=\"_blank\">RSNA 2022 1st place solution</a><br>\n<a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/365115\" target=\"_blank\">RSNA 2022 2nd place solution</a><br>\n<a href=\"https://github.com/pascal-pfeiffer/kaggle-rsna-2022-5th-place\" target=\"_blank\">RSNA 2022 5th place solution</a><br>\n<a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/435053\" target=\"_blank\">descussion</a></p>\n<h2>Code :</h2>\n<p><a href=\"https://github.com/siwooyong/RSNA-2023-Abdominal-Trauma-Detection\" target=\"_blank\">https://github.com/siwooyong/RSNA-2023-Abdominal-Trauma-Detection</a><br>\nInference notebook : <a href=\"https://www.kaggle.com/code/ahmedelfazouan/rsna-atd-channel2-512-inference-ensemble?scriptVersionId=146616538\" target=\"_blank\">https://www.kaggle.com/code/ahmedelfazouan/rsna-atd-channel2-512-inference-ensemble?scriptVersionId=146616538</a></p>",
  "messages": [
    {
      "id": 2484242,
      "postDate": "2023-10-16T10:43:35.117Z",
      "content": "<p>Thanks kaggle and the hosts for this interesting competition, <br>\nBig thanks to kagglers out there for their great ideas and engaging descussions.<br>\nThanks a lot as well to my great teammate <a href=\"https://www.kaggle.com/siwooyong\" target=\"_blank\">@siwooyong</a> </p>\n<h1>Summary</h1>\n<p>Our solution is an ensemble of one stage approach without segmentation and two stage approach with segmentation.</p>\n<h2>One stage approach (public LB = 0.4Thanks kaggle and the hosts for this interesting competition,</h2>\n<p>Big thanks to kagglers out there for their great ideas and engaging descussions.<br>\nThanks a lot as well to my great teammate <a href=\"https://www.kaggle.com/siwooyong\" target=\"_blank\">@siwooyong</a> </p>\n<h1>Summary</h1>\n<p>Our solution is an ensemble of one stage approach without segmentation and two stage approach with segmentation.</p>\n<h2>One stage approach (public LB = 0.5, private LB = 0.45)</h2>\n<h3>Data Pre-Processing</h3>\n<p>If the dimension for the image is greater than (512, 512), we cropped the area with a higher density of pixels to get a (512, 512) image, then the input is resized to (96, 256, 256) for each serie following the same preprocessing steps that were used by hengck23 in his great <a href=\"https://www.kaggle.com/code/hengck23/lb0-55-2-5d-3d-sample-model\" target=\"_blank\">notebook</a></p>\n<h3>Model : resnest50d + GRU Attention</h3>\n<p>We tried to predict each target independently from the others so we have 13 outputs</p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        self.seq_len = seq_len\n        self.model_arch = model_arch\n        self.model = timm.create_model(model_arch, in_chans=, pretrained=pretrained)\n\n\n        cnn_feature = self.model.fc.in_features\n        self.model.global_pool = nn.Identity()\n        self.model.fc = nn.Identity()\n        self.pooling = nn.AdaptiveAvgPool2d()\n\n\n        self.spatialdropout = SpatialDropout(CFG.dropout)\n        self.gru = nn.GRU(cnn_feature, hidden_dim, num_layers=, batch_first=, bidirectional=)\n        self.mlp_attention_layer = MLPAttentionNetwork( * hidden_dim)\n        self.logits = nn.Sequential(\n            nn.Linear( * hidden_dim, ),\n        )\n     ():\n        bs = x.size()\n        x = x.reshape(bs*self.seq_len//, , x.size(), x.size())\n        features = self.model(x)\n        features = self.pooling(features).view(bs*self.seq_len//, -)\n        features = self.spatialdropout(features) \n        \n        features = features.reshape(bs, self.seq_len//, -) \n        features, _ = self.gru(features)            \n        atten_out = self.mlp_attention_layer(features) \n        pred = self.logits(atten_out)\n        pred = pred.view(bs, -)\n         pred\n</code></pre>\n<h3>Augmentation</h3>\n<ul>\n<li>Mixup</li>\n<li>Random crop + resize</li>\n<li>Random shift, scale, rotate</li>\n<li>shuffle randomly the indexes of the sequence, but respecting the same order and keeping the dependency between each three consecutive images:</li>\n</ul>\n<pre><code>    inds = np.random.choice(np.arange(, -), , replace = )\n    inds.sort()\n    inds = np.stack([inds-, inds, inds+]).T.flatten()\n    image = image[inds]\n</code></pre>\n<p>Loss : BCEWithLogitsLoss<br>\nscheduler : CosineAnnealingLR<br>\noptimizer : AdamW<br>\nlearning rate : 5e-5</p>\n<h3>Postprocessing</h3>\n<p>We simply multiplied the output by the weights of the competition metric :</p>\n<pre><code>preds.loc[:, [, , , ]] *= \npreds.loc[:, [, , ]] *= \npreds.loc[:, []] *= \n</code></pre>\n<h2>Two stage approach (public LB = 0.45, private LB = 0.43)</h2>\n<h3>stage1 : Segmentation</h3>\n<p>Model : regnety002 + unet</p>\n<p>Even with only 160 of 200 data (1th fold) used as training data, the model has already shown good performance.</p>\n<pre><code> (nn.Module):\n     ():\n        (SegModel, self).__init__()\n\n        self.n_classes = (\n            [\n                ,\n                ,\n                ,\n                ,\n                ,\n                \n            ])\n        in_chans = \n        self.encoder = timm.create_model(\n            ,\n            pretrained=,\n            features_only=,\n            in_chans=in_chans,\n        )\n        encoder_channels = (\n            [in_chans]\n            + [\n                self.encoder.feature_info[i][]\n                 i  ((self.encoder.feature_info))\n            ]\n        )\n        self.decoder = UnetDecoder(\n            encoder_channels=encoder_channels,\n            decoder_channels=(, , , , ),\n            n_blocks=,\n            use_batchnorm=,\n            center=,\n            attention_type=,\n        )\n\n        self.segmentation_head = SegmentationHead(\n            in_channels=,\n            out_channels=self.n_classes,\n            activation=,\n            kernel_size=,\n        )\n\n        self.bce_seg = nn.BCEWithLogitsLoss()\n\n     ():\n        enc_out = self.encoder(x_in)\n\n        decoder_out = self.decoder(*[x_in] + enc_out)\n        x_seg = self.segmentation_head(decoder_out)\n\n         nn.Sigmoid()(x_seg)\n</code></pre>\n<h3>stage2 : 2.5DCNN</h3>\n<h4>Data Pre-Processing:</h4>\n<p>We used the segmentation logits obtained from stage1 to crop livers, spleen, and kidney, and then resized each to (96, 224, 224). <br>\n(We use 10-size padding when we crop the organs with segmentation logits)<br>\nIn addition, full ct data not cropped is resized to (128, 224, 224) and a total of four inputs are put into the model (full_video, crop_liver, crop_spleen, crop_kidney)</p>\n<h4>Model : regnety002 + transformer</h4>\n<p>We initially used a custom any_injury_loss function, but found that it did not improve the performance.  For the model input channel, we experimented with different values, including 2, 3, 4, and 8. <br>\nWe found that a channel size of 2 performed the best, we also initially tried using a shared CNN and transformer model for all organs,  but found that separate CNN and transformer models for each organ performed better.  we also experimented with increasing the size of the CNN (using ConvNeXt and EfficientNet models), but this resulted in a decrease in performance.  Therefore, we used the RegNet002 model, which is a smaller CNN model.</p>\n<pre><code> (nn.Module):\n     ():\n        (FeatureExtractor, self).__init__()\n\n        self.hidden = hidden\n        self.num_channel = num_channel\n\n        self.cnn = timm.create_model(model_name = ,\n                                     pretrained = ,\n                                     num_classes = ,\n                                     in_chans = num_channel)\n\n        self.fc = nn.Linear(hidden, hidden//)\n\n     ():\n        batch_size, num_frame, h, w = x.shape\n        x = x.reshape(batch_size, num_frame//self.num_channel, self.num_channel, h, w)\n        x = x.reshape(-, self.num_channel, h, w)\n        x = self.cnn(x)\n        x = x.reshape(batch_size, num_frame//self.num_channel, self.hidden)\n\n        x = self.fc(x)\n         x\n\n (nn.Module):\n     ():\n        (ContextProcessor, self).__init__()\n        self.transformer = RobertaPreLayerNormModel(\n            RobertaPreLayerNormConfig(\n                hidden_size = hidden//,\n                num_hidden_layers = ,\n                num_attention_heads = ,\n                intermediate_size = hidden*,\n                hidden_act = ,\n                )\n            )\n\n         self.transformer.embeddings.word_embeddings\n\n        self.dense = nn.Linear(hidden, hidden)\n        self.activation = nn.ReLU()\n\n\n     ():\n        x = self.transformer(inputs_embeds = x).last_hidden_state\n\n        apool = torch.mean(x, dim = )\n        mpool, _ = torch.(x, dim = )\n        x = torch.cat([mpool, apool], dim = -)\n\n        x = self.dense(x)\n        x = self.activation(x)\n         x\n\n (nn.Module):\n     ():\n        (Custom3DCNN, self).__init__()\n\n        self.full_extractor = FeatureExtractor(hidden=hidden, num_channel=num_channel)\n        self.kidney_extractor = FeatureExtractor(hidden=hidden, num_channel=num_channel)\n        self.liver_extractor = FeatureExtractor(hidden=hidden, num_channel=num_channel)\n        self.spleen_extractor = FeatureExtractor(hidden=hidden, num_channel=num_channel)\n\n        self.full_processor = ContextProcessor(hidden=hidden)\n        self.kidney_processor = ContextProcessor(hidden=hidden)\n        self.liver_processor = ContextProcessor(hidden=hidden)\n        self.spleen_processor = ContextProcessor(hidden=hidden)\n\n        self.bowel = nn.Linear(hidden, )\n        self.extravasation = nn.Linear(hidden, )\n        self.kidney = nn.Linear(hidden, )\n        self.liver = nn.Linear(hidden, )\n        self.spleen = nn.Linear(hidden, )\n\n        self.softmax = nn.Softmax(dim = -)\n\n     ():\n        full_output = self.full_extractor(full_input)\n        kidney_output = self.kidney_extractor(crop_kidney)\n        liver_output = self.liver_extractor(crop_liver)\n        spleen_output = self.spleen_extractor(crop_spleen)\n\n        full_output2 = self.full_processor(torch.cat([full_output, kidney_output, liver_output, spleen_output], dim = ))\n        kidney_output2 = self.kidney_processor(torch.cat([full_output, kidney_output], dim = ))\n        liver_output2 = self.liver_processor(torch.cat([full_output, liver_output], dim = ))\n        spleen_output2 = self.spleen_processor(torch.cat([full_output, spleen_output], dim = ))\n\n        bowel = self.bowel(full_output2)\n        extravasation = self.extravasation(full_output2)\n        kidney = self.kidney(kidney_output2)\n        liver = self.liver(liver_output2)\n        spleen = self.spleen(spleen_output2)\n\n\n        any_injury = torch.stack([\n            self.softmax(bowel)[:, ],\n            self.softmax(extravasation)[:, ],\n            self.softmax(kidney)[:, ],\n            self.softmax(liver)[:, ],\n            self.softmax(spleen)[:, ]\n        ], dim = -)\n        any_injury =  - any_injury\n        any_injury, _ = any_injury.()\n         bowel, extravasation, kidney, liver, spleen, any_injury\n</code></pre>\n<h4>Augmentation</h4>\n<pre><code> (nn.Module):\n     ():\n        (CustomAug, self).__init__()\n        self.prob = prob\n\n        self.do_random_rotate = v2.RandomRotation(\n            degrees = (-, ),\n            interpolation = torchvision.transforms.InterpolationMode.BILINEAR,\n            expand = ,\n            center = ,\n            fill = \n        )\n        self.do_random_scale = v2.ScaleJitter(\n            target_size = [s, s],\n            scale_range = (, ),\n            interpolation = torchvision.transforms.InterpolationMode.BILINEAR,\n            antialias = )\n\n        self.do_random_crop = v2.RandomCrop(\n            size = [s, s],\n            \n            pad_if_needed = ,\n            fill = ,\n            padding_mode = \n        )\n\n        self.do_horizontal_flip = v2.RandomHorizontalFlip(self.prob)\n        self.do_vertical_flip = v2.RandomVerticalFlip(self.prob)\n     ():\n         np.random.rand() &lt; self.prob:\n            x = self.do_random_rotate(x)\n\n         np.random.rand() &lt; self.prob:\n            x = self.do_random_scale(x)\n            x = self.do_random_crop(x)\n\n        x = self.do_horizontal_flip(x)\n        x = self.do_vertical_flip(x)\n         x\n</code></pre>\n<p>Loss : nn.CrossEntropyLoss(no class weight)<br>\nscheduler : cosine_schedule_with_warmup<br>\noptimizer : AdamW<br>\nlearning rate :2e-4</p>\n<h4>Postprocessing</h4>\n<p>We multiplied by the value that maximizes the validation score for each pred_df obtained for each fold.</p>\n<pre><code>weights = [\n    [, , , , , , , ],\n    [, , , , , , , ],\n    [, , , , , , , ],\n    [, , , , , , , ],\n    [, , , , , , , ]\n]\n\ny_pred = pred_df.copy().groupby().mean().reset_index()\n\nw1, w2, w3, w4, w5, w6, w7, w8 = weights[i]\n\ny_pred[] *= w1\ny_pred[] *= w2\ny_pred[] *= w3\ny_pred[] *= w4\ny_pred[] *= w5\ny_pred[] *= w6\ny_pred[] *= w7\ny_pred[] *= w8\n\ny_pred = y_pred ** \n</code></pre>\n<h3>Reference</h3>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362607\" target=\"_blank\">RSNA 2022 1st place solution</a><br>\n<a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/365115\" target=\"_blank\">RSNA 2022 2nd place solution</a><br>\n<a href=\"https://github.com/pascal-pfeiffer/kaggle-rsna-2022-5th-place\" target=\"_blank\">RSNA 2022 5th place solution</a><br>\n<a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>'s <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/435053\" target=\"_blank\">descussion</a></p>\n<h2>Code :</h2>\n<p><a href=\"https://github.com/siwooyong/RSNA-2023-Abdominal-Trauma-Detection\" target=\"_blank\">https://github.com/siwooyong/RSNA-2023-Abdominal-Trauma-Detection</a><br>\nInference notebook : <a href=\"https://www.kaggle.com/code/ahmedelfazouan/rsna-atd-channel2-512-inference-ensemble?scriptVersionId=146616538\" target=\"_blank\">https://www.kaggle.com/code/ahmedelfazouan/rsna-atd-channel2-512-inference-ensemble?scriptVersionId=146616538</a></p>",
      "rawMarkdown": "Thanks kaggle and the hosts for this interesting competition, \nBig thanks to kagglers out there for their great ideas and engaging descussions.\nThanks a lot as well to my great teammate @siwooyong \n\n# Summary\n\nOur solution is an ensemble of one stage approach without segmentation and two stage approach with segmentation.\n\n## One stage approach (public LB = 0.4Thanks kaggle and the hosts for this interesting competition, \nBig thanks to kagglers out there for their great ideas and engaging descussions.\nThanks a lot as well to my great teammate @siwooyong \n\n# Summary\n\nOur solution is an ensemble of one stage approach without segmentation and two stage approach with segmentation.\n\n## One stage approach (public LB = 0.5, private LB = 0.45)\n\n### Data Pre-Processing\nIf the dimension for the image is greater than (512, 512), we cropped the area with a higher density of pixels to get a (512, 512) image, then the input is resized to (96, 256, 256) for each serie following the same preprocessing steps that were used by hengck23 in his great [notebook](https://www.kaggle.com/code/hengck23/lb0-55-2-5d-3d-sample-model)\n\n### Model : resnest50d + GRU Attention\nWe tried to predict each target independently from the others so we have 13 outputs\n\n```python\nclass RSNAClassifier(nn.Module):\n    def __init__(self, model_arch, hidden_dim=128, seq_len=3, pretrained=False):\n        super().__init__()\n        self.seq_len = seq_len\n        self.model_arch = model_arch\n        self.model = timm.create_model(model_arch, in_chans=3, pretrained=pretrained)\n\n\n        cnn_feature = self.model.fc.in_features\n        self.model.global_pool = nn.Identity()\n        self.model.fc = nn.Identity()\n        self.pooling = nn.AdaptiveAvgPool2d(1)\n\n\n        self.spatialdropout = SpatialDropout(CFG.dropout)\n        self.gru = nn.GRU(cnn_feature, hidden_dim, num_layers=2, batch_first=True, bidirectional=True)\n        self.mlp_attention_layer = MLPAttentionNetwork(2 * hidden_dim)\n        self.logits = nn.Sequential(\n            nn.Linear(2 * hidden_dim, 13),\n        )\n    def forward(self, x):\n        bs = x.size(0)\n        x = x.reshape(bs*self.seq_len//3, 3, x.size(2), x.size(3))\n        features = self.model(x)\n        features = self.pooling(features).view(bs*self.seq_len//3, -1)\n        features = self.spatialdropout(features) \n        # print(features.shape)\n        features = features.reshape(bs, self.seq_len//3, -1) \n        features, _ = self.gru(features)            \n        atten_out = self.mlp_attention_layer(features) \n        pred = self.logits(atten_out)\n        pred = pred.view(bs, -1)\n        return pred\n```\n\n### Augmentation\n * Mixup\n * Random crop + resize\n * Random shift, scale, rotate\n * shuffle randomly the indexes of the sequence, but respecting the same order and keeping the dependency between each three consecutive images:\n```python\n\tinds = np.random.choice(np.arange(1, 96-1), 32, replace = False)\n\tinds.sort()\n\tinds = np.stack([inds-1, inds, inds+1]).T.flatten()\n\timage = image[inds]\n\n```\nLoss : BCEWithLogitsLoss\nscheduler : CosineAnnealingLR\noptimizer : AdamW\nlearning rate : 5e-5\n\n### Postprocessing \nWe simply multiplied the output by the weights of the competition metric :\n```python\npreds.loc[:, ['bowel_injury', 'kidney_low', 'liver_low', 'spleen_low']] *= 2\npreds.loc[:, ['kidney_high', 'liver_high', 'spleen_high']] *= 4\npreds.loc[:, ['extravasation_injury']] *= 6\n```\n\n## Two stage approach (public LB = 0.45, private LB = 0.43)\n### stage1 : Segmentation\n\nModel : regnety002 + unet\n\nEven with only 160 of 200 data (1th fold) used as training data, the model has already shown good performance.\n\n```python\nclass SegModel(nn.Module):\n    def __init__(self):\n        super(SegModel, self).__init__()\n   \n        self.n_classes = len(\n            [\n                'background',\n                'liver',\n                'spleen',\n                'left kidney',\n                'right kidney',\n                'bowel'\n            ])\n        in_chans = 1\n        self.encoder = timm.create_model(\n            'regnety_002',\n            pretrained=False,\n            features_only=True,\n            in_chans=in_chans,\n        )\n        encoder_channels = tuple(\n            [in_chans]\n            + [\n                self.encoder.feature_info[i][\"num_chs\"]\n                for i in range(len(self.encoder.feature_info))\n            ]\n        )\n        self.decoder = UnetDecoder(\n            encoder_channels=encoder_channels,\n            decoder_channels=(256, 128, 64, 32, 16),\n            n_blocks=5,\n            use_batchnorm=True,\n            center=False,\n            attention_type=None,\n        )\n\n        self.segmentation_head = SegmentationHead(\n            in_channels=16,\n            out_channels=self.n_classes,\n            activation=None,\n            kernel_size=3,\n        )\n\n        self.bce_seg = nn.BCEWithLogitsLoss()\n\n    def forward(self, x_in):\n        enc_out = self.encoder(x_in)\n\n        decoder_out = self.decoder(*[x_in] + enc_out)\n        x_seg = self.segmentation_head(decoder_out)\n\n        return nn.Sigmoid()(x_seg)\n```\n\n### stage2 : 2.5DCNN \n\n#### Data Pre-Processing:\nWe used the segmentation logits obtained from stage1 to crop livers, spleen, and kidney, and then resized each to (96, 224, 224). \n(We use 10-size padding when we crop the organs with segmentation logits)\nIn addition, full ct data not cropped is resized to (128, 224, 224) and a total of four inputs are put into the model (full_video, crop_liver, crop_spleen, crop_kidney)\n\n#### Model : regnety002 + transformer\n\nWe initially used a custom any_injury_loss function, but found that it did not improve the performance.  For the model input channel, we experimented with different values, including 2, 3, 4, and 8. \nWe found that a channel size of 2 performed the best, we also initially tried using a shared CNN and transformer model for all organs,  but found that separate CNN and transformer models for each organ performed better.  we also experimented with increasing the size of the CNN (using ConvNeXt and EfficientNet models), but this resulted in a decrease in performance.  Therefore, we used the RegNet002 model, which is a smaller CNN model.\n\n```python\nclass FeatureExtractor(nn.Module):\n    def __init__(self, hidden, num_channel):\n        super(FeatureExtractor, self).__init__()\n\n        self.hidden = hidden\n        self.num_channel = num_channel\n\n        self.cnn = timm.create_model(model_name = 'regnety_002',\n                                     pretrained = True,\n                                     num_classes = 0,\n                                     in_chans = num_channel)\n\n        self.fc = nn.Linear(hidden, hidden//2)\n\n    def forward(self, x):\n        batch_size, num_frame, h, w = x.shape\n        x = x.reshape(batch_size, num_frame//self.num_channel, self.num_channel, h, w)\n        x = x.reshape(-1, self.num_channel, h, w)\n        x = self.cnn(x)\n        x = x.reshape(batch_size, num_frame//self.num_channel, self.hidden)\n\n        x = self.fc(x)\n        return x\n\nclass ContextProcessor(nn.Module):\n    def __init__(self, hidden):\n        super(ContextProcessor, self).__init__()\n        self.transformer = RobertaPreLayerNormModel(\n            RobertaPreLayerNormConfig(\n                hidden_size = hidden//2,\n                num_hidden_layers = 1,\n                num_attention_heads = 4,\n                intermediate_size = hidden*2,\n                hidden_act = 'gelu_new',\n                )\n            )\n\n        del self.transformer.embeddings.word_embeddings\n\n        self.dense = nn.Linear(hidden, hidden)\n        self.activation = nn.ReLU()\n\n\n    def forward(self, x):\n        x = self.transformer(inputs_embeds = x).last_hidden_state\n\n        apool = torch.mean(x, dim = 1)\n        mpool, _ = torch.max(x, dim = 1)\n        x = torch.cat([mpool, apool], dim = -1)\n\n        x = self.dense(x)\n        x = self.activation(x)\n        return x\n\nclass Custom3DCNN(nn.Module):\n    def __init__(self, hidden = 368, num_channel = 2):\n        super(Custom3DCNN, self).__init__()\n\n        self.full_extractor = FeatureExtractor(hidden=hidden, num_channel=num_channel)\n        self.kidney_extractor = FeatureExtractor(hidden=hidden, num_channel=num_channel)\n        self.liver_extractor = FeatureExtractor(hidden=hidden, num_channel=num_channel)\n        self.spleen_extractor = FeatureExtractor(hidden=hidden, num_channel=num_channel)\n\n        self.full_processor = ContextProcessor(hidden=hidden)\n        self.kidney_processor = ContextProcessor(hidden=hidden)\n        self.liver_processor = ContextProcessor(hidden=hidden)\n        self.spleen_processor = ContextProcessor(hidden=hidden)\n\n        self.bowel = nn.Linear(hidden, 2)\n        self.extravasation = nn.Linear(hidden, 2)\n        self.kidney = nn.Linear(hidden, 3)\n        self.liver = nn.Linear(hidden, 3)\n        self.spleen = nn.Linear(hidden, 3)\n\n        self.softmax = nn.Softmax(dim = -1)\n\n    def forward(self, full_input, crop_liver, crop_spleen, crop_kidney, mask, mode):\n        full_output = self.full_extractor(full_input)\n        kidney_output = self.kidney_extractor(crop_kidney)\n        liver_output = self.liver_extractor(crop_liver)\n        spleen_output = self.spleen_extractor(crop_spleen)\n\n        full_output2 = self.full_processor(torch.cat([full_output, kidney_output, liver_output, spleen_output], dim = 1))\n        kidney_output2 = self.kidney_processor(torch.cat([full_output, kidney_output], dim = 1))\n        liver_output2 = self.liver_processor(torch.cat([full_output, liver_output], dim = 1))\n        spleen_output2 = self.spleen_processor(torch.cat([full_output, spleen_output], dim = 1))\n\n        bowel = self.bowel(full_output2)\n        extravasation = self.extravasation(full_output2)\n        kidney = self.kidney(kidney_output2)\n        liver = self.liver(liver_output2)\n        spleen = self.spleen(spleen_output2)\n\n\n        any_injury = torch.stack([\n            self.softmax(bowel)[:, 0],\n            self.softmax(extravasation)[:, 0],\n            self.softmax(kidney)[:, 0],\n            self.softmax(liver)[:, 0],\n            self.softmax(spleen)[:, 0]\n        ], dim = -1)\n        any_injury = 1 - any_injury\n        any_injury, _ = any_injury.max(1)\n        return bowel, extravasation, kidney, liver, spleen, any_injury\n\n```\n\n#### Augmentation\n\n```python\nclass CustomAug(nn.Module):\n    def __init__(self, prob = 0.5, s = 224):\n        super(CustomAug, self).__init__()\n        self.prob = prob\n\n        self.do_random_rotate = v2.RandomRotation(\n            degrees = (-45, 45),\n            interpolation = torchvision.transforms.InterpolationMode.BILINEAR,\n            expand = False,\n            center = None,\n            fill = 0\n        )\n        self.do_random_scale = v2.ScaleJitter(\n            target_size = [s, s],\n            scale_range = (0.8, 1.2),\n            interpolation = torchvision.transforms.InterpolationMode.BILINEAR,\n            antialias = True)\n\n        self.do_random_crop = v2.RandomCrop(\n            size = [s, s],\n            #padding = None,\n            pad_if_needed = True,\n            fill = 0,\n            padding_mode = 'constant'\n        )\n\n        self.do_horizontal_flip = v2.RandomHorizontalFlip(self.prob)\n        self.do_vertical_flip = v2.RandomVerticalFlip(self.prob)\n    def forward(self, x):\n        if np.random.rand() < self.prob:\n            x = self.do_random_rotate(x)\n\n        if np.random.rand() < self.prob:\n            x = self.do_random_scale(x)\n            x = self.do_random_crop(x)\n\n        x = self.do_horizontal_flip(x)\n        x = self.do_vertical_flip(x)\n        return x\n```\n\nLoss : nn.CrossEntropyLoss(no class weight)\nscheduler : cosine_schedule_with_warmup\noptimizer : AdamW\nlearning rate :2e-4\n\n#### Postprocessing\nWe multiplied by the value that maximizes the validation score for each pred_df obtained for each fold.\n\n```python\nweights = [\n    [0.9, 4, 2, 4, 2, 6, 6, 6],\n    [0.9, 1, 4, 3, 2, 5, 5, 6],\n    [0.2, 3, 2, 1, 2, 4, 2, 6],\n    [0.5, 2, 2, 2, 2, 2, 6, 6],\n    [1, 2, 3, 2, 6, 3, 6, 5]\n]\n\ny_pred = pred_df.copy().groupby('patient_id').mean().reset_index()\n\nw1, w2, w3, w4, w5, w6, w7, w8 = weights[i]\n\ny_pred['bowel_injury'] *= w1\ny_pred['kidney_low'] *= w2\ny_pred['liver_low'] *= w3\ny_pred['spleen_low'] *= w4\ny_pred['kidney_high'] *= w5\ny_pred['liver_high'] *= w6\ny_pred['spleen_high'] *= w7\ny_pred['extravasation_injury'] *= w8\n\ny_pred = y_pred ** 0.8\n\n```\n\n### Reference\n[RSNA 2022 1st place solution](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362607)\n[RSNA 2022 2nd place solution](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/365115)\n[RSNA 2022 5th place solution](https://github.com/pascal-pfeiffer/kaggle-rsna-2022-5th-place)\n@hengck23's [descussion](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/435053)\n\n## Code :\n https://github.com/siwooyong/RSNA-2023-Abdominal-Trauma-Detection\nInference notebook : https://www.kaggle.com/code/ahmedelfazouan/rsna-atd-channel2-512-inference-ensemble?scriptVersionId=146616538",
      "votes": 23
    },
    {
      "id": 2486316,
      "postDate": "2023-10-17T20:14:39.483Z",
      "content": "<p>Congrats and thanks for the source code, it is a great resource to learn from top solution, as a beginner.</p>",
      "rawMarkdown": "Congrats and thanks for the source code, it is a great resource to learn from top solution, as a beginner.",
      "votes": 1
    },
    {
      "id": 2485222,
      "postDate": "2023-10-17T03:28:15.207Z",
      "content": "<p>Congratulations, and thanks to my great teammate <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a>.</p>",
      "rawMarkdown": "Congratulations, and thanks to my great teammate @ahmedelfazouan.",
      "votes": 1,
      "replies": [
        {
          "id": 2485407,
          "postDate": "2023-10-17T07:01:35.403Z",
          "content": "<p>Thanks, It wasn't possible without your hard work !!</p>",
          "rawMarkdown": "Thanks, It wasn't possible without your hard work !!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2484630,
      "postDate": "2023-10-16T15:47:12.280Z",
      "content": "<p>Congrats on the gold. Great job <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> !</p>",
      "rawMarkdown": "Congrats on the gold. Great job @ahmedelfazouan !",
      "votes": 1,
      "replies": [
        {
          "id": 2484692,
          "postDate": "2023-10-16T16:12:13.510Z",
          "content": "<p>Thanks for your kind words, good luck on Bengali competition !!!</p>",
          "rawMarkdown": "Thanks for your kind words, good luck on Bengali competition !!!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2484516,
      "postDate": "2023-10-16T14:38:30.490Z",
      "content": "<p>Congrats to both of you. <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a>, I have a feeling you're going to be a GM soon 😉</p>",
      "rawMarkdown": "Congrats to both of you. @ahmedelfazouan, I have a feeling you're going to be a GM soon 😉",
      "votes": 1,
      "replies": [
        {
          "id": 2484537,
          "postDate": "2023-10-16T14:56:23.980Z",
          "content": "<p>😂😂 It's still a long way</p>",
          "rawMarkdown": "😂😂 It's still a long way",
          "votes": 1
        }
      ]
    },
    {
      "id": 2484272,
      "postDate": "2023-10-16T11:00:04.823Z",
      "content": "<p>Congratulations on ranking 12th on the leaderboard. Thanks for sharing the details of your solution. </p>",
      "rawMarkdown": "Congratulations on ranking 12th on the leaderboard. Thanks for sharing the details of your solution. ",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2486316,
      "author_name": "tharun_01",
      "author_url": "",
      "post_date": "2023-10-17T20:14:39.483000",
      "content": "<p>Congrats and thanks for the source code, it is a great resource to learn from top solution, as a beginner.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2485222,
      "author_name": "siwooyong",
      "author_url": "",
      "post_date": "2023-10-17T03:28:15.207000",
      "content": "<p>Congratulations, and thanks to my great teammate <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a>.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2485407,
          "author_name": "Ahmed El Fazouani",
          "author_url": "",
          "post_date": "2023-10-17T07:01:35.403000",
          "content": "<p>Thanks, It wasn't possible without your hard work !!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2484630,
      "author_name": "qdv206",
      "author_url": "",
      "post_date": "2023-10-16T15:47:12.280000",
      "content": "<p>Congrats on the gold. Great job <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> !</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2484692,
          "author_name": "Ahmed El Fazouani",
          "author_url": "",
          "post_date": "2023-10-16T16:12:13.510000",
          "content": "<p>Thanks for your kind words, good luck on Bengali competition !!!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2484516,
      "author_name": "jcerpent",
      "author_url": "",
      "post_date": "2023-10-16T14:38:30.490000",
      "content": "<p>Congrats to both of you. <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a>, I have a feeling you're going to be a GM soon 😉</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2484537,
          "author_name": "Ahmed El Fazouani",
          "author_url": "",
          "post_date": "2023-10-16T14:56:23.980000",
          "content": "<p>😂😂 It's still a long way</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2484272,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-10-16T11:00:04.823000",
      "content": "<p>Congratulations on ranking 12th on the leaderboard. Thanks for sharing the details of your solution. </p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2484242": "Thanks kaggle and the hosts for this interesting competition, \nBig thanks to kagglers out there for their great ideas and engaging descussions.\nThanks a lot as well to my great teammate @siwooyong \n\n# Summary\n\nOur solution is an ensemble of one stage approach without segmentation and two stage approach with segmentation.\n\n## One stage approach (public LB = 0.4Thanks kaggle and the hosts for this interesting competition, \nBig thanks to kagglers out there for their great ideas and engaging descussions.\nThanks a lot as well to my great teammate @siwooyong \n\n# Summary\n\nOur solution is an ensemble of one stage approach without segmentation and two stage approach with segmentation.\n\n## One stage approach (public LB = 0.5, private LB = 0.45)\n\n### Data Pre-Processing\nIf the dimension for the image is greater than (512, 512), we cropped the area with a higher density of pixels to get a (512, 512) image, then the input is resized to (96, 256, 256) for each serie following the same preprocessing steps that were used by hengck23 in his great [notebook](https://www.kaggle.com/code/hengck23/lb0-55-2-5d-3d-sample-model)\n\n### Model : resnest50d + GRU Attention\nWe tried to predict each target independently from the others so we have 13 outputs\n\n```python\nclass RSNAClassifier(nn.Module):\n    def __init__(self, model_arch, hidden_dim=128, seq_len=3, pretrained=False):\n        super().__init__()\n        self.seq_len = seq_len\n        self.model_arch = model_arch\n        self.model = timm.create_model(model_arch, in_chans=3, pretrained=pretrained)\n\n\n        cnn_feature = self.model.fc.in_features\n        self.model.global_pool = nn.Identity()\n        self.model.fc = nn.Identity()\n        self.pooling = nn.AdaptiveAvgPool2d(1)\n\n\n        self.spatialdropout = SpatialDropout(CFG.dropout)\n        self.gru = nn.GRU(cnn_feature, hidden_dim, num_layers=2, batch_first=True, bidirectional=True)\n        self.mlp_attention_layer = MLPAttentionNetwork(2 * hidden_dim)\n        self.logits = nn.Sequential(\n            nn.Linear(2 * hidden_dim, 13),\n        )\n    def forward(self, x):\n        bs = x.size(0)\n        x = x.reshape(bs*self.seq_len//3, 3, x.size(2), x.size(3))\n        features = self.model(x)\n        features = self.pooling(features).view(bs*self.seq_len//3, -1)\n        features = self.spatialdropout(features) \n        # print(features.shape)\n        features = features.reshape(bs, self.seq_len//3, -1) \n        features, _ = self.gru(features)            \n        atten_out = self.mlp_attention_layer(features) \n        pred = self.logits(atten_out)\n        pred = pred.view(bs, -1)\n        return pred\n```\n\n### Augmentation\n * Mixup\n * Random crop + resize\n * Random shift, scale, rotate\n * shuffle randomly the indexes of the sequence, but respecting the same order and keeping the dependency between each three consecutive images:\n```python\n\tinds = np.random.choice(np.arange(1, 96-1), 32, replace = False)\n\tinds.sort()\n\tinds = np.stack([inds-1, inds, inds+1]).T.flatten()\n\timage = image[inds]\n\n```\nLoss : BCEWithLogitsLoss\nscheduler : CosineAnnealingLR\noptimizer : AdamW\nlearning rate : 5e-5\n\n### Postprocessing \nWe simply multiplied the output by the weights of the competition metric :\n```python\npreds.loc[:, ['bowel_injury', 'kidney_low', 'liver_low', 'spleen_low']] *= 2\npreds.loc[:, ['kidney_high', 'liver_high', 'spleen_high']] *= 4\npreds.loc[:, ['extravasation_injury']] *= 6\n```\n\n## Two stage approach (public LB = 0.45, private LB = 0.43)\n### stage1 : Segmentation\n\nModel : regnety002 + unet\n\nEven with only 160 of 200 data (1th fold) used as training data, the model has already shown good performance.\n\n```python\nclass SegModel(nn.Module):\n    def __init__(self):\n        super(SegModel, self).__init__()\n   \n        self.n_classes = len(\n            [\n                'background',\n                'liver',\n                'spleen',\n                'left kidney',\n                'right kidney',\n                'bowel'\n            ])\n        in_chans = 1\n        self.encoder = timm.create_model(\n            'regnety_002',\n            pretrained=False,\n            features_only=True,\n            in_chans=in_chans,\n        )\n        encoder_channels = tuple(\n            [in_chans]\n            + [\n                self.encoder.feature_info[i][\"num_chs\"]\n                for i in range(len(self.encoder.feature_info))\n            ]\n        )\n        self.decoder = UnetDecoder(\n            encoder_channels=encoder_channels,\n            decoder_channels=(256, 128, 64, 32, 16),\n            n_blocks=5,\n            use_batchnorm=True,\n            center=False,\n            attention_type=None,\n        )\n\n        self.segmentation_head = SegmentationHead(\n            in_channels=16,\n            out_channels=self.n_classes,\n            activation=None,\n            kernel_size=3,\n        )\n\n        self.bce_seg = nn.BCEWithLogitsLoss()\n\n    def forward(self, x_in):\n        enc_out = self.encoder(x_in)\n\n        decoder_out = self.decoder(*[x_in] + enc_out)\n        x_seg = self.segmentation_head(decoder_out)\n\n        return nn.Sigmoid()(x_seg)\n```\n\n### stage2 : 2.5DCNN \n\n#### Data Pre-Processing:\nWe used the segmentation logits obtained from stage1 to crop livers, spleen, and kidney, and then resized each to (96, 224, 224). \n(We use 10-size padding when we crop the organs with segmentation logits)\nIn addition, full ct data not cropped is resized to (128, 224, 224) and a total of four inputs are put into the model (full_video, crop_liver, crop_spleen, crop_kidney)\n\n#### Model : regnety002 + transformer\n\nWe initially used a custom any_injury_loss function, but found that it did not improve the performance.  For the model input channel, we experimented with different values, including 2, 3, 4, and 8. \nWe found that a channel size of 2 performed the best, we also initially tried using a shared CNN and transformer model for all organs,  but found that separate CNN and transformer models for each organ performed better.  we also experimented with increasing the size of the CNN (using ConvNeXt and EfficientNet models), but this resulted in a decrease in performance.  Therefore, we used the RegNet002 model, which is a smaller CNN model.\n\n```python\nclass FeatureExtractor(nn.Module):\n    def __init__(self, hidden, num_channel):\n        super(FeatureExtractor, self).__init__()\n\n        self.hidden = hidden\n        self.num_channel = num_channel\n\n        self.cnn = timm.create_model(model_name = 'regnety_002',\n                                     pretrained = True,\n                                     num_classes = 0,\n                                     in_chans = num_channel)\n\n        self.fc = nn.Linear(hidden, hidden//2)\n\n    def forward(self, x):\n        batch_size, num_frame, h, w = x.shape\n        x = x.reshape(batch_size, num_frame//self.num_channel, self.num_channel, h, w)\n        x = x.reshape(-1, self.num_channel, h, w)\n        x = self.cnn(x)\n        x = x.reshape(batch_size, num_frame//self.num_channel, self.hidden)\n\n        x = self.fc(x)\n        return x\n\nclass ContextProcessor(nn.Module):\n    def __init__(self, hidden):\n        super(ContextProcessor, self).__init__()\n        self.transformer = RobertaPreLayerNormModel(\n            RobertaPreLayerNormConfig(\n                hidden_size = hidden//2,\n                num_hidden_layers = 1,\n                num_attention_heads = 4,\n                intermediate_size = hidden*2,\n                hidden_act = 'gelu_new',\n                )\n            )\n\n        del self.transformer.embeddings.word_embeddings\n\n        self.dense = nn.Linear(hidden, hidden)\n        self.activation = nn.ReLU()\n\n\n    def forward(self, x):\n        x = self.transformer(inputs_embeds = x).last_hidden_state\n\n        apool = torch.mean(x, dim = 1)\n        mpool, _ = torch.max(x, dim = 1)\n        x = torch.cat([mpool, apool], dim = -1)\n\n        x = self.dense(x)\n        x = self.activation(x)\n        return x\n\nclass Custom3DCNN(nn.Module):\n    def __init__(self, hidden = 368, num_channel = 2):\n        super(Custom3DCNN, self).__init__()\n\n        self.full_extractor = FeatureExtractor(hidden=hidden, num_channel=num_channel)\n        self.kidney_extractor = FeatureExtractor(hidden=hidden, num_channel=num_channel)\n        self.liver_extractor = FeatureExtractor(hidden=hidden, num_channel=num_channel)\n        self.spleen_extractor = FeatureExtractor(hidden=hidden, num_channel=num_channel)\n\n        self.full_processor = ContextProcessor(hidden=hidden)\n        self.kidney_processor = ContextProcessor(hidden=hidden)\n        self.liver_processor = ContextProcessor(hidden=hidden)\n        self.spleen_processor = ContextProcessor(hidden=hidden)\n\n        self.bowel = nn.Linear(hidden, 2)\n        self.extravasation = nn.Linear(hidden, 2)\n        self.kidney = nn.Linear(hidden, 3)\n        self.liver = nn.Linear(hidden, 3)\n        self.spleen = nn.Linear(hidden, 3)\n\n        self.softmax = nn.Softmax(dim = -1)\n\n    def forward(self, full_input, crop_liver, crop_spleen, crop_kidney, mask, mode):\n        full_output = self.full_extractor(full_input)\n        kidney_output = self.kidney_extractor(crop_kidney)\n        liver_output = self.liver_extractor(crop_liver)\n        spleen_output = self.spleen_extractor(crop_spleen)\n\n        full_output2 = self.full_processor(torch.cat([full_output, kidney_output, liver_output, spleen_output], dim = 1))\n        kidney_output2 = self.kidney_processor(torch.cat([full_output, kidney_output], dim = 1))\n        liver_output2 = self.liver_processor(torch.cat([full_output, liver_output], dim = 1))\n        spleen_output2 = self.spleen_processor(torch.cat([full_output, spleen_output], dim = 1))\n\n        bowel = self.bowel(full_output2)\n        extravasation = self.extravasation(full_output2)\n        kidney = self.kidney(kidney_output2)\n        liver = self.liver(liver_output2)\n        spleen = self.spleen(spleen_output2)\n\n\n        any_injury = torch.stack([\n            self.softmax(bowel)[:, 0],\n            self.softmax(extravasation)[:, 0],\n            self.softmax(kidney)[:, 0],\n            self.softmax(liver)[:, 0],\n            self.softmax(spleen)[:, 0]\n        ], dim = -1)\n        any_injury = 1 - any_injury\n        any_injury, _ = any_injury.max(1)\n        return bowel, extravasation, kidney, liver, spleen, any_injury\n\n```\n\n#### Augmentation\n\n```python\nclass CustomAug(nn.Module):\n    def __init__(self, prob = 0.5, s = 224):\n        super(CustomAug, self).__init__()\n        self.prob = prob\n\n        self.do_random_rotate = v2.RandomRotation(\n            degrees = (-45, 45),\n            interpolation = torchvision.transforms.InterpolationMode.BILINEAR,\n            expand = False,\n            center = None,\n            fill = 0\n        )\n        self.do_random_scale = v2.ScaleJitter(\n            target_size = [s, s],\n            scale_range = (0.8, 1.2),\n            interpolation = torchvision.transforms.InterpolationMode.BILINEAR,\n            antialias = True)\n\n        self.do_random_crop = v2.RandomCrop(\n            size = [s, s],\n            #padding = None,\n            pad_if_needed = True,\n            fill = 0,\n            padding_mode = 'constant'\n        )\n\n        self.do_horizontal_flip = v2.RandomHorizontalFlip(self.prob)\n        self.do_vertical_flip = v2.RandomVerticalFlip(self.prob)\n    def forward(self, x):\n        if np.random.rand() < self.prob:\n            x = self.do_random_rotate(x)\n\n        if np.random.rand() < self.prob:\n            x = self.do_random_scale(x)\n            x = self.do_random_crop(x)\n\n        x = self.do_horizontal_flip(x)\n        x = self.do_vertical_flip(x)\n        return x\n```\n\nLoss : nn.CrossEntropyLoss(no class weight)\nscheduler : cosine_schedule_with_warmup\noptimizer : AdamW\nlearning rate :2e-4\n\n#### Postprocessing\nWe multiplied by the value that maximizes the validation score for each pred_df obtained for each fold.\n\n```python\nweights = [\n    [0.9, 4, 2, 4, 2, 6, 6, 6],\n    [0.9, 1, 4, 3, 2, 5, 5, 6],\n    [0.2, 3, 2, 1, 2, 4, 2, 6],\n    [0.5, 2, 2, 2, 2, 2, 6, 6],\n    [1, 2, 3, 2, 6, 3, 6, 5]\n]\n\ny_pred = pred_df.copy().groupby('patient_id').mean().reset_index()\n\nw1, w2, w3, w4, w5, w6, w7, w8 = weights[i]\n\ny_pred['bowel_injury'] *= w1\ny_pred['kidney_low'] *= w2\ny_pred['liver_low'] *= w3\ny_pred['spleen_low'] *= w4\ny_pred['kidney_high'] *= w5\ny_pred['liver_high'] *= w6\ny_pred['spleen_high'] *= w7\ny_pred['extravasation_injury'] *= w8\n\ny_pred = y_pred ** 0.8\n\n```\n\n### Reference\n[RSNA 2022 1st place solution](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362607)\n[RSNA 2022 2nd place solution](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/365115)\n[RSNA 2022 5th place solution](https://github.com/pascal-pfeiffer/kaggle-rsna-2022-5th-place)\n@hengck23's [descussion](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/435053)\n\n## Code :\n https://github.com/siwooyong/RSNA-2023-Abdominal-Trauma-Detection\nInference notebook : https://www.kaggle.com/code/ahmedelfazouan/rsna-atd-channel2-512-inference-ensemble?scriptVersionId=146616538",
    "2486316": "Congrats and thanks for the source code, it is a great resource to learn from top solution, as a beginner.",
    "2485222": "Congratulations, and thanks to my great teammate @ahmedelfazouan.",
    "2484630": "Congrats on the gold. Great job @ahmedelfazouan !",
    "2484516": "Congrats to both of you. @ahmedelfazouan, I have a feeling you're going to be a GM soon 😉",
    "2484272": "Congratulations on ranking 12th on the leaderboard. Thanks for sharing the details of your solution. "
  }
}