{
  "id": 444299,
  "title": "3D Decoder Architectures",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/444299",
  "author_name": "Mark",
  "post_date": "2023-10-01T10:15:16.137000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Looking to get some experienced opinions on choosing/experimenting with decoder architecture in this competition. </p>\n<p>I have gone for an approach similar to much of the code I have seen shared in this competition - feature extraction for stack of 2D images followed by some decoder + head for each organ. I haven't had a great success in training a model yet that beats weighted mean baseline though. </p>\n<p>Here is a proposed architecture. I'd love to hear any thoughts on why this might be a bad architecture and what improvements to the model could/should be made. I've tried using pretrained and from scratch for the encoder too.</p>\n<p>```python<br>\n        self.in_channels=in_channels<br>\n        self.depth=depth<br>\n        self.image_size=image_size</p>\n<pre><code>    self.encoder = timm.create_model(\n        backbone,\n        in_chans=,\n        num_classes=,\n        features_only=,\n        drop_rate=drop_rate,\n        drop_path_rate=drop_path_rate,\n        pretrained=pretrained\n    )\n\n       backbone:\n        hdim = self.encoder.conv_head.out_channels\n       backbone:\n        hdim = self.encoder.head.fc.in_features\n       backbone:\n        hdim = \n\n     freeze_backbone:\n         param  self.encoder.parameters():\n            param.requires_grad = \n\n    \n    self.middle1 = nn.Sequential(\n        nn.Conv3d(hdim, , kernel_size=(,,),stride=(,,), padding=(,,)),\n        nn.BatchNorm3d(),\n        nn.ReLU(inplace=),\n    )\n    self.middle2 = nn.Sequential(\n        nn.Conv3d(, , kernel_size=(,,),stride=(,,), padding=(,,)),\n        nn.BatchNorm3d(),\n        nn.ReLU(inplace=),\n        nn.MaxPool3d(kernel_size=, stride=)\n    )\n\n    \n    self.decoder1 = nn.Sequential(\n        nn.ConvTranspose3d(, , kernel_size=, stride=),\n        nn.BatchNorm3d(),\n        nn.ReLU(inplace=),\n    )\n    self.decoder2 = nn.Sequential(\n        nn.Conv3d(, , kernel_size=, padding=),\n        nn.BatchNorm3d(),\n        nn.ReLU(inplace=),\n        nn.ConvTranspose3d(, , kernel_size=, stride=)\n    )\n    self.lhead = nn.Sequential(\n        nn.Linear(, ),\n        nn.BatchNorm1d(),\n        nn.Dropout(drop_rate_last),\n        nn.LeakyReLU(),\n        nn.Linear(,),\n    )\n\n ():  \n\n    batch_size, D, H, W = x.shape\n    x = x.view(batch_size*D//self.in_channels, self.in_channels, H, W )\n    feat = self.encoder.forward_features(x)\n    _,d,h,w = feat.shape\n    f1 = feat.reshape(batch_size, D // self.in_channels, d, h, w)  \n    f1 = f1.permute((, , , , )).contiguous()  \n    x = self.middle1(f1)\n    x = self.middle2(x)\n    x = self.decoder1(x)\n    x = self.decoder2(x)\n    f3 = nn.functional.adaptive_avg_pool3d(x, )  \n    pool2 = f3.reshape(batch_size, -) \n\n    liver = self.lhead(pool2)\n\n     liver\n</code></pre>\n<p>`</p>",
  "messages": [
    {
      "id": 2463418,
      "postDate": "2023-10-01T10:15:16.137Z",
      "content": "<p>Looking to get some experienced opinions on choosing/experimenting with decoder architecture in this competition. </p>\n<p>I have gone for an approach similar to much of the code I have seen shared in this competition - feature extraction for stack of 2D images followed by some decoder + head for each organ. I haven't had a great success in training a model yet that beats weighted mean baseline though. </p>\n<p>Here is a proposed architecture. I'd love to hear any thoughts on why this might be a bad architecture and what improvements to the model could/should be made. I've tried using pretrained and from scratch for the encoder too.</p>\n<p>```python<br>\n        self.in_channels=in_channels<br>\n        self.depth=depth<br>\n        self.image_size=image_size</p>\n<pre><code>    self.encoder = timm.create_model(\n        backbone,\n        in_chans=,\n        num_classes=,\n        features_only=,\n        drop_rate=drop_rate,\n        drop_path_rate=drop_path_rate,\n        pretrained=pretrained\n    )\n\n       backbone:\n        hdim = self.encoder.conv_head.out_channels\n       backbone:\n        hdim = self.encoder.head.fc.in_features\n       backbone:\n        hdim = \n\n     freeze_backbone:\n         param  self.encoder.parameters():\n            param.requires_grad = \n\n    \n    self.middle1 = nn.Sequential(\n        nn.Conv3d(hdim, , kernel_size=(,,),stride=(,,), padding=(,,)),\n        nn.BatchNorm3d(),\n        nn.ReLU(inplace=),\n    )\n    self.middle2 = nn.Sequential(\n        nn.Conv3d(, , kernel_size=(,,),stride=(,,), padding=(,,)),\n        nn.BatchNorm3d(),\n        nn.ReLU(inplace=),\n        nn.MaxPool3d(kernel_size=, stride=)\n    )\n\n    \n    self.decoder1 = nn.Sequential(\n        nn.ConvTranspose3d(, , kernel_size=, stride=),\n        nn.BatchNorm3d(),\n        nn.ReLU(inplace=),\n    )\n    self.decoder2 = nn.Sequential(\n        nn.Conv3d(, , kernel_size=, padding=),\n        nn.BatchNorm3d(),\n        nn.ReLU(inplace=),\n        nn.ConvTranspose3d(, , kernel_size=, stride=)\n    )\n    self.lhead = nn.Sequential(\n        nn.Linear(, ),\n        nn.BatchNorm1d(),\n        nn.Dropout(drop_rate_last),\n        nn.LeakyReLU(),\n        nn.Linear(,),\n    )\n\n ():  \n\n    batch_size, D, H, W = x.shape\n    x = x.view(batch_size*D//self.in_channels, self.in_channels, H, W )\n    feat = self.encoder.forward_features(x)\n    _,d,h,w = feat.shape\n    f1 = feat.reshape(batch_size, D // self.in_channels, d, h, w)  \n    f1 = f1.permute((, , , , )).contiguous()  \n    x = self.middle1(f1)\n    x = self.middle2(x)\n    x = self.decoder1(x)\n    x = self.decoder2(x)\n    f3 = nn.functional.adaptive_avg_pool3d(x, )  \n    pool2 = f3.reshape(batch_size, -) \n\n    liver = self.lhead(pool2)\n\n     liver\n</code></pre>\n<p>`</p>",
      "rawMarkdown": "Looking to get some experienced opinions on choosing/experimenting with decoder architecture in this competition. \n\nI have gone for an approach similar to much of the code I have seen shared in this competition - feature extraction for stack of 2D images followed by some decoder + head for each organ. I haven't had a great success in training a model yet that beats weighted mean baseline though. \n\nHere is a proposed architecture. I'd love to hear any thoughts on why this might be a bad architecture and what improvements to the model could/should be made. I've tried using pretrained and from scratch for the encoder too.\n\n```python\n        self.in_channels=in_channels\n        self.depth=depth\n        self.image_size=image_size\n\n        self.encoder = timm.create_model(\n            backbone,\n            in_chans=1,\n            num_classes=1,\n            features_only=False,\n            drop_rate=drop_rate,\n            drop_path_rate=drop_path_rate,\n            pretrained=pretrained\n        )\n\n        if 'efficient' in backbone:\n            hdim = self.encoder.conv_head.out_channels\n        elif 'convnext' in backbone:\n            hdim = self.encoder.head.fc.in_features\n        elif 'resnet' in backbone:\n            hdim = 512\n\n        if freeze_backbone:\n            for param in self.encoder.parameters():\n                param.requires_grad = False\n        \n        # Middle layers\n        self.middle1 = nn.Sequential(\n            nn.Conv3d(hdim, 256, kernel_size=(3,3,3),stride=(2,1,1), padding=(1,1,1)),\n            nn.BatchNorm3d(256),\n            nn.ReLU(inplace=True),\n        )\n        self.middle2 = nn.Sequential(\n            nn.Conv3d(256, 256, kernel_size=(3,3,3),stride=(2,1,1), padding=(1,1,1)),\n            nn.BatchNorm3d(256),\n            nn.ReLU(inplace=True),\n            nn.MaxPool3d(kernel_size=2, stride=2)\n        )\n\n        # Decoder layers\n        self.decoder1 = nn.Sequential(\n            nn.ConvTranspose3d(256, 64, kernel_size=2, stride=2),\n            nn.BatchNorm3d(64),\n            nn.ReLU(inplace=True),\n        )\n        self.decoder2 = nn.Sequential(\n            nn.Conv3d(64, 64, kernel_size=3, padding=1),\n            nn.BatchNorm3d(64),\n            nn.ReLU(inplace=True),\n            nn.ConvTranspose3d(64, 32, kernel_size=2, stride=2)\n        )\n        self.lhead = nn.Sequential(\n            nn.Linear(32, 32),\n            nn.BatchNorm1d(32),\n            nn.Dropout(drop_rate_last),\n            nn.LeakyReLU(0.1),\n            nn.Linear(32,3),\n        )\n\n    def forward(self, x):  # (bs, ch, sz, sz)\n\n        batch_size, D, H, W = x.shape\n        x = x.view(batch_size*D//self.in_channels, self.in_channels, H, W )\n        feat = self.encoder.forward_features(x)\n        _,d,h,w = feat.shape\n        f1 = feat.reshape(batch_size, D // self.in_channels, d, h, w)  # 8, 24, 512, 8, 8\n        f1 = f1.permute((0, 2, 1, 3, 4)).contiguous()  # 8, 512, 96, 8, 8\n        x = self.middle1(f1)\n        x = self.middle2(x)\n        x = self.decoder1(x)\n        x = self.decoder2(x)\n        f3 = nn.functional.adaptive_avg_pool3d(x, 1)  # 8, 64, 1, 1, 1\n        pool2 = f3.reshape(batch_size, -1) # 8, 128, 1, 1, 1\n\n        liver = self.lhead(pool2)\n        \n        return liver\n`",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2463418": "Looking to get some experienced opinions on choosing/experimenting with decoder architecture in this competition. \n\nI have gone for an approach similar to much of the code I have seen shared in this competition - feature extraction for stack of 2D images followed by some decoder + head for each organ. I haven't had a great success in training a model yet that beats weighted mean baseline though. \n\nHere is a proposed architecture. I'd love to hear any thoughts on why this might be a bad architecture and what improvements to the model could/should be made. I've tried using pretrained and from scratch for the encoder too.\n\n```python\n        self.in_channels=in_channels\n        self.depth=depth\n        self.image_size=image_size\n\n        self.encoder = timm.create_model(\n            backbone,\n            in_chans=1,\n            num_classes=1,\n            features_only=False,\n            drop_rate=drop_rate,\n            drop_path_rate=drop_path_rate,\n            pretrained=pretrained\n        )\n\n        if 'efficient' in backbone:\n            hdim = self.encoder.conv_head.out_channels\n        elif 'convnext' in backbone:\n            hdim = self.encoder.head.fc.in_features\n        elif 'resnet' in backbone:\n            hdim = 512\n\n        if freeze_backbone:\n            for param in self.encoder.parameters():\n                param.requires_grad = False\n        \n        # Middle layers\n        self.middle1 = nn.Sequential(\n            nn.Conv3d(hdim, 256, kernel_size=(3,3,3),stride=(2,1,1), padding=(1,1,1)),\n            nn.BatchNorm3d(256),\n            nn.ReLU(inplace=True),\n        )\n        self.middle2 = nn.Sequential(\n            nn.Conv3d(256, 256, kernel_size=(3,3,3),stride=(2,1,1), padding=(1,1,1)),\n            nn.BatchNorm3d(256),\n            nn.ReLU(inplace=True),\n            nn.MaxPool3d(kernel_size=2, stride=2)\n        )\n\n        # Decoder layers\n        self.decoder1 = nn.Sequential(\n            nn.ConvTranspose3d(256, 64, kernel_size=2, stride=2),\n            nn.BatchNorm3d(64),\n            nn.ReLU(inplace=True),\n        )\n        self.decoder2 = nn.Sequential(\n            nn.Conv3d(64, 64, kernel_size=3, padding=1),\n            nn.BatchNorm3d(64),\n            nn.ReLU(inplace=True),\n            nn.ConvTranspose3d(64, 32, kernel_size=2, stride=2)\n        )\n        self.lhead = nn.Sequential(\n            nn.Linear(32, 32),\n            nn.BatchNorm1d(32),\n            nn.Dropout(drop_rate_last),\n            nn.LeakyReLU(0.1),\n            nn.Linear(32,3),\n        )\n\n    def forward(self, x):  # (bs, ch, sz, sz)\n\n        batch_size, D, H, W = x.shape\n        x = x.view(batch_size*D//self.in_channels, self.in_channels, H, W )\n        feat = self.encoder.forward_features(x)\n        _,d,h,w = feat.shape\n        f1 = feat.reshape(batch_size, D // self.in_channels, d, h, w)  # 8, 24, 512, 8, 8\n        f1 = f1.permute((0, 2, 1, 3, 4)).contiguous()  # 8, 512, 96, 8, 8\n        x = self.middle1(f1)\n        x = self.middle2(x)\n        x = self.decoder1(x)\n        x = self.decoder2(x)\n        f3 = nn.functional.adaptive_avg_pool3d(x, 1)  # 8, 64, 1, 1, 1\n        pool2 = f3.reshape(batch_size, -1) # 8, 128, 1, 1, 1\n\n        liver = self.lhead(pool2)\n        \n        return liver\n`"
  }
}