{
  "id": 352940,
  "title": "Split network in PyTorch",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/352940",
  "author_name": "Samuel Cortinhas",
  "post_date": "2022-09-16T10:01:03.760000",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I wasn't sure how to word this in google so I thought I'd draw a picture and ask here. </p>\n<p>\n<img src=\"https://i.postimg.cc/jSprwjN1/23234.jpg\">\n</p>\n<p>Is there a way to implement a network with this kind of structure in PyTorch? I'd appreciate it if someone could share some beginner friendly resources, thanks.</p>",
  "messages": [
    {
      "id": 1942008,
      "postDate": "2022-09-16T11:46:04.243Z",
      "content": "<p>An <code>nn.Module</code> in pytorch can take any number of inputs and return anything as output.</p>\n<p>Here is a pseudo code example of your drawing which takes a tuple with img and tabular data as input and simply averages the two outputs:</p>\n<pre><code>class HybridAvgNetwork(nn.Module):\n    def __init__(self, cfg):\n        super(HybridAvgNetwork, self).__init__()\n        self.image_classifier = define_cnn(n_output=cfg.n_output)\n        self.tabular_classifier = defince_tabular(n_output=cfg.n_output)\n\n    def forward(self, data):\n        img_data, tabular_data = data\n        preds_cnn = self.image_classfier(img_data)\n        preds_tabular = self.tabular_classifier(tabular_data)\n        output = (preds_cnn + preds_tabular) / 2\n        return output  \n</code></pre>\n<p>Averaging the outputs is a very basic solution and you could decide to do something different using a linear layer to work from embeddings like this:</p>\n<pre><code>class HybridEmbedNetwork(nn.Module):\n    def __init__(self, cfg):\n        super(HybridEmbedNetwork, self).__init__()\n        self.image_classifier = define_cnn(n_output=cfg.embed_dim)\n        self.tabular_classifier = defince_tabular(n_output=cfg.embed_dim)\n        self.linear_layer = Linear(2*cfg.embed_dim, cfg.n_output)\n    def forward(self, data):\n        img_data, tabular_data = data\n        preds_cnn = self.image_classfier(img_data)\n        preds_tabular = self.tabular_classifier(tabular_data)\n        embeddings = torch.cat([preds_cnn, preds_tabular])\n        output = self.linear_layer(embeddings)\n        return output  \n</code></pre>\n<p>Those are just examples and you can really try any architecture that does make sense to combine the different inputs in an end to end manner.</p>",
      "rawMarkdown": "An `nn.Module` in pytorch can take any number of inputs and return anything as output.\n\nHere is a pseudo code example of your drawing which takes a tuple with img and tabular data as input and simply averages the two outputs:\n\n```\nclass HybridAvgNetwork(nn.Module):\n    def __init__(self, cfg):\n        super(HybridAvgNetwork, self).__init__()\n        self.image_classifier = define_cnn(n_output=cfg.n_output)\n        self.tabular_classifier = defince_tabular(n_output=cfg.n_output)\n\n    def forward(self, data):\n        img_data, tabular_data = data\n        preds_cnn = self.image_classfier(img_data)\n        preds_tabular = self.tabular_classifier(tabular_data)\n        output = (preds_cnn + preds_tabular) / 2\n        return output  \n```\n\nAveraging the outputs is a very basic solution and you could decide to do something different using a linear layer to work from embeddings like this:\n\n```\nclass HybridEmbedNetwork(nn.Module):\n    def __init__(self, cfg):\n        super(HybridEmbedNetwork, self).__init__()\n        self.image_classifier = define_cnn(n_output=cfg.embed_dim)\n        self.tabular_classifier = defince_tabular(n_output=cfg.embed_dim)\n        self.linear_layer = Linear(2*cfg.embed_dim, cfg.n_output)\n    def forward(self, data):\n        img_data, tabular_data = data\n        preds_cnn = self.image_classfier(img_data)\n        preds_tabular = self.tabular_classifier(tabular_data)\n        embeddings = torch.cat([preds_cnn, preds_tabular])\n        output = self.linear_layer(embeddings)\n        return output  \n```\n\nThose are just examples and you can really try any architecture that does make sense to combine the different inputs in an end to end manner.",
      "votes": 5,
      "replies": [
        {
          "id": 1942089,
          "postDate": "2022-09-16T12:35:56.957Z",
          "content": "<p>Thank you, this helps a lot!</p>",
          "rawMarkdown": "Thank you, this helps a lot!"
        }
      ]
    },
    {
      "id": 1941828,
      "postDate": "2022-09-16T10:01:03.760Z",
      "content": "<p>I wasn't sure how to word this in google so I thought I'd draw a picture and ask here. </p>\n<p>\n<img src=\"https://i.postimg.cc/jSprwjN1/23234.jpg\">\n</p>\n<p>Is there a way to implement a network with this kind of structure in PyTorch? I'd appreciate it if someone could share some beginner friendly resources, thanks.</p>",
      "rawMarkdown": "I wasn't sure how to word this in google so I thought I'd draw a picture and ask here. \n\n<center>\n<img src='https://i.postimg.cc/jSprwjN1/23234.jpg' width=700>\n</center>\n\nIs there a way to implement a network with this kind of structure in PyTorch? I'd appreciate it if someone could share some beginner friendly resources, thanks.",
      "votes": 3
    },
    {
      "id": 1941887,
      "postDate": "2022-09-16T10:19:23.383Z",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/175412\" target=\"_blank\">https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/175412</a></p>",
      "rawMarkdown": "https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/175412",
      "votes": 2,
      "replies": [
        {
          "id": 1942044,
          "postDate": "2022-09-16T12:10:40.940Z",
          "content": "<p>thanks for mentioning our previous solution :)</p>",
          "rawMarkdown": "thanks for mentioning our previous solution :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1941974,
      "postDate": "2022-09-16T11:16:05.797Z",
      "content": "<p>under transformer framework, every input is just a token</p>",
      "rawMarkdown": "under transformer framework, every input is just a token",
      "votes": -1
    }
  ],
  "comments": [
    {
      "id": 1942008,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2022-09-16T11:46:04.243000",
      "content": "<p>An <code>nn.Module</code> in pytorch can take any number of inputs and return anything as output.</p>\n<p>Here is a pseudo code example of your drawing which takes a tuple with img and tabular data as input and simply averages the two outputs:</p>\n<pre><code>class HybridAvgNetwork(nn.Module):\n    def __init__(self, cfg):\n        super(HybridAvgNetwork, self).__init__()\n        self.image_classifier = define_cnn(n_output=cfg.n_output)\n        self.tabular_classifier = defince_tabular(n_output=cfg.n_output)\n\n    def forward(self, data):\n        img_data, tabular_data = data\n        preds_cnn = self.image_classfier(img_data)\n        preds_tabular = self.tabular_classifier(tabular_data)\n        output = (preds_cnn + preds_tabular) / 2\n        return output  \n</code></pre>\n<p>Averaging the outputs is a very basic solution and you could decide to do something different using a linear layer to work from embeddings like this:</p>\n<pre><code>class HybridEmbedNetwork(nn.Module):\n    def __init__(self, cfg):\n        super(HybridEmbedNetwork, self).__init__()\n        self.image_classifier = define_cnn(n_output=cfg.embed_dim)\n        self.tabular_classifier = defince_tabular(n_output=cfg.embed_dim)\n        self.linear_layer = Linear(2*cfg.embed_dim, cfg.n_output)\n    def forward(self, data):\n        img_data, tabular_data = data\n        preds_cnn = self.image_classfier(img_data)\n        preds_tabular = self.tabular_classifier(tabular_data)\n        embeddings = torch.cat([preds_cnn, preds_tabular])\n        output = self.linear_layer(embeddings)\n        return output  \n</code></pre>\n<p>Those are just examples and you can really try any architecture that does make sense to combine the different inputs in an end to end manner.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1942089,
          "author_name": "Samuel Cortinhas",
          "author_url": "",
          "post_date": "2022-09-16T12:35:56.957000",
          "content": "<p>Thank you, this helps a lot!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1941887,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2022-09-16T10:19:23.383000",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/175412\" target=\"_blank\">https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/175412</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 1942044,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2022-09-16T12:10:40.940000",
          "content": "<p>thanks for mentioning our previous solution :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1941974,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-09-16T11:16:05.797000",
      "content": "<p>under transformer framework, every input is just a token</p>",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1942008": "An `nn.Module` in pytorch can take any number of inputs and return anything as output.\n\nHere is a pseudo code example of your drawing which takes a tuple with img and tabular data as input and simply averages the two outputs:\n\n```\nclass HybridAvgNetwork(nn.Module):\n    def __init__(self, cfg):\n        super(HybridAvgNetwork, self).__init__()\n        self.image_classifier = define_cnn(n_output=cfg.n_output)\n        self.tabular_classifier = defince_tabular(n_output=cfg.n_output)\n\n    def forward(self, data):\n        img_data, tabular_data = data\n        preds_cnn = self.image_classfier(img_data)\n        preds_tabular = self.tabular_classifier(tabular_data)\n        output = (preds_cnn + preds_tabular) / 2\n        return output  \n```\n\nAveraging the outputs is a very basic solution and you could decide to do something different using a linear layer to work from embeddings like this:\n\n```\nclass HybridEmbedNetwork(nn.Module):\n    def __init__(self, cfg):\n        super(HybridEmbedNetwork, self).__init__()\n        self.image_classifier = define_cnn(n_output=cfg.embed_dim)\n        self.tabular_classifier = defince_tabular(n_output=cfg.embed_dim)\n        self.linear_layer = Linear(2*cfg.embed_dim, cfg.n_output)\n    def forward(self, data):\n        img_data, tabular_data = data\n        preds_cnn = self.image_classfier(img_data)\n        preds_tabular = self.tabular_classifier(tabular_data)\n        embeddings = torch.cat([preds_cnn, preds_tabular])\n        output = self.linear_layer(embeddings)\n        return output  \n```\n\nThose are just examples and you can really try any architecture that does make sense to combine the different inputs in an end to end manner.",
    "1941828": "I wasn't sure how to word this in google so I thought I'd draw a picture and ask here. \n\n<center>\n<img src='https://i.postimg.cc/jSprwjN1/23234.jpg' width=700>\n</center>\n\nIs there a way to implement a network with this kind of structure in PyTorch? I'd appreciate it if someone could share some beginner friendly resources, thanks.",
    "1941887": "https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/175412",
    "1941974": "under transformer framework, every input is just a token"
  }
}