{
  "id": 378769,
  "title": "Can models be ensembled using logits?",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/378769",
  "author_name": "moth",
  "post_date": "2023-01-17T01:13:23.359000",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I was wondering if models could be ensembled using logits. </p>\n<p>I have seen models being ensembled using model probabilities, like averaging output probabilities of two models. I've also seen ensembles which concatenate each individual model backbone's final layer and then use a classification head (though idk if this is wrong if the models were not trained in this way).</p>\n<p>Is there a way to ensemble models by summing/averaging logits? Is this a good approach?</p>",
  "messages": [
    {
      "id": 2103055,
      "postDate": "2023-01-17T01:13:23.360Z",
      "content": "<p>I was wondering if models could be ensembled using logits. </p>\n<p>I have seen models being ensembled using model probabilities, like averaging output probabilities of two models. I've also seen ensembles which concatenate each individual model backbone's final layer and then use a classification head (though idk if this is wrong if the models were not trained in this way).</p>\n<p>Is there a way to ensemble models by summing/averaging logits? Is this a good approach?</p>",
      "rawMarkdown": "I was wondering if models could be ensembled using logits. \n\nI have seen models being ensembled using model probabilities, like averaging output probabilities of two models. I've also seen ensembles which concatenate each individual model backbone's final layer and then use a classification head (though idk if this is wrong if the models were not trained in this way).\n\nIs there a way to ensemble models by summing/averaging logits? Is this a good approach?",
      "votes": 4
    },
    {
      "id": 2111427,
      "postDate": "2023-01-22T22:20:44.380Z",
      "content": "<p>Well, you can do it this way or take the sigmoid/softmax before you take the average. Try both approaches and see which one gives better results (on validation).</p>\n<p>I agree with <a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a>, the magnitudes of each models' logits might be quite different so having the probabilities is safer before averaging.</p>\n<p>Good luck!</p>",
      "rawMarkdown": "Well, you can do it this way or take the sigmoid/softmax before you take the average. Try both approaches and see which one gives better results (on validation).\n\nI agree with @yosukeyama, the magnitudes of each models' logits might be quite different so having the probabilities is safer before averaging.\n\nGood luck!",
      "votes": 1
    },
    {
      "id": 2103342,
      "postDate": "2023-01-17T06:33:39.593Z",
      "content": "<p>yes it can be done. The best practice is always check 😊</p>",
      "rawMarkdown": "yes it can be done. The best practice is always check 😊",
      "votes": -2,
      "replies": [
        {
          "id": 2103894,
          "postDate": "2023-01-17T12:50:19.880Z",
          "content": "<p>I know it can be done, but I was asking <em>how</em> it's done</p>",
          "rawMarkdown": "I know it can be done, but I was asking *how* it's done",
          "replies": [
            {
              "id": 2103929,
              "postDate": "2023-01-17T13:21:28.633Z",
              "content": "<p>assuming that your model output are logits …</p>\n<pre><code>model1 = resnet()\nmodel2 = resnet()\n\nlogits1 = model1(image)\nlogits2 = model2(image)\n\naverage_logits = torch.mean(torch.stack([logits1, logits2]), dim=0)\n\nprediction = torch.softmax(average_logits, dim=1)  # or sigmoid -&gt; depends on task\n</code></pre>",
              "rawMarkdown": "assuming that your model output are logits ...\n\n```\nmodel1 = resnet()\nmodel2 = resnet()\n\nlogits1 = model1(image)\nlogits2 = model2(image)\n\naverage_logits = torch.mean(torch.stack([logits1, logits2]), dim=0)\n\nprediction = torch.softmax(average_logits, dim=1)  # or sigmoid -> depends on task\n\n```\n\n",
              "votes": 6
            },
            {
              "id": 2105401,
              "postDate": "2023-01-18T13:26:45.687Z",
              "content": "<p>With logits, the scale is different, so it might be effective to scale the output of each model with sigmoid (or softmax) before ensemble.</p>",
              "rawMarkdown": "With logits, the scale is different, so it might be effective to scale the output of each model with sigmoid (or softmax) before ensemble.",
              "votes": 3
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2111427,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2023-01-22T22:20:44.380000",
      "content": "<p>Well, you can do it this way or take the sigmoid/softmax before you take the average. Try both approaches and see which one gives better results (on validation).</p>\n<p>I agree with <a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a>, the magnitudes of each models' logits might be quite different so having the probabilities is safer before averaging.</p>\n<p>Good luck!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2103342,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2023-01-17T06:33:39.593000",
      "content": "<p>yes it can be done. The best practice is always check 😊</p>",
      "votes": -2,
      "replies": [
        {
          "id": 2103894,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2023-01-17T12:50:19.880000",
          "content": "<p>I know it can be done, but I was asking <em>how</em> it's done</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2103929,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-01-17T13:21:28.633000",
              "content": "<p>assuming that your model output are logits …</p>\n<pre><code>model1 = resnet()\nmodel2 = resnet()\n\nlogits1 = model1(image)\nlogits2 = model2(image)\n\naverage_logits = torch.mean(torch.stack([logits1, logits2]), dim=0)\n\nprediction = torch.softmax(average_logits, dim=1)  # or sigmoid -&gt; depends on task\n</code></pre>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2105401,
              "author_name": "YYama",
              "author_url": "",
              "post_date": "2023-01-18T13:26:45.687000",
              "content": "<p>With logits, the scale is different, so it might be effective to scale the output of each model with sigmoid (or softmax) before ensemble.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2103055": "I was wondering if models could be ensembled using logits. \n\nI have seen models being ensembled using model probabilities, like averaging output probabilities of two models. I've also seen ensembles which concatenate each individual model backbone's final layer and then use a classification head (though idk if this is wrong if the models were not trained in this way).\n\nIs there a way to ensemble models by summing/averaging logits? Is this a good approach?",
    "2111427": "Well, you can do it this way or take the sigmoid/softmax before you take the average. Try both approaches and see which one gives better results (on validation).\n\nI agree with @yosukeyama, the magnitudes of each models' logits might be quite different so having the probabilities is safer before averaging.\n\nGood luck!",
    "2103342": "yes it can be done. The best practice is always check 😊"
  }
}