{
  "id": 429475,
  "title": "Fold Ensemble vs Model Ensemble?",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/429475",
  "author_name": "Bartley",
  "post_date": "2023-08-05T19:19:24.222000",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I am considering two ways to ensemble models for this competition.</p>\n<p>In the figures below, m1 = model 1, m2 = model 2, and m3 = model 3. Each model is considered to have a different structure and achieves similar accuracy on a holdout set.</p>\n<h3>Option 1. Fold Ensemble</h3>\n<p>Aggregate predictions in each fold before the final ensemble.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F530dc8567f2044bbac20d0f81b7a3088%2Fens-fold_ensemble.drawio%20(1).png?generation=1691263031233671&amp;alt=media\" alt=\"\"></p>\n<h3>Option 2. Model Ensemble</h3>\n<p>Aggregate predictions for each model type before the final ensemble.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F1397a2a808d984e0f490195631de9555%2Fens-model%20ensemble.drawio%20(1).png?generation=1691263014510718&amp;alt=media\" alt=\"\"></p>\n<p>Due to the computational requirements of training each model, I have been unable to test the performance of each ensemble type using CV. I suspect that the difference is negligible, but I am interested to see what people think. Has anyone tried these methods in previous competitions?</p>\n<p>Note:  <strong>Each ensemble uses the median for aggregation</strong></p>",
  "messages": [
    {
      "id": 2375729,
      "postDate": "2023-08-05T21:07:38.493Z",
      "content": "<p>Hello, I am not sure both aren't the exact same thing…<br>\nFrom my understanding, each fold-model would have 1/9th of weight in both cases.</p>",
      "rawMarkdown": "Hello, I am not sure both aren't the exact same thing...\nFrom my understanding, each fold-model would have 1/9th of weight in both cases.",
      "votes": 4,
      "replies": [
        {
          "id": 2375770,
          "postDate": "2023-08-05T22:02:01.237Z",
          "content": "<p>I checked all possible predictions for 9 models, and 85.9% (440/512) of the time the final result is the same. Consider this counter example. </p>\n<p>Here I am considering 0:3 in the 1st fold, 3:6 in the 2nd fold, and 6:9 in the final fold. Every third model has the same backbone [0, 3, 6], [1, 4, 7], [2, 5, 8].</p>\n<p>preds = [1, 1, 1, 1, 0, 0, 0, 0, 1]</p>\n<p>Model Ensemble.<br>\nm1 = [1,1,0] = 1<br>\nm2 = [1,0,0] = 0<br>\nm3 = [1,0,1] = 1<br>\n= 1</p>\n<p>Fold Ensemble.<br>\nf1 = [1,1,1] = 1<br>\nf2 = [1,0,0] = 0<br>\nf3 = [0,0,1] = 0<br>\n= 0</p>",
          "rawMarkdown": "I checked all possible predictions for 9 models, and 85.9% (440/512) of the time the final result is the same. Consider this counter example. \n\nHere I am considering 0:3 in the 1st fold, 3:6 in the 2nd fold, and 6:9 in the final fold. Every third model has the same backbone [0, 3, 6], [1, 4, 7], [2, 5, 8].\n\npreds = [1, 1, 1, 1, 0, 0, 0, 0, 1]\n\nModel Ensemble.\nm1 = [1,1,0] = 1\nm2 = [1,0,0] = 0\nm3 = [1,0,1] = 1\n= 1\n\nFold Ensemble.\nf1 = [1,1,1] = 1\nf2 = [1,0,0] = 0\nf3 = [0,0,1] = 0\n= 0",
          "replies": [
            {
              "id": 2375785,
              "postDate": "2023-08-05T22:23:42.610Z",
              "content": "<p>I think that there is an error in the \"m1 = [1,1,0] = 0\" if I got it right, is it possible? Also the thing here is that you are applying the threshold at each sub prediction and most of the times it is better to ensemble the raw probabilities with mean or weighted average and just apply the threshold at the final step to have the final mask with 0s and 1s.</p>",
              "rawMarkdown": "I think that there is an error in the \"m1 = [1,1,0] = 0\" if I got it right, is it possible? Also the thing here is that you are applying the threshold at each sub prediction and most of the times it is better to ensemble the raw probabilities with mean or weighted average and just apply the threshold at the final step to have the final mask with 0s and 1s.",
              "votes": 1
            },
            {
              "id": 2377032,
              "postDate": "2023-08-06T21:46:50.340Z",
              "content": "<p>The way I do things:</p>\n<p>preds = [1, 1, 1, 1, 0, 0, 0, 0, 1]</p>\n<pre><code>Model Ensemble.\nm1 = [,,] = \nm2 = [,,] = \nm3 = [,,] = \n=  /  = , -&gt; \n</code></pre>\n<pre><code>Fold Ensemble.\nf1 = [,,] = \nf2 = [,,] = \nf3 = [,,] = \n=  / = , -&gt; \n</code></pre>",
              "rawMarkdown": "The way I do things:\n\npreds = [1, 1, 1, 1, 0, 0, 0, 0, 1]\n\n```python\nModel Ensemble.\nm1 = [1,1,0] = 0.66\nm2 = [1,0,0] = 0.33\nm3 = [1,0,1] = 0.66\n= 1.66 / 3 = 0,55 -> 1\n```\n\n```python\nFold Ensemble.\nf1 = [1,1,1] = 1\nf2 = [1,0,0] = 0.33\nf3 = [0,0,1] = 0.33\n= 1.66 /3 = 0,55 -> 1\n```\n\n",
              "votes": 1
            },
            {
              "id": 2377078,
              "postDate": "2023-08-06T23:25:32.997Z",
              "content": "<p>Ok I understand where you are coming from now. This is actually quite a smart solution! </p>\n<p>I thought that you were ensembling prior to converting predictions to a 0 or 1.</p>",
              "rawMarkdown": "Ok I understand where you are coming from now. This is actually quite a smart solution! \n\nI thought that you were ensembling prior to converting predictions to a 0 or 1.",
              "votes": 1
            }
          ]
        },
        {
          "id": 2375783,
          "postDate": "2023-08-05T22:19:49.223Z",
          "content": "<p>I also understand them both as the same thing, at the end you have 9 different checkpoints from [3 architectures x 3 train-sets] or viceversa. If the ensemble is done by a simple mean of probabilities at each level, both options should have same results. One could even simply leave the 9 predictions stacked and do the mean of all them together.</p>",
          "rawMarkdown": "I also understand them both as the same thing, at the end you have 9 different checkpoints from [3 architectures x 3 train-sets] or viceversa. If the ensemble is done by a simple mean of probabilities at each level, both options should have same results. One could even simply leave the 9 predictions stacked and do the mean of all them together.",
          "votes": 1,
          "replies": [
            {
              "id": 2375808,
              "postDate": "2023-08-05T23:17:00.720Z",
              "content": "<p>I should have been clear that I was assuming median aggregation.. In this case the predictions will differ 72/512 times (with 9 models)?</p>\n<p>Yes, I agree that for mean aggregation the weight of each model will be 1/9th of the total, and the two options will identical to a 9 prediction stack.</p>\n<hr>\n<p>Here is the code to reproduce the results for median aggregation.</p>\n<pre><code> itertools\n numpy as np\n\n generate_binary_arrays(length):\n     length &lt;= :\n         ValueError()\n     =\n    \n\n = \n = generate_binary_arrays(length)\n = \n = \n\n arr in all_arrays:\n    \n     = np.median([np.median(arr[:]), np.median(arr[:]), np.median(arr[:])])\n    \n     = np.median([np.median([arr[], arr[], arr[]]),np.median([arr[], arr[], arr[]]), np.median([arr[], arr[], arr[]])])\n    (arr, p1, p2)\n    +=\n     p1 == p2:\n        +=\n(sc, c, sc/c)\n</code></pre>",
              "rawMarkdown": "I should have been clear that I was assuming median aggregation.. In this case the predictions will differ 72/512 times (with 9 models)?\n\nYes, I agree that for mean aggregation the weight of each model will be 1/9th of the total, and the two options will identical to a 9 prediction stack.\n\n----\n\nHere is the code to reproduce the results for median aggregation.\n\n```\nimport itertools\nimport numpy as np\n\ndef generate_binary_arrays(length):\n    if length <= 0:\n        raise ValueError(\"Length must be greater than 0.\")\n    possible_values = [0, 1]\n    return [list(array) for array in itertools.product(possible_values, repeat=length)]\n\nlength = 9\nall_arrays = generate_binary_arrays(length)\nc = 0\nsc = 0\n\nfor arr in all_arrays:\n    # Fold Ensemble\n    p1 = np.median([np.median(arr[0:3]), np.median(arr[3:6]), np.median(arr[6:9])])\n    # Model Ensemble\n    p2 = np.median([np.median([arr[0], arr[3], arr[6]]),np.median([arr[1], arr[4], arr[7]]), np.median([arr[2], arr[5], arr[8]])])\n    print(arr, p1, p2)\n    c+=1\n    if p1 == p2:\n        sc+=1\nprint(sc, c, sc/c)\n```",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2375635,
      "postDate": "2023-08-05T19:19:24.223Z",
      "content": "<p>I am considering two ways to ensemble models for this competition.</p>\n<p>In the figures below, m1 = model 1, m2 = model 2, and m3 = model 3. Each model is considered to have a different structure and achieves similar accuracy on a holdout set.</p>\n<h3>Option 1. Fold Ensemble</h3>\n<p>Aggregate predictions in each fold before the final ensemble.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F530dc8567f2044bbac20d0f81b7a3088%2Fens-fold_ensemble.drawio%20(1).png?generation=1691263031233671&amp;alt=media\" alt=\"\"></p>\n<h3>Option 2. Model Ensemble</h3>\n<p>Aggregate predictions for each model type before the final ensemble.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F1397a2a808d984e0f490195631de9555%2Fens-model%20ensemble.drawio%20(1).png?generation=1691263014510718&amp;alt=media\" alt=\"\"></p>\n<p>Due to the computational requirements of training each model, I have been unable to test the performance of each ensemble type using CV. I suspect that the difference is negligible, but I am interested to see what people think. Has anyone tried these methods in previous competitions?</p>\n<p>Note:  <strong>Each ensemble uses the median for aggregation</strong></p>",
      "rawMarkdown": "I am considering two ways to ensemble models for this competition.\n\nIn the figures below, m1 = model 1, m2 = model 2, and m3 = model 3. Each model is considered to have a different structure and achieves similar accuracy on a holdout set.\n\n### Option 1. Fold Ensemble\n\nAggregate predictions in each fold before the final ensemble.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F530dc8567f2044bbac20d0f81b7a3088%2Fens-fold_ensemble.drawio%20(1).png?generation=1691263031233671&alt=media)\n\n### Option 2. Model Ensemble\n\nAggregate predictions for each model type before the final ensemble.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F1397a2a808d984e0f490195631de9555%2Fens-model%20ensemble.drawio%20(1).png?generation=1691263014510718&alt=media)\n\nDue to the computational requirements of training each model, I have been unable to test the performance of each ensemble type using CV. I suspect that the difference is negligible, but I am interested to see what people think. Has anyone tried these methods in previous competitions?\n\nNote:  **Each ensemble uses the median for aggregation**",
      "votes": 1
    },
    {
      "id": 2375755,
      "postDate": "2023-08-05T21:30:26.290Z",
      "rawMarkdown": "",
      "votes": -6,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2375729,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2023-08-05T21:07:38.493000",
      "content": "<p>Hello, I am not sure both aren't the exact same thing…<br>\nFrom my understanding, each fold-model would have 1/9th of weight in both cases.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2375770,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2023-08-05T22:02:01.237000",
          "content": "<p>I checked all possible predictions for 9 models, and 85.9% (440/512) of the time the final result is the same. Consider this counter example. </p>\n<p>Here I am considering 0:3 in the 1st fold, 3:6 in the 2nd fold, and 6:9 in the final fold. Every third model has the same backbone [0, 3, 6], [1, 4, 7], [2, 5, 8].</p>\n<p>preds = [1, 1, 1, 1, 0, 0, 0, 0, 1]</p>\n<p>Model Ensemble.<br>\nm1 = [1,1,0] = 1<br>\nm2 = [1,0,0] = 0<br>\nm3 = [1,0,1] = 1<br>\n= 1</p>\n<p>Fold Ensemble.<br>\nf1 = [1,1,1] = 1<br>\nf2 = [1,0,0] = 0<br>\nf3 = [0,0,1] = 0<br>\n= 0</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2375785,
              "author_name": "Enric Domingo",
              "author_url": "",
              "post_date": "2023-08-05T22:23:42.610000",
              "content": "<p>I think that there is an error in the \"m1 = [1,1,0] = 0\" if I got it right, is it possible? Also the thing here is that you are applying the threshold at each sub prediction and most of the times it is better to ensemble the raw probabilities with mean or weighted average and just apply the threshold at the final step to have the final mask with 0s and 1s.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2377032,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-08-06T21:46:50.340000",
              "content": "<p>The way I do things:</p>\n<p>preds = [1, 1, 1, 1, 0, 0, 0, 0, 1]</p>\n<pre><code>Model Ensemble.\nm1 = [,,] = \nm2 = [,,] = \nm3 = [,,] = \n=  /  = , -&gt; \n</code></pre>\n<pre><code>Fold Ensemble.\nf1 = [,,] = \nf2 = [,,] = \nf3 = [,,] = \n=  / = , -&gt; \n</code></pre>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2377078,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2023-08-06T23:25:32.997000",
              "content": "<p>Ok I understand where you are coming from now. This is actually quite a smart solution! </p>\n<p>I thought that you were ensembling prior to converting predictions to a 0 or 1.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2375783,
          "author_name": "Enric Domingo",
          "author_url": "",
          "post_date": "2023-08-05T22:19:49.223000",
          "content": "<p>I also understand them both as the same thing, at the end you have 9 different checkpoints from [3 architectures x 3 train-sets] or viceversa. If the ensemble is done by a simple mean of probabilities at each level, both options should have same results. One could even simply leave the 9 predictions stacked and do the mean of all them together.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2375808,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2023-08-05T23:17:00.720000",
              "content": "<p>I should have been clear that I was assuming median aggregation.. In this case the predictions will differ 72/512 times (with 9 models)?</p>\n<p>Yes, I agree that for mean aggregation the weight of each model will be 1/9th of the total, and the two options will identical to a 9 prediction stack.</p>\n<hr>\n<p>Here is the code to reproduce the results for median aggregation.</p>\n<pre><code> itertools\n numpy as np\n\n generate_binary_arrays(length):\n     length &lt;= :\n         ValueError()\n     =\n    \n\n = \n = generate_binary_arrays(length)\n = \n = \n\n arr in all_arrays:\n    \n     = np.median([np.median(arr[:]), np.median(arr[:]), np.median(arr[:])])\n    \n     = np.median([np.median([arr[], arr[], arr[]]),np.median([arr[], arr[], arr[]]), np.median([arr[], arr[], arr[]])])\n    (arr, p1, p2)\n    +=\n     p1 == p2:\n        +=\n(sc, c, sc/c)\n</code></pre>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2375755,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-05T21:30:26.290000",
      "content": "",
      "votes": -6,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2375729": "Hello, I am not sure both aren't the exact same thing...\nFrom my understanding, each fold-model would have 1/9th of weight in both cases.",
    "2375635": "I am considering two ways to ensemble models for this competition.\n\nIn the figures below, m1 = model 1, m2 = model 2, and m3 = model 3. Each model is considered to have a different structure and achieves similar accuracy on a holdout set.\n\n### Option 1. Fold Ensemble\n\nAggregate predictions in each fold before the final ensemble.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F530dc8567f2044bbac20d0f81b7a3088%2Fens-fold_ensemble.drawio%20(1).png?generation=1691263031233671&alt=media)\n\n### Option 2. Model Ensemble\n\nAggregate predictions for each model type before the final ensemble.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F1397a2a808d984e0f490195631de9555%2Fens-model%20ensemble.drawio%20(1).png?generation=1691263014510718&alt=media)\n\nDue to the computational requirements of training each model, I have been unable to test the performance of each ensemble type using CV. I suspect that the difference is negligible, but I am interested to see what people think. Has anyone tried these methods in previous competitions?\n\nNote:  **Each ensemble uses the median for aggregation**",
    "2375755": ""
  }
}