{
  "id": 428922,
  "title": "Surprisingly bad performance of BCE loss",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/428922",
  "author_name": "Raki",
  "post_date": "2023-08-03T10:19:35.121000",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I iterated through a range of learning rates for BCE loss and Dice Loss and compared both with the best lrs. <br>\nThe difference for me was enormous, I have the feeling that I made some mistake in the implementation, as I don't see how the difference between these losses could be so large. </p>\n<p>weight = torch.FloatTensor([10.0]).cuda()  # Move to GPU if available<br>\ncriterion = smp.losses.SoftBCEWithLogitsLoss(pos_weight=weight)<br>\nvs<br>\ncriterion = smp.losses.DiceLoss(mode=\"binary\", smooth=1.0)</p>\n<p><a href=\"https://wandb.ai/ralf-c-kinkel/Contrail/reports/Dice-vs-WBCE--Vmlldzo1MDQ0MDQ2\" target=\"_blank\">https://wandb.ai/ralf-c-kinkel/Contrail/reports/Dice-vs-WBCE--Vmlldzo1MDQ0MDQ2</a></p>\n<p>EDIT: I found my error and now WBCE looks better (by around +0.01), my mistake was that I did not adjust the threshold. When optimizing with WBCE the model assigns overly high confidence to contrails. When training a simple model with Dice Loss (ResNeSt 26d) I get around 0.62 val dice. When training with WBCE, weight=10 and iterating over the thresholds my best threshold at 0.8 gets 0.632. Some of that comes from overfitting the threshold, but even at 0.75 and 0.85 threshold Val Dice still almost at 0.63 (still +0.01 vs Dice Loss trained).</p>\n<p>EDIT: I integrated soft-labels, meaning that I define groundtruth as mean of individual masks, I attached an example. It turns out that Dice Loss benefits from that much more than WBCE, though only with a high threshold (0.99+ was ideal for me), I will make a post later today where I look at interaction of Dice Loss and these averaged labels in more detail. My Val Dice is now at 0.65 with this setup for ResNest26d. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F03dac235a6cf386bd242942fbec14d4d%2FSoft%20Labels.png?generation=1691232543701634&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2373608,
      "postDate": "2023-08-04T10:38:35.230Z",
      "content": "<p>During the initials stages of experimenting I have run some hyper-parameter optimization and found that for bce loss pos_weight selection is very important because there is a narrow optimum around ~10 value. And it seems that it yields worse performance vs only dice or sum of bce and dice. However, for sum of bce and dice with 0.5 weight each the optimum is lower (~1e-2) and seems to be much more flatten.</p>\n<p>See the figure below: 4 different models, bce / bce + dice loss and pos_weight in [1e-3, 1e3] interval selected randomly. I highlighted the bce by red color and bce + dice by green color, different models seem to behave similarly.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fc7d9a37e2087cbdb41fda4e3da915a87%2Fbest_vdice_vs_pos_weight.png?generation=1691145428101753&amp;alt=media\" alt=\"\"></p>\n<p>So my experiment results also suggest that bce is worse than dice.</p>",
      "rawMarkdown": "During the initials stages of experimenting I have run some hyper-parameter optimization and found that for bce loss pos_weight selection is very important because there is a narrow optimum around ~10 value. And it seems that it yields worse performance vs only dice or sum of bce and dice. However, for sum of bce and dice with 0.5 weight each the optimum is lower (~1e-2) and seems to be much more flatten.\n\nSee the figure below: 4 different models, bce / bce + dice loss and pos_weight in [1e-3, 1e3] interval selected randomly. I highlighted the bce by red color and bce + dice by green color, different models seem to behave similarly.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fc7d9a37e2087cbdb41fda4e3da915a87%2Fbest_vdice_vs_pos_weight.png?generation=1691145428101753&alt=media)\n\nSo my experiment results also suggest that bce is worse than dice.",
      "votes": 3,
      "replies": [
        {
          "id": 2373611,
          "postDate": "2023-08-04T10:41:29.423Z",
          "content": "<p>Note: there is only few points for bce loss option, only single fold (out of 5) and low number of epochs (20) was used, so decide it for yourself whether to trust it</p>",
          "rawMarkdown": "Note: there is only few points for bce loss option, only single fold (out of 5) and low number of epochs (20) was used, so decide it for yourself whether to trust it"
        }
      ]
    },
    {
      "id": 2371888,
      "postDate": "2023-08-03T10:19:35.123Z",
      "content": "<p>I iterated through a range of learning rates for BCE loss and Dice Loss and compared both with the best lrs. <br>\nThe difference for me was enormous, I have the feeling that I made some mistake in the implementation, as I don't see how the difference between these losses could be so large. </p>\n<p>weight = torch.FloatTensor([10.0]).cuda()  # Move to GPU if available<br>\ncriterion = smp.losses.SoftBCEWithLogitsLoss(pos_weight=weight)<br>\nvs<br>\ncriterion = smp.losses.DiceLoss(mode=\"binary\", smooth=1.0)</p>\n<p><a href=\"https://wandb.ai/ralf-c-kinkel/Contrail/reports/Dice-vs-WBCE--Vmlldzo1MDQ0MDQ2\" target=\"_blank\">https://wandb.ai/ralf-c-kinkel/Contrail/reports/Dice-vs-WBCE--Vmlldzo1MDQ0MDQ2</a></p>\n<p>EDIT: I found my error and now WBCE looks better (by around +0.01), my mistake was that I did not adjust the threshold. When optimizing with WBCE the model assigns overly high confidence to contrails. When training a simple model with Dice Loss (ResNeSt 26d) I get around 0.62 val dice. When training with WBCE, weight=10 and iterating over the thresholds my best threshold at 0.8 gets 0.632. Some of that comes from overfitting the threshold, but even at 0.75 and 0.85 threshold Val Dice still almost at 0.63 (still +0.01 vs Dice Loss trained).</p>\n<p>EDIT: I integrated soft-labels, meaning that I define groundtruth as mean of individual masks, I attached an example. It turns out that Dice Loss benefits from that much more than WBCE, though only with a high threshold (0.99+ was ideal for me), I will make a post later today where I look at interaction of Dice Loss and these averaged labels in more detail. My Val Dice is now at 0.65 with this setup for ResNest26d. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F03dac235a6cf386bd242942fbec14d4d%2FSoft%20Labels.png?generation=1691232543701634&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I iterated through a range of learning rates for BCE loss and Dice Loss and compared both with the best lrs. \nThe difference for me was enormous, I have the feeling that I made some mistake in the implementation, as I don't see how the difference between these losses could be so large. \n\nweight = torch.FloatTensor([10.0]).cuda()  # Move to GPU if available\ncriterion = smp.losses.SoftBCEWithLogitsLoss(pos_weight=weight)\nvs\ncriterion = smp.losses.DiceLoss(mode=\"binary\", smooth=1.0)\n\nhttps://wandb.ai/ralf-c-kinkel/Contrail/reports/Dice-vs-WBCE--Vmlldzo1MDQ0MDQ2\n\nEDIT: I found my error and now WBCE looks better (by around +0.01), my mistake was that I did not adjust the threshold. When optimizing with WBCE the model assigns overly high confidence to contrails. When training a simple model with Dice Loss (ResNeSt 26d) I get around 0.62 val dice. When training with WBCE, weight=10 and iterating over the thresholds my best threshold at 0.8 gets 0.632. Some of that comes from overfitting the threshold, but even at 0.75 and 0.85 threshold Val Dice still almost at 0.63 (still +0.01 vs Dice Loss trained).\n\nEDIT: I integrated soft-labels, meaning that I define groundtruth as mean of individual masks, I attached an example. It turns out that Dice Loss benefits from that much more than WBCE, though only with a high threshold (0.99+ was ideal for me), I will make a post later today where I look at interaction of Dice Loss and these averaged labels in more detail. My Val Dice is now at 0.65 with this setup for ResNest26d. ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F03dac235a6cf386bd242942fbec14d4d%2FSoft%20Labels.png?generation=1691232543701634&alt=media)",
      "votes": 3
    },
    {
      "id": 2374085,
      "postDate": "2023-08-04T16:52:04.727Z",
      "content": "<p>Here is the Dice Loss trained vs WBCE trained model global val dice comparison for different thresholds, while WBCE peaks a bit higher, the Dice Loss trained model is much more stable with non-optimal thresholds.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F18734ca960d92fd96cde750431538402%2FDice%20vs%20WBCE%20trained%20threshold.png?generation=1691167916813679&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Here is the Dice Loss trained vs WBCE trained model global val dice comparison for different thresholds, while WBCE peaks a bit higher, the Dice Loss trained model is much more stable with non-optimal thresholds.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F18734ca960d92fd96cde750431538402%2FDice%20vs%20WBCE%20trained%20threshold.png?generation=1691167916813679&alt=media)"
    },
    {
      "id": 2373253,
      "postDate": "2023-08-04T07:25:27.087Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13974146%2F612e6fed31e39f1d44558714843fbb38%2F2023-08-04%2015.24.43.png?generation=1691133924754866&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13974146%2F612e6fed31e39f1d44558714843fbb38%2F2023-08-04%2015.24.43.png?generation=1691133924754866&alt=media)"
    }
  ],
  "comments": [
    {
      "id": 2373608,
      "author_name": "Mikhail Kotyushev",
      "author_url": "",
      "post_date": "2023-08-04T10:38:35.230000",
      "content": "<p>During the initials stages of experimenting I have run some hyper-parameter optimization and found that for bce loss pos_weight selection is very important because there is a narrow optimum around ~10 value. And it seems that it yields worse performance vs only dice or sum of bce and dice. However, for sum of bce and dice with 0.5 weight each the optimum is lower (~1e-2) and seems to be much more flatten.</p>\n<p>See the figure below: 4 different models, bce / bce + dice loss and pos_weight in [1e-3, 1e3] interval selected randomly. I highlighted the bce by red color and bce + dice by green color, different models seem to behave similarly.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fc7d9a37e2087cbdb41fda4e3da915a87%2Fbest_vdice_vs_pos_weight.png?generation=1691145428101753&amp;alt=media\" alt=\"\"></p>\n<p>So my experiment results also suggest that bce is worse than dice.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2373611,
          "author_name": "Mikhail Kotyushev",
          "author_url": "",
          "post_date": "2023-08-04T10:41:29.423000",
          "content": "<p>Note: there is only few points for bce loss option, only single fold (out of 5) and low number of epochs (20) was used, so decide it for yourself whether to trust it</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2374085,
      "author_name": "Raki",
      "author_url": "",
      "post_date": "2023-08-04T16:52:04.727000",
      "content": "<p>Here is the Dice Loss trained vs WBCE trained model global val dice comparison for different thresholds, while WBCE peaks a bit higher, the Dice Loss trained model is much more stable with non-optimal thresholds.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F18734ca960d92fd96cde750431538402%2FDice%20vs%20WBCE%20trained%20threshold.png?generation=1691167916813679&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2373253,
      "author_name": "DL",
      "author_url": "",
      "post_date": "2023-08-04T07:25:27.087000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13974146%2F612e6fed31e39f1d44558714843fbb38%2F2023-08-04%2015.24.43.png?generation=1691133924754866&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2373608": "During the initials stages of experimenting I have run some hyper-parameter optimization and found that for bce loss pos_weight selection is very important because there is a narrow optimum around ~10 value. And it seems that it yields worse performance vs only dice or sum of bce and dice. However, for sum of bce and dice with 0.5 weight each the optimum is lower (~1e-2) and seems to be much more flatten.\n\nSee the figure below: 4 different models, bce / bce + dice loss and pos_weight in [1e-3, 1e3] interval selected randomly. I highlighted the bce by red color and bce + dice by green color, different models seem to behave similarly.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fc7d9a37e2087cbdb41fda4e3da915a87%2Fbest_vdice_vs_pos_weight.png?generation=1691145428101753&alt=media)\n\nSo my experiment results also suggest that bce is worse than dice.",
    "2371888": "I iterated through a range of learning rates for BCE loss and Dice Loss and compared both with the best lrs. \nThe difference for me was enormous, I have the feeling that I made some mistake in the implementation, as I don't see how the difference between these losses could be so large. \n\nweight = torch.FloatTensor([10.0]).cuda()  # Move to GPU if available\ncriterion = smp.losses.SoftBCEWithLogitsLoss(pos_weight=weight)\nvs\ncriterion = smp.losses.DiceLoss(mode=\"binary\", smooth=1.0)\n\nhttps://wandb.ai/ralf-c-kinkel/Contrail/reports/Dice-vs-WBCE--Vmlldzo1MDQ0MDQ2\n\nEDIT: I found my error and now WBCE looks better (by around +0.01), my mistake was that I did not adjust the threshold. When optimizing with WBCE the model assigns overly high confidence to contrails. When training a simple model with Dice Loss (ResNeSt 26d) I get around 0.62 val dice. When training with WBCE, weight=10 and iterating over the thresholds my best threshold at 0.8 gets 0.632. Some of that comes from overfitting the threshold, but even at 0.75 and 0.85 threshold Val Dice still almost at 0.63 (still +0.01 vs Dice Loss trained).\n\nEDIT: I integrated soft-labels, meaning that I define groundtruth as mean of individual masks, I attached an example. It turns out that Dice Loss benefits from that much more than WBCE, though only with a high threshold (0.99+ was ideal for me), I will make a post later today where I look at interaction of Dice Loss and these averaged labels in more detail. My Val Dice is now at 0.65 with this setup for ResNest26d. ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F03dac235a6cf386bd242942fbec14d4d%2FSoft%20Labels.png?generation=1691232543701634&alt=media)",
    "2374085": "Here is the Dice Loss trained vs WBCE trained model global val dice comparison for different thresholds, while WBCE peaks a bit higher, the Dice Loss trained model is much more stable with non-optimal thresholds.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F18734ca960d92fd96cde750431538402%2FDice%20vs%20WBCE%20trained%20threshold.png?generation=1691167916813679&alt=media)",
    "2373253": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13974146%2F612e6fed31e39f1d44558714843fbb38%2F2023-08-04%2015.24.43.png?generation=1691133924754866&alt=media)"
  }
}