{
  "id": 412554,
  "title": "Loss Fuction: Dice Loss or WCE?",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/412554",
  "author_name": "LUPIN11",
  "post_date": "2023-05-24T07:56:58.708000",
  "votes": 17,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi, kagglers. <br>\nI've tried 2 kinds of loss functions: Dice Loss and WCE.<br>\nIn my cases, WCE performs better than Dice Loss.<br>\nMy model training code: <a href=\"https://www.kaggle.com/code/lupin11/u-net-baseline-training\" target=\"_blank\">https://www.kaggle.com/code/lupin11/u-net-baseline-training</a><br>\nWCE: Weighted Cross Entropy</p>\n<pre><code>def ce_loss(y_p, y_t):\n    weight = torch.Tensor([0.57, 4.17]).to('cuda')\n    criterion = nn.CrossEntropyLoss(weight)\n    loss = criterion(y_p, y_t)\n    return loss\n</code></pre>\n<p>Dice Loss:</p>\n<pre><code>def dice_loss(y_p, y_t, smooth=1e-6):\n    y_p = y_p.reshape(-1)\n    y_t = y_t.reshape(-1)\n    i = (y_p * y_t).sum()\n    return 1 - (2. * i + smooth) / (y_p.sum() + y_t.sum() + smooth)\n</code></pre>\n<p><strong>Experiments:</strong><br>\nThe case using Dice Loss:<br>\noutput without activation/with Sigmoid: divergence<br>\noutput with ReLu: suffer from neuron death<br>\nWhen using WCE, the loss decreases without fluctuations. And I also track the dice coeff(dice loss = 1 - dice coeff) at the same time, finding it increases with significant fluctuations.<br>\nSome predictions by a model trained by WCE on my validation set(5th epoch and ValidDice=0.3917):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F885c19ff7183a20f9e261ad356f7dc95%2Fm2.png?generation=1684910042108786&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F0e98c9e4cc7a5e5b6a45a00a004e734d%2Fm3.png?generation=1684910103250212&amp;alt=media\" alt=\"![\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F74a07c83e345e8e217b9afe857f2e9d6%2Fm4.png?generation=1684910169420398&amp;alt=media\" alt=\"\"><br>\n<strong>Try to explain why Dice Loss is hard to train:</strong><br>\nTo simplify the problem, let's consider only one pixel: Xi（X for pixels in prediction, Y for pixels in groundtruth）<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F5feeb8e1f83e2ec9b453d7118f955c25%2Ff2.png?generation=1684912463014238&amp;alt=media\" alt=\"\"><br>\nWe notice that, regardless of whether Yi=0 or Yi=1, the absolute value of the gradient of Xi follows this:<br>\nA larger Xi always has a smaller abs(gradient) and a smaller Xi always has a larger abs(gradient).<br>\nThis leads to 2 problems:<br>\n<strong>1. Gradient Vanishing</strong><br>\nThe model tends to stagnate on pixels having very large ouput owing to very small abs(gradient).<br>\nAnd too many pixels like this make C2 larger, then abs(gradient) on every pixels become smaller.<br>\n<strong>2. Pointless Gradient</strong><br>\nSuppose that 0&lt;=Xi&lt;=1.<br>\nWhen Yi=0 (holds on most pixels), the closer Xi is to 0, the larger abs(gradient), and the closer Xi is to 1, the smaller abs(gradient).<br>\nThis makes no sense in that such gradient make model tend to escape right and stick to wrong.<br>\n<strong>Try to explain why Dice Loss suffer from fluctuations while WCE loss is decreasing:</strong><br>\nAn example may provide some insights:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F57197922e271c61337bf59f58f0bde72%2Ffu.PNG?generation=1684939408853534&amp;alt=media\" alt=\"\"><br>\ndice coeff of pred1 = 0.1/(1+0.3*8+0.1) = 0.02857</p>\n<p>dice coeff of pred2 = 0.05/(1+0.15*8+0.05) = 0.0222<br>\nwce of pred1 = -(log0.1 + log0.7)m = -log0.07<br>\nwce of pred2 = -(log0.05 + log0.85) = -log0.0425<br>\nAlthough the WCE loss of pred2 is smaller, its dice coeff is lower</p>",
  "messages": [
    {
      "id": 3233020,
      "postDate": "2025-06-26T12:18:18.890Z",
      "content": "<p>it's exactly the case with me in concrete defect segmentation.</p>",
      "rawMarkdown": "it's exactly the case with me in concrete defect segmentation.",
      "votes": 1
    },
    {
      "id": 2271957,
      "postDate": "2023-05-24T07:56:58.710Z",
      "content": "<p>Hi, kagglers. <br>\nI've tried 2 kinds of loss functions: Dice Loss and WCE.<br>\nIn my cases, WCE performs better than Dice Loss.<br>\nMy model training code: <a href=\"https://www.kaggle.com/code/lupin11/u-net-baseline-training\" target=\"_blank\">https://www.kaggle.com/code/lupin11/u-net-baseline-training</a><br>\nWCE: Weighted Cross Entropy</p>\n<pre><code>def ce_loss(y_p, y_t):\n    weight = torch.Tensor([0.57, 4.17]).to('cuda')\n    criterion = nn.CrossEntropyLoss(weight)\n    loss = criterion(y_p, y_t)\n    return loss\n</code></pre>\n<p>Dice Loss:</p>\n<pre><code>def dice_loss(y_p, y_t, smooth=1e-6):\n    y_p = y_p.reshape(-1)\n    y_t = y_t.reshape(-1)\n    i = (y_p * y_t).sum()\n    return 1 - (2. * i + smooth) / (y_p.sum() + y_t.sum() + smooth)\n</code></pre>\n<p><strong>Experiments:</strong><br>\nThe case using Dice Loss:<br>\noutput without activation/with Sigmoid: divergence<br>\noutput with ReLu: suffer from neuron death<br>\nWhen using WCE, the loss decreases without fluctuations. And I also track the dice coeff(dice loss = 1 - dice coeff) at the same time, finding it increases with significant fluctuations.<br>\nSome predictions by a model trained by WCE on my validation set(5th epoch and ValidDice=0.3917):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F885c19ff7183a20f9e261ad356f7dc95%2Fm2.png?generation=1684910042108786&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F0e98c9e4cc7a5e5b6a45a00a004e734d%2Fm3.png?generation=1684910103250212&amp;alt=media\" alt=\"![\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F74a07c83e345e8e217b9afe857f2e9d6%2Fm4.png?generation=1684910169420398&amp;alt=media\" alt=\"\"><br>\n<strong>Try to explain why Dice Loss is hard to train:</strong><br>\nTo simplify the problem, let's consider only one pixel: Xi（X for pixels in prediction, Y for pixels in groundtruth）<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F5feeb8e1f83e2ec9b453d7118f955c25%2Ff2.png?generation=1684912463014238&amp;alt=media\" alt=\"\"><br>\nWe notice that, regardless of whether Yi=0 or Yi=1, the absolute value of the gradient of Xi follows this:<br>\nA larger Xi always has a smaller abs(gradient) and a smaller Xi always has a larger abs(gradient).<br>\nThis leads to 2 problems:<br>\n<strong>1. Gradient Vanishing</strong><br>\nThe model tends to stagnate on pixels having very large ouput owing to very small abs(gradient).<br>\nAnd too many pixels like this make C2 larger, then abs(gradient) on every pixels become smaller.<br>\n<strong>2. Pointless Gradient</strong><br>\nSuppose that 0&lt;=Xi&lt;=1.<br>\nWhen Yi=0 (holds on most pixels), the closer Xi is to 0, the larger abs(gradient), and the closer Xi is to 1, the smaller abs(gradient).<br>\nThis makes no sense in that such gradient make model tend to escape right and stick to wrong.<br>\n<strong>Try to explain why Dice Loss suffer from fluctuations while WCE loss is decreasing:</strong><br>\nAn example may provide some insights:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F57197922e271c61337bf59f58f0bde72%2Ffu.PNG?generation=1684939408853534&amp;alt=media\" alt=\"\"><br>\ndice coeff of pred1 = 0.1/(1+0.3*8+0.1) = 0.02857</p>\n<p>dice coeff of pred2 = 0.05/(1+0.15*8+0.05) = 0.0222<br>\nwce of pred1 = -(log0.1 + log0.7)m = -log0.07<br>\nwce of pred2 = -(log0.05 + log0.85) = -log0.0425<br>\nAlthough the WCE loss of pred2 is smaller, its dice coeff is lower</p>",
      "rawMarkdown": "Hi, kagglers. \nI've tried 2 kinds of loss functions: Dice Loss and WCE.\nIn my cases, WCE performs better than Dice Loss.\nMy model training code: [https://www.kaggle.com/code/lupin11/u-net-baseline-training](https://www.kaggle.com/code/lupin11/u-net-baseline-training)\nWCE: Weighted Cross Entropy\n```\ndef ce_loss(y_p, y_t):\n    weight = torch.Tensor([0.57, 4.17]).to('cuda')\n    criterion = nn.CrossEntropyLoss(weight)\n    loss = criterion(y_p, y_t)\n    return loss\n```\nDice Loss:\n```\ndef dice_loss(y_p, y_t, smooth=1e-6):\n    y_p = y_p.reshape(-1)\n    y_t = y_t.reshape(-1)\n    i = (y_p * y_t).sum()\n    return 1 - (2. * i + smooth) / (y_p.sum() + y_t.sum() + smooth)\n```\n**Experiments:**\nThe case using Dice Loss:\noutput without activation/with Sigmoid: divergence\noutput with ReLu: suffer from neuron death\nWhen using WCE, the loss decreases without fluctuations. And I also track the dice coeff(dice loss = 1 - dice coeff) at the same time, finding it increases with significant fluctuations.\nSome predictions by a model trained by WCE on my validation set(5th epoch and ValidDice=0.3917):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F885c19ff7183a20f9e261ad356f7dc95%2Fm2.png?generation=1684910042108786&alt=media)\n![![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F0e98c9e4cc7a5e5b6a45a00a004e734d%2Fm3.png?generation=1684910103250212&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F74a07c83e345e8e217b9afe857f2e9d6%2Fm4.png?generation=1684910169420398&alt=media)\n**Try to explain why Dice Loss is hard to train:**\nTo simplify the problem, let's consider only one pixel: Xi（X for pixels in prediction, Y for pixels in groundtruth）\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F5feeb8e1f83e2ec9b453d7118f955c25%2Ff2.png?generation=1684912463014238&alt=media)\nWe notice that, regardless of whether Yi=0 or Yi=1, the absolute value of the gradient of Xi follows this:\nA larger Xi always has a smaller abs(gradient) and a smaller Xi always has a larger abs(gradient).\nThis leads to 2 problems:\n**1. Gradient Vanishing**\nThe model tends to stagnate on pixels having very large ouput owing to very small abs(gradient).\nAnd too many pixels like this make C2 larger, then abs(gradient) on every pixels become smaller.\n**2. Pointless Gradient**\nSuppose that 0<=Xi<=1.\nWhen Yi=0 (holds on most pixels), the closer Xi is to 0, the larger abs(gradient), and the closer Xi is to 1, the smaller abs(gradient).\nThis makes no sense in that such gradient make model tend to escape right and stick to wrong.\n**Try to explain why Dice Loss suffer from fluctuations while WCE loss is decreasing:**\nAn example may provide some insights:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F57197922e271c61337bf59f58f0bde72%2Ffu.PNG?generation=1684939408853534&alt=media)\ndice coeff of pred1 = 0.1/(1+0.3*8+0.1) = 0.02857\n\ndice coeff of pred2 = 0.05/(1+0.15*8+0.05) = 0.0222\nwce of pred1 = -(log0.1 + log0.7)m = -log0.07\nwce of pred2 = -(log0.05 + log0.85) = -log0.0425\nAlthough the WCE loss of pred2 is smaller, its dice coeff is lower\n",
      "votes": 17
    },
    {
      "id": 2280365,
      "postDate": "2023-05-30T04:41:13.203Z",
      "content": "<p>Looking back, I wonder why you use WCE and not WBCE.<br>\nUsing WBCE you won't have to have a 2 output channels, then apply softmax accros the second dim and pick the second channel only, you will just have to apply a sigmoid.</p>\n<pre><code>pos_weight = torch.Tensor([]).to(device)\ncriterion = nn.BCEWithLogitsLoss(pos_weight=pos_weight)\n</code></pre>",
      "rawMarkdown": "Looking back, I wonder why you use WCE and not WBCE.\nUsing WBCE you won't have to have a 2 output channels, then apply softmax accros the second dim and pick the second channel only, you will just have to apply a sigmoid.\n\n```python\npos_weight = torch.Tensor([7.31]).to(device)\ncriterion = nn.BCEWithLogitsLoss(pos_weight=pos_weight)\n```",
      "votes": 3,
      "replies": [
        {
          "id": 2360744,
          "postDate": "2023-07-27T01:34:14.337Z",
          "content": "<p>I'm thinking how the bce loss is calculated on the 2 (n, 256, 256) matrix? It's equal to calculating on 2 (n, 256×256) matrix, I think. You are so smart on the loss function selection! You won't suffer from tuning the threshold. btw, how did you figure out 7.31 for the positive weight? </p>",
          "rawMarkdown": "I'm thinking how the bce loss is calculated on the 2 (n, 256, 256) matrix? It's equal to calculating on 2 (n, 256×256) matrix, I think. You are so smart on the loss function selection! You won't suffer from tuning the threshold. btw, how did you figure out 7.31 for the positive weight? ",
          "replies": [
            {
              "id": 2360793,
              "postDate": "2023-07-27T02:59:30.953Z",
              "content": "<p>7.31 is just based on the loss function that <a href=\"https://www.kaggle.com/lupin11\" target=\"_blank\">@lupin11</a> used in the post. 4.17 / 0.57 = ~ 7.315</p>",
              "rawMarkdown": "7.31 is just based on the loss function that @lupin11 used in the post. 4.17 / 0.57 = ~ 7.315"
            },
            {
              "id": 2360839,
              "postDate": "2023-07-27T03:47:27.033Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/vexxingbanana\" target=\"_blank\">@vexxingbanana</a> , I found that as well. So, I also asked <a href=\"https://www.kaggle.com/lupin11\" target=\"_blank\">@lupin11</a> how he came up with these 2 numbers</p>",
              "rawMarkdown": "Thanks @vexxingbanana , I found that as well. So, I also asked @lupin11 how he came up with these 2 numbers"
            }
          ]
        }
      ]
    },
    {
      "id": 2272029,
      "postDate": "2023-05-24T08:49:11.150Z",
      "content": "<p>Amazing, thanks for you explain of the loss function. How is the perfermance with BCE and WCE?</p>",
      "rawMarkdown": "Amazing, thanks for you explain of the loss function. How is the perfermance with BCE and WCE?",
      "votes": 1
    },
    {
      "id": 2273384,
      "postDate": "2023-05-25T06:18:59.550Z",
      "content": "<p>I have some similar problems with a dice loss function. It doesn't seem to diverge, but it looks to be finding a local minima. I'm looking to try out your loss function, maybe I can get some improvement.</p>",
      "rawMarkdown": "I have some similar problems with a dice loss function. It doesn't seem to diverge, but it looks to be finding a local minima. I'm looking to try out your loss function, maybe I can get some improvement.",
      "votes": 2
    },
    {
      "id": 2360748,
      "postDate": "2023-07-27T01:37:17.173Z",
      "content": "<p>Thanks for sharing! How did figure out 0.57 and 4.17 for the weights?</p>",
      "rawMarkdown": "Thanks for sharing! How did figure out 0.57 and 4.17 for the weights?"
    },
    {
      "id": 2278871,
      "postDate": "2023-05-29T03:33:59.223Z",
      "content": "<p>Hey, have you tried to use the two losses combined ? I'll try today and update this message</p>",
      "rawMarkdown": "Hey, have you tried to use the two losses combined ? I'll try today and update this message",
      "replies": [
        {
          "id": 2278945,
          "postDate": "2023-05-29T05:35:05.427Z",
          "content": "<p>Not yet, but I believe a better loss function might be the key to this competition.</p>",
          "rawMarkdown": "Not yet, but I believe a better loss function might be the key to this competition.",
          "votes": 3,
          "replies": [
            {
              "id": 2279642,
              "postDate": "2023-05-29T13:52:55.510Z",
              "content": "<p>Update: Training with Dice loss + WCE didn't work at all for me but I'm getting very good dice results when training with WCE alone</p>",
              "rawMarkdown": "Update: Training with Dice loss + WCE didn't work at all for me but I'm getting very good dice results when training with WCE alone",
              "votes": 3
            },
            {
              "id": 2360741,
              "postDate": "2023-07-27T01:24:49.880Z",
              "content": "<p>You mean wbce only?</p>",
              "rawMarkdown": "You mean wbce only?",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3233020,
      "author_name": "raid athmane",
      "author_url": "",
      "post_date": "2025-06-26T12:18:18.890000",
      "content": "<p>it's exactly the case with me in concrete defect segmentation.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2280365,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2023-05-30T04:41:13.203000",
      "content": "<p>Looking back, I wonder why you use WCE and not WBCE.<br>\nUsing WBCE you won't have to have a 2 output channels, then apply softmax accros the second dim and pick the second channel only, you will just have to apply a sigmoid.</p>\n<pre><code>pos_weight = torch.Tensor([]).to(device)\ncriterion = nn.BCEWithLogitsLoss(pos_weight=pos_weight)\n</code></pre>",
      "votes": 3,
      "replies": [
        {
          "id": 2360744,
          "author_name": "william.wu",
          "author_url": "",
          "post_date": "2023-07-27T01:34:14.337000",
          "content": "<p>I'm thinking how the bce loss is calculated on the 2 (n, 256, 256) matrix? It's equal to calculating on 2 (n, 256×256) matrix, I think. You are so smart on the loss function selection! You won't suffer from tuning the threshold. btw, how did you figure out 7.31 for the positive weight? </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2360793,
              "author_name": "Ari",
              "author_url": "",
              "post_date": "2023-07-27T02:59:30.953000",
              "content": "<p>7.31 is just based on the loss function that <a href=\"https://www.kaggle.com/lupin11\" target=\"_blank\">@lupin11</a> used in the post. 4.17 / 0.57 = ~ 7.315</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2360839,
              "author_name": "william.wu",
              "author_url": "",
              "post_date": "2023-07-27T03:47:27.033000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/vexxingbanana\" target=\"_blank\">@vexxingbanana</a> , I found that as well. So, I also asked <a href=\"https://www.kaggle.com/lupin11\" target=\"_blank\">@lupin11</a> how he came up with these 2 numbers</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2272029,
      "author_name": "Dewei Chen",
      "author_url": "",
      "post_date": "2023-05-24T08:49:11.150000",
      "content": "<p>Amazing, thanks for you explain of the loss function. How is the perfermance with BCE and WCE?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2273384,
      "author_name": "Ole-Magnus Høiback",
      "author_url": "",
      "post_date": "2023-05-25T06:18:59.550000",
      "content": "<p>I have some similar problems with a dice loss function. It doesn't seem to diverge, but it looks to be finding a local minima. I'm looking to try out your loss function, maybe I can get some improvement.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2360748,
      "author_name": "william.wu",
      "author_url": "",
      "post_date": "2023-07-27T01:37:17.173000",
      "content": "<p>Thanks for sharing! How did figure out 0.57 and 4.17 for the weights?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2278871,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2023-05-29T03:33:59.223000",
      "content": "<p>Hey, have you tried to use the two losses combined ? I'll try today and update this message</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2278945,
          "author_name": "LUPIN11",
          "author_url": "",
          "post_date": "2023-05-29T05:35:05.427000",
          "content": "<p>Not yet, but I believe a better loss function might be the key to this competition.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2279642,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-05-29T13:52:55.510000",
              "content": "<p>Update: Training with Dice loss + WCE didn't work at all for me but I'm getting very good dice results when training with WCE alone</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2360741,
              "author_name": "william.wu",
              "author_url": "",
              "post_date": "2023-07-27T01:24:49.880000",
              "content": "<p>You mean wbce only?</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3233020": "it's exactly the case with me in concrete defect segmentation.",
    "2271957": "Hi, kagglers. \nI've tried 2 kinds of loss functions: Dice Loss and WCE.\nIn my cases, WCE performs better than Dice Loss.\nMy model training code: [https://www.kaggle.com/code/lupin11/u-net-baseline-training](https://www.kaggle.com/code/lupin11/u-net-baseline-training)\nWCE: Weighted Cross Entropy\n```\ndef ce_loss(y_p, y_t):\n    weight = torch.Tensor([0.57, 4.17]).to('cuda')\n    criterion = nn.CrossEntropyLoss(weight)\n    loss = criterion(y_p, y_t)\n    return loss\n```\nDice Loss:\n```\ndef dice_loss(y_p, y_t, smooth=1e-6):\n    y_p = y_p.reshape(-1)\n    y_t = y_t.reshape(-1)\n    i = (y_p * y_t).sum()\n    return 1 - (2. * i + smooth) / (y_p.sum() + y_t.sum() + smooth)\n```\n**Experiments:**\nThe case using Dice Loss:\noutput without activation/with Sigmoid: divergence\noutput with ReLu: suffer from neuron death\nWhen using WCE, the loss decreases without fluctuations. And I also track the dice coeff(dice loss = 1 - dice coeff) at the same time, finding it increases with significant fluctuations.\nSome predictions by a model trained by WCE on my validation set(5th epoch and ValidDice=0.3917):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F885c19ff7183a20f9e261ad356f7dc95%2Fm2.png?generation=1684910042108786&alt=media)\n![![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F0e98c9e4cc7a5e5b6a45a00a004e734d%2Fm3.png?generation=1684910103250212&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F74a07c83e345e8e217b9afe857f2e9d6%2Fm4.png?generation=1684910169420398&alt=media)\n**Try to explain why Dice Loss is hard to train:**\nTo simplify the problem, let's consider only one pixel: Xi（X for pixels in prediction, Y for pixels in groundtruth）\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F5feeb8e1f83e2ec9b453d7118f955c25%2Ff2.png?generation=1684912463014238&alt=media)\nWe notice that, regardless of whether Yi=0 or Yi=1, the absolute value of the gradient of Xi follows this:\nA larger Xi always has a smaller abs(gradient) and a smaller Xi always has a larger abs(gradient).\nThis leads to 2 problems:\n**1. Gradient Vanishing**\nThe model tends to stagnate on pixels having very large ouput owing to very small abs(gradient).\nAnd too many pixels like this make C2 larger, then abs(gradient) on every pixels become smaller.\n**2. Pointless Gradient**\nSuppose that 0<=Xi<=1.\nWhen Yi=0 (holds on most pixels), the closer Xi is to 0, the larger abs(gradient), and the closer Xi is to 1, the smaller abs(gradient).\nThis makes no sense in that such gradient make model tend to escape right and stick to wrong.\n**Try to explain why Dice Loss suffer from fluctuations while WCE loss is decreasing:**\nAn example may provide some insights:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9893459%2F57197922e271c61337bf59f58f0bde72%2Ffu.PNG?generation=1684939408853534&alt=media)\ndice coeff of pred1 = 0.1/(1+0.3*8+0.1) = 0.02857\n\ndice coeff of pred2 = 0.05/(1+0.15*8+0.05) = 0.0222\nwce of pred1 = -(log0.1 + log0.7)m = -log0.07\nwce of pred2 = -(log0.05 + log0.85) = -log0.0425\nAlthough the WCE loss of pred2 is smaller, its dice coeff is lower\n",
    "2280365": "Looking back, I wonder why you use WCE and not WBCE.\nUsing WBCE you won't have to have a 2 output channels, then apply softmax accros the second dim and pick the second channel only, you will just have to apply a sigmoid.\n\n```python\npos_weight = torch.Tensor([7.31]).to(device)\ncriterion = nn.BCEWithLogitsLoss(pos_weight=pos_weight)\n```",
    "2272029": "Amazing, thanks for you explain of the loss function. How is the perfermance with BCE and WCE?",
    "2273384": "I have some similar problems with a dice loss function. It doesn't seem to diverge, but it looks to be finding a local minima. I'm looking to try out your loss function, maybe I can get some improvement.",
    "2360748": "Thanks for sharing! How did figure out 0.57 and 4.17 for the weights?",
    "2278871": "Hey, have you tried to use the two losses combined ? I'll try today and update this message"
  }
}