{
  "id": 369704,
  "title": "Recommend Loss function for F1_Score Task. It's mainly Medical Image Analysis.",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/369704",
  "author_name": "Wongi Park",
  "post_date": "2022-12-01T05:57:32.909000",
  "votes": 19,
  "comment_count": 2,
  "views": 0,
  "content": "<h2>Recommend loss function</h2>\n<p>First. Problem of the F1-Score is cannot be used as loss function to compute gradient. <br><br>\nTherefore we use \"Dice loss\" or \"soft f1 socre\"as loss function. But, be careful It's approximate for f1 score.</p>\n<p>This code will helpful in competition.</p>\n<h2>Dice Loss implement pytorch</h2>\n<pre><code>def dice_loss(pred, target, smooth = 1e-5):\n    # binary cross entropy loss\n    bce = F.binary_cross_entropy_with_logits(pred, target, reduction='sum')\n\n    pred = torch.sigmoid(pred)\n    intersection = (pred * target).sum(dim=(2,3))\n    union = pred.sum(dim=(2,3)) + target.sum(dim=(2,3))\n\n    # dice coefficient\n    dice = 2.0 * (intersection + smooth) / (union + smooth)\n\n    # dice loss\n    dice_loss = 1.0 - dice\n\n    # total loss\n    loss = bce + dice_loss\n\n    return loss.sum(), dice.sum()\n</code></pre>\n<h2>soft f1 Loss implement</h2>\n<pre><code>def macro_soft_f1(y, y_hat):\n    \"\"\"Compute the macro soft F1-score as a cost.\n    Average (1 - soft-F1) across all labels.\n    Use probability values instead of binary predictions.\n\n    Args:\n        y (int32 Tensor): targets array of shape (BATCH_SIZE, N_LABELS)\n        y_hat (float32 Tensor): probability matrix of shape (BATCH_SIZE, N_LABELS)\n\n    Returns:\n        cost (scalar Tensor): value of the cost function for the batch\n    \"\"\"\n\n    y = tf.cast(y, tf.float32)\n    y_hat = tf.cast(y_hat, tf.float32)\n    tp = tf.reduce_sum(y_hat * y, axis=0)\n    fp = tf.reduce_sum(y_hat * (1 - y), axis=0)\n    fn = tf.reduce_sum((1 - y_hat) * y, axis=0)\n    soft_f1 = 2*tp / (2*tp + fn + fp + 1e-16)\n    cost = 1 - soft_f1 # reduce 1 - soft-f1 in order to increase soft-f1\n    macro_cost = tf.reduce_mean(cost) # average on all labels\n\n    return macro_cost\n</code></pre>",
  "messages": [
    {
      "id": 2050968,
      "postDate": "2022-12-01T05:57:32.910Z",
      "content": "<h2>Recommend loss function</h2>\n<p>First. Problem of the F1-Score is cannot be used as loss function to compute gradient. <br><br>\nTherefore we use \"Dice loss\" or \"soft f1 socre\"as loss function. But, be careful It's approximate for f1 score.</p>\n<p>This code will helpful in competition.</p>\n<h2>Dice Loss implement pytorch</h2>\n<pre><code>def dice_loss(pred, target, smooth = 1e-5):\n    # binary cross entropy loss\n    bce = F.binary_cross_entropy_with_logits(pred, target, reduction='sum')\n\n    pred = torch.sigmoid(pred)\n    intersection = (pred * target).sum(dim=(2,3))\n    union = pred.sum(dim=(2,3)) + target.sum(dim=(2,3))\n\n    # dice coefficient\n    dice = 2.0 * (intersection + smooth) / (union + smooth)\n\n    # dice loss\n    dice_loss = 1.0 - dice\n\n    # total loss\n    loss = bce + dice_loss\n\n    return loss.sum(), dice.sum()\n</code></pre>\n<h2>soft f1 Loss implement</h2>\n<pre><code>def macro_soft_f1(y, y_hat):\n    \"\"\"Compute the macro soft F1-score as a cost.\n    Average (1 - soft-F1) across all labels.\n    Use probability values instead of binary predictions.\n\n    Args:\n        y (int32 Tensor): targets array of shape (BATCH_SIZE, N_LABELS)\n        y_hat (float32 Tensor): probability matrix of shape (BATCH_SIZE, N_LABELS)\n\n    Returns:\n        cost (scalar Tensor): value of the cost function for the batch\n    \"\"\"\n\n    y = tf.cast(y, tf.float32)\n    y_hat = tf.cast(y_hat, tf.float32)\n    tp = tf.reduce_sum(y_hat * y, axis=0)\n    fp = tf.reduce_sum(y_hat * (1 - y), axis=0)\n    fn = tf.reduce_sum((1 - y_hat) * y, axis=0)\n    soft_f1 = 2*tp / (2*tp + fn + fp + 1e-16)\n    cost = 1 - soft_f1 # reduce 1 - soft-f1 in order to increase soft-f1\n    macro_cost = tf.reduce_mean(cost) # average on all labels\n\n    return macro_cost\n</code></pre>",
      "rawMarkdown": "## Recommend loss function\nFirst. Problem of the F1-Score is cannot be used as loss function to compute gradient. </br>\nTherefore we use \"Dice loss\" or \"soft f1 socre\"as loss function. But, be careful It's approximate for f1 score.\n\nThis code will helpful in competition.\n\n## Dice Loss implement pytorch\n```\ndef dice_loss(pred, target, smooth = 1e-5):\n    # binary cross entropy loss\n    bce = F.binary_cross_entropy_with_logits(pred, target, reduction='sum')\n    \n    pred = torch.sigmoid(pred)\n    intersection = (pred * target).sum(dim=(2,3))\n    union = pred.sum(dim=(2,3)) + target.sum(dim=(2,3))\n    \n    # dice coefficient\n    dice = 2.0 * (intersection + smooth) / (union + smooth)\n    \n    # dice loss\n    dice_loss = 1.0 - dice\n    \n    # total loss\n    loss = bce + dice_loss\n    \n    return loss.sum(), dice.sum()\n```\n\n## soft f1 Loss implement\n```\ndef macro_soft_f1(y, y_hat):\n    \"\"\"Compute the macro soft F1-score as a cost.\n    Average (1 - soft-F1) across all labels.\n    Use probability values instead of binary predictions.\n    \n    Args:\n        y (int32 Tensor): targets array of shape (BATCH_SIZE, N_LABELS)\n        y_hat (float32 Tensor): probability matrix of shape (BATCH_SIZE, N_LABELS)\n        \n    Returns:\n        cost (scalar Tensor): value of the cost function for the batch\n    \"\"\"\n    \n    y = tf.cast(y, tf.float32)\n    y_hat = tf.cast(y_hat, tf.float32)\n    tp = tf.reduce_sum(y_hat * y, axis=0)\n    fp = tf.reduce_sum(y_hat * (1 - y), axis=0)\n    fn = tf.reduce_sum((1 - y_hat) * y, axis=0)\n    soft_f1 = 2*tp / (2*tp + fn + fp + 1e-16)\n    cost = 1 - soft_f1 # reduce 1 - soft-f1 in order to increase soft-f1\n    macro_cost = tf.reduce_mean(cost) # average on all labels\n    \n    return macro_cost\n```",
      "votes": 19
    },
    {
      "id": 2051180,
      "postDate": "2022-12-01T09:10:14.843Z",
      "content": "<p>That is a good point, <a href=\"https://www.kaggle.com/kalelpark\" target=\"_blank\">@kalelpark</a>! Curious how well they will help us address the class imbalance problem… </p>",
      "rawMarkdown": "That is a good point, @kalelpark! Curious how well they will help us address the class imbalance problem... ",
      "votes": 1
    },
    {
      "id": 2054849,
      "postDate": "2022-12-04T13:41:03.380Z",
      "content": "<p>you can also see this implementation: <a href=\"https://gist.github.com/SuperShinyEyes/dcc68a08ff8b615442e3bc6a9b55a354?permalink_comment_id=3055663#gistcomment-3055663\" target=\"_blank\">https://gist.github.com/SuperShinyEyes/dcc68a08ff8b615442e3bc6a9b55a354?permalink_comment_id=3055663#gistcomment-3055663</a></p>",
      "rawMarkdown": "you can also see this implementation: https://gist.github.com/SuperShinyEyes/dcc68a08ff8b615442e3bc6a9b55a354?permalink_comment_id=3055663#gistcomment-3055663",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2051180,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2022-12-01T09:10:14.843000",
      "content": "<p>That is a good point, <a href=\"https://www.kaggle.com/kalelpark\" target=\"_blank\">@kalelpark</a>! Curious how well they will help us address the class imbalance problem… </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2054849,
      "author_name": "Jirka",
      "author_url": "",
      "post_date": "2022-12-04T13:41:03.380000",
      "content": "<p>you can also see this implementation: <a href=\"https://gist.github.com/SuperShinyEyes/dcc68a08ff8b615442e3bc6a9b55a354?permalink_comment_id=3055663#gistcomment-3055663\" target=\"_blank\">https://gist.github.com/SuperShinyEyes/dcc68a08ff8b615442e3bc6a9b55a354?permalink_comment_id=3055663#gistcomment-3055663</a></p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2050968": "## Recommend loss function\nFirst. Problem of the F1-Score is cannot be used as loss function to compute gradient. </br>\nTherefore we use \"Dice loss\" or \"soft f1 socre\"as loss function. But, be careful It's approximate for f1 score.\n\nThis code will helpful in competition.\n\n## Dice Loss implement pytorch\n```\ndef dice_loss(pred, target, smooth = 1e-5):\n    # binary cross entropy loss\n    bce = F.binary_cross_entropy_with_logits(pred, target, reduction='sum')\n    \n    pred = torch.sigmoid(pred)\n    intersection = (pred * target).sum(dim=(2,3))\n    union = pred.sum(dim=(2,3)) + target.sum(dim=(2,3))\n    \n    # dice coefficient\n    dice = 2.0 * (intersection + smooth) / (union + smooth)\n    \n    # dice loss\n    dice_loss = 1.0 - dice\n    \n    # total loss\n    loss = bce + dice_loss\n    \n    return loss.sum(), dice.sum()\n```\n\n## soft f1 Loss implement\n```\ndef macro_soft_f1(y, y_hat):\n    \"\"\"Compute the macro soft F1-score as a cost.\n    Average (1 - soft-F1) across all labels.\n    Use probability values instead of binary predictions.\n    \n    Args:\n        y (int32 Tensor): targets array of shape (BATCH_SIZE, N_LABELS)\n        y_hat (float32 Tensor): probability matrix of shape (BATCH_SIZE, N_LABELS)\n        \n    Returns:\n        cost (scalar Tensor): value of the cost function for the batch\n    \"\"\"\n    \n    y = tf.cast(y, tf.float32)\n    y_hat = tf.cast(y_hat, tf.float32)\n    tp = tf.reduce_sum(y_hat * y, axis=0)\n    fp = tf.reduce_sum(y_hat * (1 - y), axis=0)\n    fn = tf.reduce_sum((1 - y_hat) * y, axis=0)\n    soft_f1 = 2*tp / (2*tp + fn + fp + 1e-16)\n    cost = 1 - soft_f1 # reduce 1 - soft-f1 in order to increase soft-f1\n    macro_cost = tf.reduce_mean(cost) # average on all labels\n    \n    return macro_cost\n```",
    "2051180": "That is a good point, @kalelpark! Curious how well they will help us address the class imbalance problem... ",
    "2054849": "you can also see this implementation: https://gist.github.com/SuperShinyEyes/dcc68a08ff8b615442e3bc6a9b55a354?permalink_comment_id=3055663#gistcomment-3055663"
  }
}