{
  "id": 114415,
  "title": "PyTorch validation loss function [Help]",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/114415",
  "author_name": "Kalyan S.S",
  "post_date": "2019-10-26T03:27:31.481000",
  "votes": 0,
  "comment_count": 6,
  "views": 0,
  "content": "<p><code>\ndef weighted_log_loss_metric(trues, preds):\n     class_weights = torch.tensor([2., 1., 1., 1., 1., 1.]).to(device,dtype=torch.float)\n     epsilon = 1e-7 <br>\n     preds = torch.clamp(preds, epsilon, 1.0-epsilon).float()\n     loss = trues * torch.log(preds) + (1.0 - trues) * torch.log(1 - preds)\n     loss_samples = torch.sum(class_weights * loss,axis=1)/torch.sum(class_weights)\n   return - loss_samples.mean()\n</code>\nI defined following loss function for getting LB, but torch.clamp is not returning 1.0-1e-7 (it's returning 1.0) so i'm getting NaN as a result of  this(torch.log(1 - preds))</p>\n\n<p>Can someone help where I'm going wrong?</p>",
  "messages": [
    {
      "id": 658559,
      "postDate": "2019-10-26T06:43:58.527Z",
      "content": "<p>i think you can just use BCEWithLogitsLoss with weights. e.g.</p>\n\n<p>```\ndef criterion_label(logit, truth, weight=None):\n    batch_size,num_class = logit.shape[:2]\n    logit = logit.view(batch_size,num_class)\n    truth = truth.view(batch_size,num_class)</p>\n\n<pre><code>if weight is None: weight=[2,1,1,1,1,1]\nweight = torch.FloatTensor(weight).to(truth.device).view(1,-1)\n\nloss = F.binary_cross_entropy_with_logits(logit,truth,reduction='none')\nloss = loss*weight\nloss = loss.mean()\nreturn loss\n</code></pre>\n\n<p>```</p>\n\n<p>i verify the above code is the same as:</p>\n\n<p>```\ndef compute_kaggle_metric(truth_label, probability_label):\n    eps = 1e-15</p>\n\n<pre><code>t = truth_label.reshape(-1,6)\np = probability_label.reshape(-1,6)\np = np.clip(  p, eps, 1-eps)\nn = np.clip(1-p, eps, 1-eps)\nlog_p = -np.log(p)\nlog_n = -np.log(n)\n\nmetric = (1-t)*log_n + t*log_p\nmetric = metric.mean(0)\nmetric_weight = [2,1,1, 1,1,1]\nscore = (metric_weight*metric).mean()\n\nreturn score, metric, metric_weight\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "i think you can just use BCEWithLogitsLoss with weights. e.g.\n\n```\ndef criterion_label(logit, truth, weight=None):\n    batch_size,num_class = logit.shape[:2]\n    logit = logit.view(batch_size,num_class)\n    truth = truth.view(batch_size,num_class)\n\n    if weight is None: weight=[2,1,1,1,1,1]\n    weight = torch.FloatTensor(weight).to(truth.device).view(1,-1)\n\n    loss = F.binary_cross_entropy_with_logits(logit,truth,reduction='none')\n    loss = loss*weight\n    loss = loss.mean()\n    return loss\n\n\n```\n\ni verify the above code is the same as:\n\n```\ndef compute_kaggle_metric(truth_label, probability_label):\n    eps = 1e-15\n\n    t = truth_label.reshape(-1,6)\n    p = probability_label.reshape(-1,6)\n    p = np.clip(  p, eps, 1-eps)\n    n = np.clip(1-p, eps, 1-eps)\n    log_p = -np.log(p)\n    log_n = -np.log(n)\n\n    metric = (1-t)*log_n + t*log_p\n    metric = metric.mean(0)\n    metric_weight = [2,1,1, 1,1,1]\n    score = (metric_weight*metric).mean()\n\n    return score, metric, metric_weight\n\n```",
      "votes": 1,
      "replies": [
        {
          "id": 658659,
          "postDate": "2019-10-26T10:14:18.017Z",
          "content": "<p>Thanks</p>",
          "rawMarkdown": "Thanks"
        }
      ]
    },
    {
      "id": 658479,
      "postDate": "2019-10-26T03:46:20.843Z",
      "content": "<p>Note that you are losing the gradient tracking when you do clamping like that. Why do you get values of ones and zeros for <code>preds</code> to begin with? I assume it is an output from <code>torch.sigmoid</code>, and ones and zeros happen rarely for numerical reasons? In this case, that is why in torch we have two function, <a href=\"https://pytorch.org/docs/stable/nn.html#bceloss\"><code>BCELoss</code></a>and <a href=\"https://pytorch.org/docs/stable/nn.html#bcewithlogitsloss\"><code>BCEWithLogitsLoss</code></a>. Read the description for the second one, this is what you need instead of <code>torch.clamp</code>. </p>",
      "rawMarkdown": "Note that you are losing the gradient tracking when you do clamping like that. Why do you get values of ones and zeros for `preds` to begin with? I assume it is an output from `torch.sigmoid`, and ones and zeros happen rarely for numerical reasons? In this case, that is why in torch we have two function, [`BCELoss `](https://pytorch.org/docs/stable/nn.html#bceloss)and [`BCEWithLogitsLoss`](https://pytorch.org/docs/stable/nn.html#bcewithlogitsloss). Read the description for the second one, this is what you need instead of `torch.clamp`. ",
      "votes": 1,
      "replies": [
        {
          "id": 658658,
          "postDate": "2019-10-26T10:13:59.143Z",
          "content": "<p>I got my mistake, I was passing without sigmoid 😅 , Thanks!! I think BCEWithLogitsLoss with  weights will work</p>",
          "rawMarkdown": "I got my mistake, I was passing without sigmoid 😅 , Thanks!! I think BCEWithLogitsLoss with  weights will work",
          "votes": 2
        }
      ]
    },
    {
      "id": 660563,
      "postDate": "2019-10-29T10:16:25.460Z",
      "content": "<p>I think that you should replace float with float64</p>",
      "rawMarkdown": "I think that you should replace float with float64\n"
    },
    {
      "id": 658735,
      "postDate": "2019-10-26T13:04:16.623Z",
      "content": "<p>Python implementation of loss function should be marginally suboptimal to pytorch's readily implemented loss. I was very frustrated with the loss function and solved the problem recently. I made a post of it here <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/114243#latest-658461\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/114243#latest-658461</a>. You just have to change the weights since it appears you use \"any\" as your first class</p>",
      "rawMarkdown": "Python implementation of loss function should be marginally suboptimal to pytorch's readily implemented loss. I was very frustrated with the loss function and solved the problem recently. I made a post of it here https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/114243#latest-658461. You just have to change the weights since it appears you use \"any\" as your first class"
    },
    {
      "id": 658470,
      "postDate": "2019-10-26T03:27:31.480Z",
      "content": "<p><code>\ndef weighted_log_loss_metric(trues, preds):\n     class_weights = torch.tensor([2., 1., 1., 1., 1., 1.]).to(device,dtype=torch.float)\n     epsilon = 1e-7 <br>\n     preds = torch.clamp(preds, epsilon, 1.0-epsilon).float()\n     loss = trues * torch.log(preds) + (1.0 - trues) * torch.log(1 - preds)\n     loss_samples = torch.sum(class_weights * loss,axis=1)/torch.sum(class_weights)\n   return - loss_samples.mean()\n</code>\nI defined following loss function for getting LB, but torch.clamp is not returning 1.0-1e-7 (it's returning 1.0) so i'm getting NaN as a result of  this(torch.log(1 - preds))</p>\n\n<p>Can someone help where I'm going wrong?</p>",
      "rawMarkdown": "```\ndef weighted_log_loss_metric(trues, preds):\n     class_weights = torch.tensor([2., 1., 1., 1., 1., 1.]).to(device,dtype=torch.float)\n     epsilon = 1e-7   \n     preds = torch.clamp(preds, epsilon, 1.0-epsilon).float()\n     loss = trues * torch.log(preds) + (1.0 - trues) * torch.log(1 - preds)\n     loss_samples = torch.sum(class_weights * loss,axis=1)/torch.sum(class_weights)\n   return - loss_samples.mean()\n```\nI defined following loss function for getting LB, but torch.clamp is not returning 1.0-1e-7 (it's returning 1.0) so i'm getting NaN as a result of  this(torch.log(1 - preds))\n\nCan someone help where I'm going wrong?"
    }
  ],
  "comments": [
    {
      "id": 658559,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-10-26T06:43:58.527000",
      "content": "<p>i think you can just use BCEWithLogitsLoss with weights. e.g.</p>\n\n<p>```\ndef criterion_label(logit, truth, weight=None):\n    batch_size,num_class = logit.shape[:2]\n    logit = logit.view(batch_size,num_class)\n    truth = truth.view(batch_size,num_class)</p>\n\n<pre><code>if weight is None: weight=[2,1,1,1,1,1]\nweight = torch.FloatTensor(weight).to(truth.device).view(1,-1)\n\nloss = F.binary_cross_entropy_with_logits(logit,truth,reduction='none')\nloss = loss*weight\nloss = loss.mean()\nreturn loss\n</code></pre>\n\n<p>```</p>\n\n<p>i verify the above code is the same as:</p>\n\n<p>```\ndef compute_kaggle_metric(truth_label, probability_label):\n    eps = 1e-15</p>\n\n<pre><code>t = truth_label.reshape(-1,6)\np = probability_label.reshape(-1,6)\np = np.clip(  p, eps, 1-eps)\nn = np.clip(1-p, eps, 1-eps)\nlog_p = -np.log(p)\nlog_n = -np.log(n)\n\nmetric = (1-t)*log_n + t*log_p\nmetric = metric.mean(0)\nmetric_weight = [2,1,1, 1,1,1]\nscore = (metric_weight*metric).mean()\n\nreturn score, metric, metric_weight\n</code></pre>\n\n<p>```</p>",
      "votes": 1,
      "replies": [
        {
          "id": 658659,
          "author_name": "Kalyan S.S",
          "author_url": "",
          "post_date": "2019-10-26T10:14:18.017000",
          "content": "<p>Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 658479,
      "author_name": "nosound",
      "author_url": "",
      "post_date": "2019-10-26T03:46:20.843000",
      "content": "<p>Note that you are losing the gradient tracking when you do clamping like that. Why do you get values of ones and zeros for <code>preds</code> to begin with? I assume it is an output from <code>torch.sigmoid</code>, and ones and zeros happen rarely for numerical reasons? In this case, that is why in torch we have two function, <a href=\"https://pytorch.org/docs/stable/nn.html#bceloss\"><code>BCELoss</code></a>and <a href=\"https://pytorch.org/docs/stable/nn.html#bcewithlogitsloss\"><code>BCEWithLogitsLoss</code></a>. Read the description for the second one, this is what you need instead of <code>torch.clamp</code>. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 658658,
          "author_name": "Kalyan S.S",
          "author_url": "",
          "post_date": "2019-10-26T10:13:59.143000",
          "content": "<p>I got my mistake, I was passing without sigmoid 😅 , Thanks!! I think BCEWithLogitsLoss with  weights will work</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 660563,
      "author_name": "Hadar",
      "author_url": "",
      "post_date": "2019-10-29T10:16:25.460000",
      "content": "<p>I think that you should replace float with float64</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 658735,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2019-10-26T13:04:16.623000",
      "content": "<p>Python implementation of loss function should be marginally suboptimal to pytorch's readily implemented loss. I was very frustrated with the loss function and solved the problem recently. I made a post of it here <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/114243#latest-658461\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/114243#latest-658461</a>. You just have to change the weights since it appears you use \"any\" as your first class</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "658559": "i think you can just use BCEWithLogitsLoss with weights. e.g.\n\n```\ndef criterion_label(logit, truth, weight=None):\n    batch_size,num_class = logit.shape[:2]\n    logit = logit.view(batch_size,num_class)\n    truth = truth.view(batch_size,num_class)\n\n    if weight is None: weight=[2,1,1,1,1,1]\n    weight = torch.FloatTensor(weight).to(truth.device).view(1,-1)\n\n    loss = F.binary_cross_entropy_with_logits(logit,truth,reduction='none')\n    loss = loss*weight\n    loss = loss.mean()\n    return loss\n\n\n```\n\ni verify the above code is the same as:\n\n```\ndef compute_kaggle_metric(truth_label, probability_label):\n    eps = 1e-15\n\n    t = truth_label.reshape(-1,6)\n    p = probability_label.reshape(-1,6)\n    p = np.clip(  p, eps, 1-eps)\n    n = np.clip(1-p, eps, 1-eps)\n    log_p = -np.log(p)\n    log_n = -np.log(n)\n\n    metric = (1-t)*log_n + t*log_p\n    metric = metric.mean(0)\n    metric_weight = [2,1,1, 1,1,1]\n    score = (metric_weight*metric).mean()\n\n    return score, metric, metric_weight\n\n```",
    "658479": "Note that you are losing the gradient tracking when you do clamping like that. Why do you get values of ones and zeros for `preds` to begin with? I assume it is an output from `torch.sigmoid`, and ones and zeros happen rarely for numerical reasons? In this case, that is why in torch we have two function, [`BCELoss `](https://pytorch.org/docs/stable/nn.html#bceloss)and [`BCEWithLogitsLoss`](https://pytorch.org/docs/stable/nn.html#bcewithlogitsloss). Read the description for the second one, this is what you need instead of `torch.clamp`. ",
    "660563": "I think that you should replace float with float64\n",
    "658735": "Python implementation of loss function should be marginally suboptimal to pytorch's readily implemented loss. I was very frustrated with the loss function and solved the problem recently. I made a post of it here https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/114243#latest-658461. You just have to change the weights since it appears you use \"any\" as your first class",
    "658470": "```\ndef weighted_log_loss_metric(trues, preds):\n     class_weights = torch.tensor([2., 1., 1., 1., 1., 1.]).to(device,dtype=torch.float)\n     epsilon = 1e-7   \n     preds = torch.clamp(preds, epsilon, 1.0-epsilon).float()\n     loss = trues * torch.log(preds) + (1.0 - trues) * torch.log(1 - preds)\n     loss_samples = torch.sum(class_weights * loss,axis=1)/torch.sum(class_weights)\n   return - loss_samples.mean()\n```\nI defined following loss function for getting LB, but torch.clamp is not returning 1.0-1e-7 (it's returning 1.0) so i'm getting NaN as a result of  this(torch.log(1 - preds))\n\nCan someone help where I'm going wrong?"
  }
}