{
  "id": 114243,
  "title": "Loss function in Pytorch",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/114243",
  "author_name": "Nicholas Lyu",
  "post_date": "2019-10-25T05:18:19.140000",
  "votes": 9,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I have spent a lot of time trying to establish a stable validation. Thanks to my teammate I fixed this problem this morning. Here is the loss function which I am currently using. Taking the mean over validation set should yield a very decent estimation of LB:\n<code>\nreturn torch.nn.functional.binary_cross_entropy_with_logits(y_pred, y_true, pos_weight=weights)\n</code></p>\n\n<p>I used the argument 'weight' instead of 'pos_weight' for my previous models. There was a gap but overall the validation and LB correlates well. The gap is gone when I use the argument 'pos_weight' instead of 'weight'</p>",
  "messages": [
    {
      "id": 657390,
      "postDate": "2019-10-25T05:18:19.140Z",
      "content": "<p>I have spent a lot of time trying to establish a stable validation. Thanks to my teammate I fixed this problem this morning. Here is the loss function which I am currently using. Taking the mean over validation set should yield a very decent estimation of LB:\n<code>\nreturn torch.nn.functional.binary_cross_entropy_with_logits(y_pred, y_true, pos_weight=weights)\n</code></p>\n\n<p>I used the argument 'weight' instead of 'pos_weight' for my previous models. There was a gap but overall the validation and LB correlates well. The gap is gone when I use the argument 'pos_weight' instead of 'weight'</p>",
      "rawMarkdown": "I have spent a lot of time trying to establish a stable validation. Thanks to my teammate I fixed this problem this morning. Here is the loss function which I am currently using. Taking the mean over validation set should yield a very decent estimation of LB:\n`\nreturn torch.nn.functional.binary_cross_entropy_with_logits(y_pred, y_true, pos_weight=weights)\n`\n\nI used the argument 'weight' instead of 'pos_weight' for my previous models. There was a gap but overall the validation and LB correlates well. The gap is gone when I use the argument 'pos_weight' instead of 'weight'",
      "votes": 8
    },
    {
      "id": 657854,
      "postDate": "2019-10-25T13:31:57.393Z",
      "content": "<p>Interesting indeed! What do you think explains this difference ? </p>",
      "rawMarkdown": "Interesting indeed! What do you think explains this difference ? ",
      "votes": 1,
      "replies": [
        {
          "id": 658432,
          "postDate": "2019-10-26T02:01:44.743Z",
          "content": "<p><a href=\"/bdubreu\">@bdubreu</a> According to the pytorch documentation, 'pos_weight' trades precision for recall, whereas 'weight' simply changes the contribution of respective losses to the final loss</p>",
          "rawMarkdown": "@bdubreu According to the pytorch documentation, 'pos_weight' trades precision for recall, whereas 'weight' simply changes the contribution of respective losses to the final loss",
          "votes": 1
        },
        {
          "id": 661725,
          "postDate": "2019-10-30T16:33:07.083Z",
          "content": "<p>So basically the loss is multiplied by two only if the given example was of positive class ? i.e a picture where the model would predict \"any\" but there was in fact no hemorrhage wouldn't get its loss multiplied by two ? </p>",
          "rawMarkdown": "So basically the loss is multiplied by two only if the given example was of positive class ? i.e a picture where the model would predict \"any\" but there was in fact no hemorrhage wouldn't get its loss multiplied by two ? "
        }
      ]
    },
    {
      "id": 661636,
      "postDate": "2019-10-30T14:34:29.527Z",
      "content": "<p>Hi <a href=\"/roguekk007\">@roguekk007</a> ,</p>\n\n<p>are you using this loss both for training and validation?\nDid you notice any improvement in LB when using this loss for training (in case you have experimented)?</p>",
      "rawMarkdown": "Hi @roguekk007 ,\n\nare you using this loss both for training and validation?\nDid you notice any improvement in LB when using this loss for training (in case you have experimented)?",
      "replies": [
        {
          "id": 661959,
          "postDate": "2019-10-30T23:25:15.887Z",
          "content": "<p><a href=\"/bernardohenz\">@bernardohenz</a> No improvement noted, just significantly better estimate of LB using local CV</p>",
          "rawMarkdown": "@bernardohenz No improvement noted, just significantly better estimate of LB using local CV"
        }
      ]
    },
    {
      "id": 659099,
      "postDate": "2019-10-27T04:41:03.220Z",
      "content": "<p>Is there any specific way, you had split validation data? I am doing random split of 10% data, and used loss as you said, but I am not getting decent estimate of LB.\n```\ndef validation(model,device,valid_loader):\n    model.eval()</p>\n\n<pre><code>class_weights = torch.tensor([2., 1., 1., 1., 1., 1.]).to(device,dtype=torch.float)\n\nwith torch.no_grad():\n    tr_loss = 0.0\n    for dev_batch_idx, dev_batch in enumerate(tqdm(valid_loader,desc=\"validation\")):\n        inputs = dev_batch[\"image\"].to(device,dtype=torch.float)\n        labels = dev_batch[\"labels\"].to(device,dtype=torch.float)\n        answer = model(inputs)\n        tr_loss += torch.nn.functional.binary_cross_entropy_with_logits(preds, trues, pos_weight=class_weights)\n\nepoch_loss = tr_loss/len(valid_loader)\nprint(\"validation loss: %.4f\" % epoch_loss)\nreturn epoch_loss\n</code></pre>\n\n<p>```</p>\n\n<p>Can you check my implementation once?</p>",
      "rawMarkdown": "Is there any specific way, you had split validation data? I am doing random split of 10% data, and used loss as you said, but I am not getting decent estimate of LB.\n```\ndef validation(model,device,valid_loader):\n    model.eval()\n\n    class_weights = torch.tensor([2., 1., 1., 1., 1., 1.]).to(device,dtype=torch.float)\n\n    with torch.no_grad():\n        tr_loss = 0.0\n        for dev_batch_idx, dev_batch in enumerate(tqdm(valid_loader,desc=\"validation\")):\n            inputs = dev_batch[\"image\"].to(device,dtype=torch.float)\n            labels = dev_batch[\"labels\"].to(device,dtype=torch.float)\n            answer = model(inputs)\n            tr_loss += torch.nn.functional.binary_cross_entropy_with_logits(preds, trues, pos_weight=class_weights)\n\n    epoch_loss = tr_loss/len(valid_loader)\n    print(\"validation loss: %.4f\" % epoch_loss)\n    return epoch_loss\n```\n\nCan you check my implementation once?",
      "replies": [
        {
          "id": 659108,
          "postDate": "2019-10-27T05:07:48.383Z",
          "content": "<p><a href=\"/saikalyan9981\">@saikalyan9981</a> Random splitting is suboptimal for this challenge, because there are multiple slices taken for a single patient/study which are similar but might be shuffled into train/val set. Your implementation looks alright :)</p>",
          "rawMarkdown": "@saikalyan9981 Random splitting is suboptimal for this challenge, because there are multiple slices taken for a single patient/study which are similar but might be shuffled into train/val set. Your implementation looks alright :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 657674,
      "postDate": "2019-10-25T11:03:10.037Z",
      "content": "<p>Using 'weight' argument for loss instead of 'posweight' always gives an overestimate of LB, but my model yields good LB nevertheless (high-silver range). Anyone who cares to explain the difference?</p>",
      "rawMarkdown": "Using 'weight' argument for loss instead of 'posweight' always gives an overestimate of LB, but my model yields good LB nevertheless (high-silver range). Anyone who cares to explain the difference?",
      "replies": [
        {
          "id": 658151,
          "postDate": "2019-10-25T17:59:33.630Z",
          "content": "<p>I think your initial loss was fine. In my opinion, the main reason why you have higher CV loss compared to stage 1 LB is because you created folds without patient overlap like you said previously on the forums. I'm doing this as well and I also observe a 0.008-0.010 difference. The ones who report good CV/LB correlation usually just did a random fold split or study split. </p>\n\n<p>There is a high risk of overfitting on duplicated patients on stage1 LB which will likely not transfer to stage 2. Consequently, I don`t think using pos_weight is the solution. Having confidence on your CV should be the solution for stage 2.</p>",
          "rawMarkdown": "I think your initial loss was fine. In my opinion, the main reason why you have higher CV loss compared to stage 1 LB is because you created folds without patient overlap like you said previously on the forums. I'm doing this as well and I also observe a 0.008-0.010 difference. The ones who report good CV/LB correlation usually just did a random fold split or study split. \n\nThere is a high risk of overfitting on duplicated patients on stage1 LB which will likely not transfer to stage 2. Consequently, I don`t think using pos_weight is the solution. Having confidence on your CV should be the solution for stage 2.",
          "votes": 2
        },
        {
          "id": 658435,
          "postDate": "2019-10-26T02:04:09.687Z",
          "content": "<p><a href=\"/alexandrecc\">@alexandrecc</a> I think pos_weight is better? First, from numerical values it correlates well and deviates little from LB, so this might actually be the official metric on LB. optimizing 'weight' instead of 'pos_weight' is actually optimizing a subtly different function. I am doing overfitting experiments right now to see whether local overfitting is reflected in my local validation and LB.</p>",
          "rawMarkdown": "@alexandrecc I think pos_weight is better? First, from numerical values it correlates well and deviates little from LB, so this might actually be the official metric on LB. optimizing 'weight' instead of 'pos_weight' is actually optimizing a subtly different function. I am doing overfitting experiments right now to see whether local overfitting is reflected in my local validation and LB.",
          "votes": 1
        }
      ]
    },
    {
      "id": 657669,
      "postDate": "2019-10-25T11:00:34.003Z",
      "content": "<p>Here. I was not at my desktop PC earlier. If you are computing loss on GPU and the sixth class is 'any'.\nTypically, the input is e.g. [32, 3, 200, 200] and model output is [32, 6]. Target should be [32, 6].</p>\n\n<p><code>\nweights = torch.tensor([1.0, 1.0, 1.0, 1.0, 1.0, 2.0]).cuda()\ndef criterion(y_pred,y_true):\n    return torch.nn.functional.binary_cross_entropy_with_logits(y_pred,\n                               y_true, pos_weight=weights)\n</code></p>",
      "rawMarkdown": "Here. I was not at my desktop PC earlier. If you are computing loss on GPU and the sixth class is 'any'.\nTypically, the input is e.g. [32, 3, 200, 200] and model output is [32, 6]. Target should be [32, 6].\n\n```\nweights = torch.tensor([1.0, 1.0, 1.0, 1.0, 1.0, 2.0]).cuda()\ndef criterion(y_pred,y_true):\n    return torch.nn.functional.binary_cross_entropy_with_logits(y_pred,\n                               y_true, pos_weight=weights)\n```",
      "replies": [
        {
          "id": 657803,
          "postDate": "2019-10-25T12:38:51.927Z",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> <br>\n  why have given more weight to class Any which is more than all others i suppose lowest is class 0  subtype.</p>",
          "rawMarkdown": "@roguekk007     \n  why have given more weight to class Any which is more than all others i suppose lowest is class 0  subtype."
        },
        {
          "id": 658461,
          "postDate": "2019-10-26T03:14:12.033Z",
          "content": "<p>epsilon = 1e-7 <br>\npreds = torch.clamp(preds, epsilon, 1.0-epsilon).float()\nIs it right way to clamp?</p>",
          "rawMarkdown": "epsilon = 1e-7   \npreds = torch.clamp(preds, epsilon, 1.0-epsilon).float()\nIs it right way to clamp?"
        }
      ]
    },
    {
      "id": 657633,
      "postDate": "2019-10-25T10:08:51.477Z",
      "content": "<p><a href=\"/roguekk007\">@roguekk007</a>  thanks for posting..\npos_weight would have how many weights here..,basically its dimension?\nany example which u can give ?</p>",
      "rawMarkdown": "@roguekk007  thanks for posting..\npos_weight would have how many weights here..,basically its dimension?\nany example which u can give ?"
    }
  ],
  "comments": [
    {
      "id": 657854,
      "author_name": "Benjamin Dubreu",
      "author_url": "",
      "post_date": "2019-10-25T13:31:57.393000",
      "content": "<p>Interesting indeed! What do you think explains this difference ? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 658432,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2019-10-26T02:01:44.743000",
          "content": "<p><a href=\"/bdubreu\">@bdubreu</a> According to the pytorch documentation, 'pos_weight' trades precision for recall, whereas 'weight' simply changes the contribution of respective losses to the final loss</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 661725,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2019-10-30T16:33:07.083000",
          "content": "<p>So basically the loss is multiplied by two only if the given example was of positive class ? i.e a picture where the model would predict \"any\" but there was in fact no hemorrhage wouldn't get its loss multiplied by two ? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 661636,
      "author_name": "Bernardo",
      "author_url": "",
      "post_date": "2019-10-30T14:34:29.527000",
      "content": "<p>Hi <a href=\"/roguekk007\">@roguekk007</a> ,</p>\n\n<p>are you using this loss both for training and validation?\nDid you notice any improvement in LB when using this loss for training (in case you have experimented)?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 661959,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2019-10-30T23:25:15.887000",
          "content": "<p><a href=\"/bernardohenz\">@bernardohenz</a> No improvement noted, just significantly better estimate of LB using local CV</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 659099,
      "author_name": "Kalyan S.S",
      "author_url": "",
      "post_date": "2019-10-27T04:41:03.220000",
      "content": "<p>Is there any specific way, you had split validation data? I am doing random split of 10% data, and used loss as you said, but I am not getting decent estimate of LB.\n```\ndef validation(model,device,valid_loader):\n    model.eval()</p>\n\n<pre><code>class_weights = torch.tensor([2., 1., 1., 1., 1., 1.]).to(device,dtype=torch.float)\n\nwith torch.no_grad():\n    tr_loss = 0.0\n    for dev_batch_idx, dev_batch in enumerate(tqdm(valid_loader,desc=\"validation\")):\n        inputs = dev_batch[\"image\"].to(device,dtype=torch.float)\n        labels = dev_batch[\"labels\"].to(device,dtype=torch.float)\n        answer = model(inputs)\n        tr_loss += torch.nn.functional.binary_cross_entropy_with_logits(preds, trues, pos_weight=class_weights)\n\nepoch_loss = tr_loss/len(valid_loader)\nprint(\"validation loss: %.4f\" % epoch_loss)\nreturn epoch_loss\n</code></pre>\n\n<p>```</p>\n\n<p>Can you check my implementation once?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 659108,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2019-10-27T05:07:48.383000",
          "content": "<p><a href=\"/saikalyan9981\">@saikalyan9981</a> Random splitting is suboptimal for this challenge, because there are multiple slices taken for a single patient/study which are similar but might be shuffled into train/val set. Your implementation looks alright :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 657674,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2019-10-25T11:03:10.037000",
      "content": "<p>Using 'weight' argument for loss instead of 'posweight' always gives an overestimate of LB, but my model yields good LB nevertheless (high-silver range). Anyone who cares to explain the difference?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 658151,
          "author_name": "Alexandre Cadrin-Chênevert",
          "author_url": "",
          "post_date": "2019-10-25T17:59:33.630000",
          "content": "<p>I think your initial loss was fine. In my opinion, the main reason why you have higher CV loss compared to stage 1 LB is because you created folds without patient overlap like you said previously on the forums. I'm doing this as well and I also observe a 0.008-0.010 difference. The ones who report good CV/LB correlation usually just did a random fold split or study split. </p>\n\n<p>There is a high risk of overfitting on duplicated patients on stage1 LB which will likely not transfer to stage 2. Consequently, I don`t think using pos_weight is the solution. Having confidence on your CV should be the solution for stage 2.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 658435,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2019-10-26T02:04:09.687000",
          "content": "<p><a href=\"/alexandrecc\">@alexandrecc</a> I think pos_weight is better? First, from numerical values it correlates well and deviates little from LB, so this might actually be the official metric on LB. optimizing 'weight' instead of 'pos_weight' is actually optimizing a subtly different function. I am doing overfitting experiments right now to see whether local overfitting is reflected in my local validation and LB.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 657669,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2019-10-25T11:00:34.003000",
      "content": "<p>Here. I was not at my desktop PC earlier. If you are computing loss on GPU and the sixth class is 'any'.\nTypically, the input is e.g. [32, 3, 200, 200] and model output is [32, 6]. Target should be [32, 6].</p>\n\n<p><code>\nweights = torch.tensor([1.0, 1.0, 1.0, 1.0, 1.0, 2.0]).cuda()\ndef criterion(y_pred,y_true):\n    return torch.nn.functional.binary_cross_entropy_with_logits(y_pred,\n                               y_true, pos_weight=weights)\n</code></p>",
      "votes": 0,
      "replies": [
        {
          "id": 657803,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-10-25T12:38:51.927000",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> <br>\n  why have given more weight to class Any which is more than all others i suppose lowest is class 0  subtype.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 658461,
          "author_name": "Kalyan S.S",
          "author_url": "",
          "post_date": "2019-10-26T03:14:12.033000",
          "content": "<p>epsilon = 1e-7 <br>\npreds = torch.clamp(preds, epsilon, 1.0-epsilon).float()\nIs it right way to clamp?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 657633,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2019-10-25T10:08:51.477000",
      "content": "<p><a href=\"/roguekk007\">@roguekk007</a>  thanks for posting..\npos_weight would have how many weights here..,basically its dimension?\nany example which u can give ?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "657390": "I have spent a lot of time trying to establish a stable validation. Thanks to my teammate I fixed this problem this morning. Here is the loss function which I am currently using. Taking the mean over validation set should yield a very decent estimation of LB:\n`\nreturn torch.nn.functional.binary_cross_entropy_with_logits(y_pred, y_true, pos_weight=weights)\n`\n\nI used the argument 'weight' instead of 'pos_weight' for my previous models. There was a gap but overall the validation and LB correlates well. The gap is gone when I use the argument 'pos_weight' instead of 'weight'",
    "657854": "Interesting indeed! What do you think explains this difference ? ",
    "661636": "Hi @roguekk007 ,\n\nare you using this loss both for training and validation?\nDid you notice any improvement in LB when using this loss for training (in case you have experimented)?",
    "659099": "Is there any specific way, you had split validation data? I am doing random split of 10% data, and used loss as you said, but I am not getting decent estimate of LB.\n```\ndef validation(model,device,valid_loader):\n    model.eval()\n\n    class_weights = torch.tensor([2., 1., 1., 1., 1., 1.]).to(device,dtype=torch.float)\n\n    with torch.no_grad():\n        tr_loss = 0.0\n        for dev_batch_idx, dev_batch in enumerate(tqdm(valid_loader,desc=\"validation\")):\n            inputs = dev_batch[\"image\"].to(device,dtype=torch.float)\n            labels = dev_batch[\"labels\"].to(device,dtype=torch.float)\n            answer = model(inputs)\n            tr_loss += torch.nn.functional.binary_cross_entropy_with_logits(preds, trues, pos_weight=class_weights)\n\n    epoch_loss = tr_loss/len(valid_loader)\n    print(\"validation loss: %.4f\" % epoch_loss)\n    return epoch_loss\n```\n\nCan you check my implementation once?",
    "657674": "Using 'weight' argument for loss instead of 'posweight' always gives an overestimate of LB, but my model yields good LB nevertheless (high-silver range). Anyone who cares to explain the difference?",
    "657669": "Here. I was not at my desktop PC earlier. If you are computing loss on GPU and the sixth class is 'any'.\nTypically, the input is e.g. [32, 3, 200, 200] and model output is [32, 6]. Target should be [32, 6].\n\n```\nweights = torch.tensor([1.0, 1.0, 1.0, 1.0, 1.0, 2.0]).cuda()\ndef criterion(y_pred,y_true):\n    return torch.nn.functional.binary_cross_entropy_with_logits(y_pred,\n                               y_true, pos_weight=weights)\n```",
    "657633": "@roguekk007  thanks for posting..\npos_weight would have how many weights here..,basically its dimension?\nany example which u can give ?"
  }
}