{
  "id": 114087,
  "title": "Relationship between accuracy and loss?",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/114087",
  "author_name": "Ryan Epp",
  "post_date": "2019-10-24T05:17:53.644000",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I've noticed something odd when training on the full dataset without any over/under sampling. When I use a learning rate that's too large, the train accuracy improves at the expense of the loss and the validation loss/acc seems more stable.</p>\n\n<p>I did a quick experiment, here are the first 4 epoch of two training runs over the full dataset that are completely identical except for the lr schedule.</p>\n\n<h3>Run 1 - Larger LR, Worse loss, Better Acc, More Stable Validation</h3>\n\n<p>| Epoch | lr    | Train Loss | Train Acc | Val Loss | Val Acc |\n|-------|-------|------------|-----------|----------|---------|\n| 1     | .005  | 0.1350     | 0.8345    | 0.1247   | 0.6262  |\n| 2     | .003  | 0.1072     | 0.8407    | 0.1079   | 0.8345  |\n| 3     | .001  | 0.922      | 0.8538    | 0.1030   | 0.9562  |\n| 4     | .0005 | 0.0854     | 0.8650    | 0.0951   | 0.9450  |</p>\n\n<h3>Run 2 - Smaller LR, Better loss, Worse Acc, Less Stable Validation</h3>\n\n<p>| Epoch | lr    | Train Loss | Train Acc | Val Loss | Val Acc |\n|-------|-------|------------|-----------|----------|---------|\n| 1     | .001  | 0.1231     | 0.7911    | 0.1452   | 0.6319  |\n| 2     | .0003 | 0.0923     | 0.8318    | 0.1221   | 0.8623  |\n| 3     | .0001 | 0.0821     | 0.8331    | 0.0946   | 0.8361  |\n| 4     | .0001 | 0.0784     | 0.8262    | 0.1120   | 0.7661  |</p>\n\n<p>Run 2 converges faster, so (as I understand it), run 2 has a better lr, but it does worse on val loss/acc. My guess is that because of the class imbalance, run 2 is able to overfit to the negative/majority examples to drive down the loss while Run 1's higher lr makes it more sensitive to positive/minority examples. </p>\n\n<p>Does that make sense? Am I misunderstanding anything? Does anybody else have any ideas?</p>\n\n<p>Obviously there are better ways to deal with class imbalance, but setting the lr too high as a form of regularization seems mildly interesting?</p>",
  "messages": [
    {
      "id": 656279,
      "postDate": "2019-10-24T05:17:53.643Z",
      "content": "<p>I've noticed something odd when training on the full dataset without any over/under sampling. When I use a learning rate that's too large, the train accuracy improves at the expense of the loss and the validation loss/acc seems more stable.</p>\n\n<p>I did a quick experiment, here are the first 4 epoch of two training runs over the full dataset that are completely identical except for the lr schedule.</p>\n\n<h3>Run 1 - Larger LR, Worse loss, Better Acc, More Stable Validation</h3>\n\n<p>| Epoch | lr    | Train Loss | Train Acc | Val Loss | Val Acc |\n|-------|-------|------------|-----------|----------|---------|\n| 1     | .005  | 0.1350     | 0.8345    | 0.1247   | 0.6262  |\n| 2     | .003  | 0.1072     | 0.8407    | 0.1079   | 0.8345  |\n| 3     | .001  | 0.922      | 0.8538    | 0.1030   | 0.9562  |\n| 4     | .0005 | 0.0854     | 0.8650    | 0.0951   | 0.9450  |</p>\n\n<h3>Run 2 - Smaller LR, Better loss, Worse Acc, Less Stable Validation</h3>\n\n<p>| Epoch | lr    | Train Loss | Train Acc | Val Loss | Val Acc |\n|-------|-------|------------|-----------|----------|---------|\n| 1     | .001  | 0.1231     | 0.7911    | 0.1452   | 0.6319  |\n| 2     | .0003 | 0.0923     | 0.8318    | 0.1221   | 0.8623  |\n| 3     | .0001 | 0.0821     | 0.8331    | 0.0946   | 0.8361  |\n| 4     | .0001 | 0.0784     | 0.8262    | 0.1120   | 0.7661  |</p>\n\n<p>Run 2 converges faster, so (as I understand it), run 2 has a better lr, but it does worse on val loss/acc. My guess is that because of the class imbalance, run 2 is able to overfit to the negative/majority examples to drive down the loss while Run 1's higher lr makes it more sensitive to positive/minority examples. </p>\n\n<p>Does that make sense? Am I misunderstanding anything? Does anybody else have any ideas?</p>\n\n<p>Obviously there are better ways to deal with class imbalance, but setting the lr too high as a form of regularization seems mildly interesting?</p>",
      "rawMarkdown": "I've noticed something odd when training on the full dataset without any over/under sampling. When I use a learning rate that's too large, the train accuracy improves at the expense of the loss and the validation loss/acc seems more stable.\n\nI did a quick experiment, here are the first 4 epoch of two training runs over the full dataset that are completely identical except for the lr schedule.\n\n### Run 1 - Larger LR, Worse loss, Better Acc, More Stable Validation\n| Epoch | lr    | Train Loss | Train Acc | Val Loss | Val Acc |\n|-------|-------|------------|-----------|----------|---------|\n| 1     | .005  | 0.1350     | 0.8345    | 0.1247   | 0.6262  |\n| 2     | .003  | 0.1072     | 0.8407    | 0.1079   | 0.8345  |\n| 3     | .001  | 0.922      | 0.8538    | 0.1030   | 0.9562  |\n| 4     | .0005 | 0.0854     | 0.8650    | 0.0951   | 0.9450  |\n\n### Run 2 - Smaller LR, Better loss, Worse Acc, Less Stable Validation\n| Epoch | lr    | Train Loss | Train Acc | Val Loss | Val Acc |\n|-------|-------|------------|-----------|----------|---------|\n| 1     | .001  | 0.1231     | 0.7911    | 0.1452   | 0.6319  |\n| 2     | .0003 | 0.0923     | 0.8318    | 0.1221   | 0.8623  |\n| 3     | .0001 | 0.0821     | 0.8331    | 0.0946   | 0.8361  |\n| 4     | .0001 | 0.0784     | 0.8262    | 0.1120   | 0.7661  |\n\nRun 2 converges faster, so (as I understand it), run 2 has a better lr, but it does worse on val loss/acc. My guess is that because of the class imbalance, run 2 is able to overfit to the negative/majority examples to drive down the loss while Run 1's higher lr makes it more sensitive to positive/minority examples. \n\nDoes that make sense? Am I misunderstanding anything? Does anybody else have any ideas?\n\nObviously there are better ways to deal with class imbalance, but setting the lr too high as a form of regularization seems mildly interesting?\n\n",
      "votes": 5
    },
    {
      "id": 656285,
      "postDate": "2019-10-24T05:30:11.977Z",
      "content": "<p>How do you define the accuracy? Have you selected a threshold? Is it different per class? Because of the problems underlined by these questions, I think AUC ROC is a better proxy metric for accuracy here.</p>\n\n<p>Additionally, I wouldn't give much importance to validation loss or accuracy at the first epochs, it can be highly unstable (it jumps a lot for me too). Keep training for more epochs. </p>",
      "rawMarkdown": "How do you define the accuracy? Have you selected a threshold? Is it different per class? Because of the problems underlined by these questions, I think AUC ROC is a better proxy metric for accuracy here.\n\nAdditionally, I wouldn't give much importance to validation loss or accuracy at the first epochs, it can be highly unstable (it jumps a lot for me too). Keep training for more epochs. ",
      "votes": 2,
      "replies": [
        {
          "id": 656287,
          "postDate": "2019-10-24T05:42:58.927Z",
          "content": "<p>Thanks nosound! This is super helpful. I've just been using the default Keras accuracy metric which I assume is the average accuracy of all 6 classes with a threshold at 0.5.</p>",
          "rawMarkdown": "Thanks nosound! This is super helpful. I've just been using the default Keras accuracy metric which I assume is the average accuracy of all 6 classes with a threshold at 0.5.",
          "votes": 1
        },
        {
          "id": 656335,
          "postDate": "2019-10-24T06:48:31.203Z",
          "content": "<p>Yea, the prior probabilities of classes are in range <code>0.004 - 0.144</code> in train, so threshold <code>0.5</code> is inappropriate. Use <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.roc_auc_score.html\">this</a>.</p>",
          "rawMarkdown": "Yea, the prior probabilities of classes are in range `0.004 - 0.144` in train, so threshold `0.5` is inappropriate. Use [this](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.roc_auc_score.html).",
          "votes": 2
        },
        {
          "id": 656367,
          "postDate": "2019-10-24T07:09:05.107Z",
          "content": "<p>I am still curious as to why the organisers didn't choose ROC AUC as the metric. This would have made it easier to compare the results to other studies and added an interesting twist to the challenge</p>",
          "rawMarkdown": "I am still curious as to why the organisers didn't choose ROC AUC as the metric. This would have made it easier to compare the results to other studies and added an interesting twist to the challenge",
          "votes": 1
        },
        {
          "id": 656396,
          "postDate": "2019-10-24T07:52:21.310Z",
          "content": "<p>or how about f1 score?</p>",
          "rawMarkdown": "or how about f1 score?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 656285,
      "author_name": "nosound",
      "author_url": "",
      "post_date": "2019-10-24T05:30:11.977000",
      "content": "<p>How do you define the accuracy? Have you selected a threshold? Is it different per class? Because of the problems underlined by these questions, I think AUC ROC is a better proxy metric for accuracy here.</p>\n\n<p>Additionally, I wouldn't give much importance to validation loss or accuracy at the first epochs, it can be highly unstable (it jumps a lot for me too). Keep training for more epochs. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 656287,
          "author_name": "Ryan Epp",
          "author_url": "",
          "post_date": "2019-10-24T05:42:58.927000",
          "content": "<p>Thanks nosound! This is super helpful. I've just been using the default Keras accuracy metric which I assume is the average accuracy of all 6 classes with a threshold at 0.5.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 656335,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2019-10-24T06:48:31.203000",
          "content": "<p>Yea, the prior probabilities of classes are in range <code>0.004 - 0.144</code> in train, so threshold <code>0.5</code> is inappropriate. Use <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.roc_auc_score.html\">this</a>.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 656367,
          "author_name": "datasaurus",
          "author_url": "",
          "post_date": "2019-10-24T07:09:05.107000",
          "content": "<p>I am still curious as to why the organisers didn't choose ROC AUC as the metric. This would have made it easier to compare the results to other studies and added an interesting twist to the challenge</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 656396,
          "author_name": "Verne",
          "author_url": "",
          "post_date": "2019-10-24T07:52:21.310000",
          "content": "<p>or how about f1 score?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "656279": "I've noticed something odd when training on the full dataset without any over/under sampling. When I use a learning rate that's too large, the train accuracy improves at the expense of the loss and the validation loss/acc seems more stable.\n\nI did a quick experiment, here are the first 4 epoch of two training runs over the full dataset that are completely identical except for the lr schedule.\n\n### Run 1 - Larger LR, Worse loss, Better Acc, More Stable Validation\n| Epoch | lr    | Train Loss | Train Acc | Val Loss | Val Acc |\n|-------|-------|------------|-----------|----------|---------|\n| 1     | .005  | 0.1350     | 0.8345    | 0.1247   | 0.6262  |\n| 2     | .003  | 0.1072     | 0.8407    | 0.1079   | 0.8345  |\n| 3     | .001  | 0.922      | 0.8538    | 0.1030   | 0.9562  |\n| 4     | .0005 | 0.0854     | 0.8650    | 0.0951   | 0.9450  |\n\n### Run 2 - Smaller LR, Better loss, Worse Acc, Less Stable Validation\n| Epoch | lr    | Train Loss | Train Acc | Val Loss | Val Acc |\n|-------|-------|------------|-----------|----------|---------|\n| 1     | .001  | 0.1231     | 0.7911    | 0.1452   | 0.6319  |\n| 2     | .0003 | 0.0923     | 0.8318    | 0.1221   | 0.8623  |\n| 3     | .0001 | 0.0821     | 0.8331    | 0.0946   | 0.8361  |\n| 4     | .0001 | 0.0784     | 0.8262    | 0.1120   | 0.7661  |\n\nRun 2 converges faster, so (as I understand it), run 2 has a better lr, but it does worse on val loss/acc. My guess is that because of the class imbalance, run 2 is able to overfit to the negative/majority examples to drive down the loss while Run 1's higher lr makes it more sensitive to positive/minority examples. \n\nDoes that make sense? Am I misunderstanding anything? Does anybody else have any ideas?\n\nObviously there are better ways to deal with class imbalance, but setting the lr too high as a form of regularization seems mildly interesting?\n\n",
    "656285": "How do you define the accuracy? Have you selected a threshold? Is it different per class? Because of the problems underlined by these questions, I think AUC ROC is a better proxy metric for accuracy here.\n\nAdditionally, I wouldn't give much importance to validation loss or accuracy at the first epochs, it can be highly unstable (it jumps a lot for me too). Keep training for more epochs. "
  }
}