{
  "id": 463245,
  "title": "[Solved] Loss isn't correlated with the metric too much",
  "url": "/competitions/UBC-OCEAN/discussion/463245",
  "author_name": "Gunes Evitan",
  "post_date": "2023-12-24T06:12:37.603000",
  "votes": 11,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Does anyone else notice this? My best scoring checkpoints are very rarely have the best loss. I guess we can explain that by the nature of an unstable metric and use the best scoring checkpoints.</p>",
  "messages": [
    {
      "id": 2572332,
      "postDate": "2023-12-24T06:12:37.603Z",
      "content": "<p>Does anyone else notice this? My best scoring checkpoints are very rarely have the best loss. I guess we can explain that by the nature of an unstable metric and use the best scoring checkpoints.</p>",
      "rawMarkdown": "Does anyone else notice this? My best scoring checkpoints are very rarely have the best loss. I guess we can explain that by the nature of an unstable metric and use the best scoring checkpoints.",
      "votes": 11
    },
    {
      "id": 2580897,
      "postDate": "2023-12-31T10:57:47.837Z",
      "content": "<p>I solved this by increasing data points in validation sets. Epochs with best validation loss also have the best balanced accuracy now. </p>",
      "rawMarkdown": "I solved this by increasing data points in validation sets. Epochs with best validation loss also have the best balanced accuracy now. ",
      "votes": 5
    },
    {
      "id": 2573205,
      "postDate": "2023-12-24T19:30:58.177Z",
      "content": "<p>It is hard to believe the organisers spent valuable cancer research money for organising competition with  improper scoring rule as a metric. Metrics such as 'balanced accuracy' are not proper scoring rules and don't lead models to well calibrated predictions! The results of putting such models into real life setting will result in incorrect probabilities of classifying cancer!</p>",
      "rawMarkdown": "It is hard to believe the organisers spent valuable cancer research money for organising competition with  improper scoring rule as a metric. Metrics such as 'balanced accuracy' are not proper scoring rules and don't lead models to well calibrated predictions! The results of putting such models into real life setting will result in incorrect probabilities of classifying cancer!",
      "votes": 4,
      "replies": [
        {
          "id": 2575138,
          "postDate": "2023-12-26T15:55:35.123Z",
          "content": "<p>Metrics like macro F1 score would be unstable like this one. What do you suggest? Log loss?</p>",
          "rawMarkdown": "Metrics like macro F1 score would be unstable like this one. What do you suggest? Log loss?",
          "votes": 1,
          "replies": [
            {
              "id": 2575211,
              "postDate": "2023-12-26T17:08:17.673Z",
              "content": "<p>Dear <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> </p>\n<p>The log-loss  would be an obvious choice as this is what most classifiers minimize and constitutes a  <a href=\"https://sites.stat.washington.edu/raftery/Research/PDF/Gneiting2007jasa.pdf\" target=\"_blank\">strictly proper scoring rule</a>, as does the <a href=\"https://en.wikipedia.org/wiki/Brier_score#Original_definition_by_Brier\" target=\"_blank\">multi-class Brier score</a>. Indeed the problem of correctly assessing  class predictions goes back well over 70 years and was effectively addressed by Brier in the paper <a href=\"https://journals.ametsoc.org/downloadpdf/journals/mwre/78/1/1520-0493_1950_078_0001_vofeit_2_0_co_2.xml\" target=\"_blank\">Verification of Forecasts Expressed in Terms of Probability</a> in 1950 and is well worth reading.</p>\n<p>All the best,<br>\ncarl</p>",
              "rawMarkdown": "Dear @gunesevitan \n\nThe log-loss  would be an obvious choice as this is what most classifiers minimize and constitutes a  [strictly proper scoring rule](https://sites.stat.washington.edu/raftery/Research/PDF/Gneiting2007jasa.pdf), as does the [multi-class Brier score](https://en.wikipedia.org/wiki/Brier_score#Original_definition_by_Brier). Indeed the problem of correctly assessing  class predictions goes back well over 70 years and was effectively addressed by Brier in the paper [Verification of Forecasts Expressed in Terms of Probability](https://journals.ametsoc.org/downloadpdf/journals/mwre/78/1/1520-0493_1950_078_0001_vofeit_2_0_co_2.xml) in 1950 and is well worth reading.\n\nAll the best,\ncarl",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2572337,
      "postDate": "2023-12-24T06:26:25.687Z",
      "content": "<p>LOL; yet another classification competition that does not use a <a href=\"https://en.wikipedia.org/wiki/Scoring_rule\" target=\"_blank\">proper scoring rule</a>….</p>\n<p>All the best,<br>\ncarl </p>",
      "rawMarkdown": "LOL; yet another classification competition that does not use a [proper scoring rule](https://en.wikipedia.org/wiki/Scoring_rule)....\n\nAll the best,\ncarl ",
      "votes": 1
    },
    {
      "id": 2572457,
      "postDate": "2023-12-24T09:08:07.473Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> , I think if you set a proper weight of the loss, they will have better correlation.</p>",
      "rawMarkdown": "Hi @gunesevitan , I think if you set a proper weight of the loss, they will have better correlation.",
      "votes": 1,
      "replies": [
        {
          "id": 2572484,
          "postDate": "2023-12-24T09:41:32.117Z",
          "content": "<p>I tried weighted sampling but it didn't work. I'll try weighted loss too but based on my understanding of the metric, isn't each sample should have equal weights?</p>",
          "rawMarkdown": "I tried weighted sampling but it didn't work. I'll try weighted loss too but based on my understanding of the metric, isn't each sample should have equal weights?",
          "replies": [
            {
              "id": 2572664,
              "postDate": "2023-12-24T12:50:03.787Z",
              "content": "<p>I think weight loss correlates more with bacc. For example, given there are 200 <code>HGSC</code> and <code>40</code> LGSC, when you use mean loss,  if you make 1 more correct predictions for <code>LGSC</code> and 3 more wrong predictions for <code>HGSC</code>, the mean loss is worse while bacc is better. </p>",
              "rawMarkdown": "I think weight loss correlates more with bacc. For example, given there are 200 `HGSC` and `40` LGSC, when you use mean loss,  if you make 1 more correct predictions for `LGSC` and 3 more wrong predictions for `HGSC`, the mean loss is worse while bacc is better. ",
              "votes": 6
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2580897,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2023-12-31T10:57:47.837000",
      "content": "<p>I solved this by increasing data points in validation sets. Epochs with best validation loss also have the best balanced accuracy now. </p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2573205,
      "author_name": "predict_addict",
      "author_url": "",
      "post_date": "2023-12-24T19:30:58.177000",
      "content": "<p>It is hard to believe the organisers spent valuable cancer research money for organising competition with  improper scoring rule as a metric. Metrics such as 'balanced accuracy' are not proper scoring rules and don't lead models to well calibrated predictions! The results of putting such models into real life setting will result in incorrect probabilities of classifying cancer!</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2575138,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2023-12-26T15:55:35.123000",
          "content": "<p>Metrics like macro F1 score would be unstable like this one. What do you suggest? Log loss?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2575211,
              "author_name": "Carl McBride Ellis",
              "author_url": "",
              "post_date": "2023-12-26T17:08:17.673000",
              "content": "<p>Dear <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> </p>\n<p>The log-loss  would be an obvious choice as this is what most classifiers minimize and constitutes a  <a href=\"https://sites.stat.washington.edu/raftery/Research/PDF/Gneiting2007jasa.pdf\" target=\"_blank\">strictly proper scoring rule</a>, as does the <a href=\"https://en.wikipedia.org/wiki/Brier_score#Original_definition_by_Brier\" target=\"_blank\">multi-class Brier score</a>. Indeed the problem of correctly assessing  class predictions goes back well over 70 years and was effectively addressed by Brier in the paper <a href=\"https://journals.ametsoc.org/downloadpdf/journals/mwre/78/1/1520-0493_1950_078_0001_vofeit_2_0_co_2.xml\" target=\"_blank\">Verification of Forecasts Expressed in Terms of Probability</a> in 1950 and is well worth reading.</p>\n<p>All the best,<br>\ncarl</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2572337,
      "author_name": "Carl McBride Ellis",
      "author_url": "",
      "post_date": "2023-12-24T06:26:25.687000",
      "content": "<p>LOL; yet another classification competition that does not use a <a href=\"https://en.wikipedia.org/wiki/Scoring_rule\" target=\"_blank\">proper scoring rule</a>….</p>\n<p>All the best,<br>\ncarl </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2572457,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2023-12-24T09:08:07.473000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> , I think if you set a proper weight of the loss, they will have better correlation.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2572484,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2023-12-24T09:41:32.117000",
          "content": "<p>I tried weighted sampling but it didn't work. I'll try weighted loss too but based on my understanding of the metric, isn't each sample should have equal weights?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2572664,
              "author_name": "ForcewithMe",
              "author_url": "",
              "post_date": "2023-12-24T12:50:03.787000",
              "content": "<p>I think weight loss correlates more with bacc. For example, given there are 200 <code>HGSC</code> and <code>40</code> LGSC, when you use mean loss,  if you make 1 more correct predictions for <code>LGSC</code> and 3 more wrong predictions for <code>HGSC</code>, the mean loss is worse while bacc is better. </p>",
              "votes": 6,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2572332": "Does anyone else notice this? My best scoring checkpoints are very rarely have the best loss. I guess we can explain that by the nature of an unstable metric and use the best scoring checkpoints.",
    "2580897": "I solved this by increasing data points in validation sets. Epochs with best validation loss also have the best balanced accuracy now. ",
    "2573205": "It is hard to believe the organisers spent valuable cancer research money for organising competition with  improper scoring rule as a metric. Metrics such as 'balanced accuracy' are not proper scoring rules and don't lead models to well calibrated predictions! The results of putting such models into real life setting will result in incorrect probabilities of classifying cancer!",
    "2572337": "LOL; yet another classification competition that does not use a [proper scoring rule](https://en.wikipedia.org/wiki/Scoring_rule)....\n\nAll the best,\ncarl ",
    "2572457": "Hi @gunesevitan , I think if you set a proper weight of the loss, they will have better correlation."
  }
}