{
  "id": 446760,
  "title": "Error in LB metric",
  "url": "/competitions/UBC-OCEAN/discussion/446760",
  "author_name": "David Austin",
  "post_date": "2023-10-13T00:11:24.779000",
  "votes": 35,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I believe there's an error in the LB metric.  <br>\nAccording to the data description there are 6 labels in the test set: <code>CC, EC, HGSC, LGSC, MC, Other</code>, so if we make a submission just predicting a single one of these classes for the whole test set, our score should be 0.167 (1/6 * (1+0+0+0+0+0)).  However, the score given for a single class prediction is 0.14 which is what we would expect for 7 classes, but not 6.  I think what may be happening is there's a 7th class in the GT data that's throwing things off.  <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> can you look into this?</p>",
  "messages": [
    {
      "id": 2479849,
      "postDate": "2023-10-13T00:11:24.780Z",
      "content": "<p>I believe there's an error in the LB metric.  <br>\nAccording to the data description there are 6 labels in the test set: <code>CC, EC, HGSC, LGSC, MC, Other</code>, so if we make a submission just predicting a single one of these classes for the whole test set, our score should be 0.167 (1/6 * (1+0+0+0+0+0)).  However, the score given for a single class prediction is 0.14 which is what we would expect for 7 classes, but not 6.  I think what may be happening is there's a 7th class in the GT data that's throwing things off.  <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> can you look into this?</p>",
      "rawMarkdown": "I believe there's an error in the LB metric.  \nAccording to the data description there are 6 labels in the test set: `CC, EC, HGSC, LGSC, MC, Other`, so if we make a submission just predicting a single one of these classes for the whole test set, our score should be 0.167 (1/6 * (1+0+0+0+0+0)).  However, the score given for a single class prediction is 0.14 which is what we would expect for 7 classes, but not 6.  I think what may be happening is there's a 7th class in the GT data that's throwing things off.  @sohier can you look into this?",
      "votes": 35
    },
    {
      "id": 2487662,
      "postDate": "2023-10-18T17:47:37.187Z",
      "content": "<p>Thank you for flagging this. There were a few labels where inconsistent capitalization caused them to be treated as a distinct class. I'll fix this now.</p>",
      "rawMarkdown": "Thank you for flagging this. There were a few labels where inconsistent capitalization caused them to be treated as a distinct class. I'll fix this now.",
      "votes": 7
    },
    {
      "id": 2482644,
      "postDate": "2023-10-15T07:47:21.800Z",
      "content": "<p>Nice catch. It's probably a typo on one of the labels.</p>",
      "rawMarkdown": "Nice catch. It's probably a typo on one of the labels.",
      "votes": 5
    },
    {
      "id": 2483015,
      "postDate": "2023-10-15T12:28:39.423Z",
      "content": "<p>Maybe this can also explain why the gap between CV and LB is so big. <br>\n<a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a> <a href=\"https://www.kaggle.com/homesmac\" target=\"_blank\">@homesmac</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Can you check it?</p>",
      "rawMarkdown": "Maybe this can also explain why the gap between CV and LB is so big. \n@ashleychow @homesmac @sohier Can you check it?",
      "votes": 6
    },
    {
      "id": 2483246,
      "postDate": "2023-10-15T15:30:31.783Z",
      "content": "<p>Maybe the distribution of classes in the test set is not uniform, and the proportion of that single class in the test set is 14%, that's why you're getting a score of 0.14</p>",
      "rawMarkdown": "Maybe the distribution of classes in the test set is not uniform, and the proportion of that single class in the test set is 14%, that's why you're getting a score of 0.14",
      "replies": [
        {
          "id": 2483249,
          "postDate": "2023-10-15T15:32:58.323Z",
          "content": "<p>It would be true for regular accuracy. For balanced accuracy, predicting only one class yields 1 / n_class score.</p>",
          "rawMarkdown": "It would be true for regular accuracy. For balanced accuracy, predicting only one class yields 1 / n_class score.",
          "votes": 8
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2487662,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2023-10-18T17:47:37.187000",
      "content": "<p>Thank you for flagging this. There were a few labels where inconsistent capitalization caused them to be treated as a distinct class. I'll fix this now.</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 2482644,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2023-10-15T07:47:21.800000",
      "content": "<p>Nice catch. It's probably a typo on one of the labels.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2483015,
      "author_name": "Johnny Lee",
      "author_url": "",
      "post_date": "2023-10-15T12:28:39.423000",
      "content": "<p>Maybe this can also explain why the gap between CV and LB is so big. <br>\n<a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a> <a href=\"https://www.kaggle.com/homesmac\" target=\"_blank\">@homesmac</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Can you check it?</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2483246,
      "author_name": "MD Mushfirat Mohaimin",
      "author_url": "",
      "post_date": "2023-10-15T15:30:31.783000",
      "content": "<p>Maybe the distribution of classes in the test set is not uniform, and the proportion of that single class in the test set is 14%, that's why you're getting a score of 0.14</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2483249,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2023-10-15T15:32:58.323000",
          "content": "<p>It would be true for regular accuracy. For balanced accuracy, predicting only one class yields 1 / n_class score.</p>",
          "votes": 8,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2479849": "I believe there's an error in the LB metric.  \nAccording to the data description there are 6 labels in the test set: `CC, EC, HGSC, LGSC, MC, Other`, so if we make a submission just predicting a single one of these classes for the whole test set, our score should be 0.167 (1/6 * (1+0+0+0+0+0)).  However, the score given for a single class prediction is 0.14 which is what we would expect for 7 classes, but not 6.  I think what may be happening is there's a 7th class in the GT data that's throwing things off.  @sohier can you look into this?",
    "2487662": "Thank you for flagging this. There were a few labels where inconsistent capitalization caused them to be treated as a distinct class. I'll fix this now.",
    "2482644": "Nice catch. It's probably a typo on one of the labels.",
    "2483015": "Maybe this can also explain why the gap between CV and LB is so big. \n@ashleychow @homesmac @sohier Can you check it?",
    "2483246": "Maybe the distribution of classes in the test set is not uniform, and the proportion of that single class in the test set is 14%, that's why you're getting a score of 0.14"
  }
}