{
  "id": 461431,
  "title": "Model only trains or predicts 'CC' for every scan",
  "url": "/competitions/UBC-OCEAN/discussion/461431",
  "author_name": "Varun Manoj Gupta",
  "post_date": "2023-12-14T12:23:24.620000",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I'm fairly new to python and machine learning, I have a problem where my model only seems to be either training or predicting for the label - 'CC',  I am using the label_encoder.pkl file shared in the top public scoring notebooks.</p>\n<p>I tested my theory on a partition of the training dataset which i used to validate, and it predicted cc for everything, and the public lb score wouldn't go above 0.16 for all submissions</p>\n<p>if my model only predicts the label cc for every image in the submission notebook, and it scored 0.16 everytime, does that mean that 16% of the public hidden test set images are cc (ik it can predict cc incorrectly too) but is my understanding correct here.</p>\n<p>Explanation of my understanding or a recommended fix would be appreciated👍  </p>",
  "messages": [
    {
      "id": 2561340,
      "postDate": "2023-12-14T12:23:24.620Z",
      "content": "<p>I'm fairly new to python and machine learning, I have a problem where my model only seems to be either training or predicting for the label - 'CC',  I am using the label_encoder.pkl file shared in the top public scoring notebooks.</p>\n<p>I tested my theory on a partition of the training dataset which i used to validate, and it predicted cc for everything, and the public lb score wouldn't go above 0.16 for all submissions</p>\n<p>if my model only predicts the label cc for every image in the submission notebook, and it scored 0.16 everytime, does that mean that 16% of the public hidden test set images are cc (ik it can predict cc incorrectly too) but is my understanding correct here.</p>\n<p>Explanation of my understanding or a recommended fix would be appreciated👍  </p>",
      "rawMarkdown": "I'm fairly new to python and machine learning, I have a problem where my model only seems to be either training or predicting for the label - 'CC',  I am using the label_encoder.pkl file shared in the top public scoring notebooks.\n\nI tested my theory on a partition of the training dataset which i used to validate, and it predicted cc for everything, and the public lb score wouldn't go above 0.16 for all submissions\n\nif my model only predicts the label cc for every image in the submission notebook, and it scored 0.16 everytime, does that mean that 16% of the public hidden test set images are cc (ik it can predict cc incorrectly too) but is my understanding correct here.\n\nExplanation of my understanding or a recommended fix would be appreciated👍  ",
      "votes": 3
    },
    {
      "id": 2570428,
      "postDate": "2023-12-22T07:58:48.530Z",
      "content": "<p>I guess this is because the evaluation metric is balanced accuracy (macro-averaged accuracy). If all samples are predicted 'CC', all samples in the 'CC' category are correctly predicted and all samples in other categories are predicted wrong. Since there are 6 classes in the competition, the LB score will be 1/6 $\\approx$ 0.16.</p>",
      "rawMarkdown": "I guess this is because the evaluation metric is balanced accuracy (macro-averaged accuracy). If all samples are predicted 'CC', all samples in the 'CC' category are correctly predicted and all samples in other categories are predicted wrong. Since there are 6 classes in the competition, the LB score will be 1/6 $\\approx$ 0.16.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2570428,
      "author_name": "Zijie Fang",
      "author_url": "",
      "post_date": "2023-12-22T07:58:48.530000",
      "content": "<p>I guess this is because the evaluation metric is balanced accuracy (macro-averaged accuracy). If all samples are predicted 'CC', all samples in the 'CC' category are correctly predicted and all samples in other categories are predicted wrong. Since there are 6 classes in the competition, the LB score will be 1/6 $\\approx$ 0.16.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2561340": "I'm fairly new to python and machine learning, I have a problem where my model only seems to be either training or predicting for the label - 'CC',  I am using the label_encoder.pkl file shared in the top public scoring notebooks.\n\nI tested my theory on a partition of the training dataset which i used to validate, and it predicted cc for everything, and the public lb score wouldn't go above 0.16 for all submissions\n\nif my model only predicts the label cc for every image in the submission notebook, and it scored 0.16 everytime, does that mean that 16% of the public hidden test set images are cc (ik it can predict cc incorrectly too) but is my understanding correct here.\n\nExplanation of my understanding or a recommended fix would be appreciated👍  ",
    "2570428": "I guess this is because the evaluation metric is balanced accuracy (macro-averaged accuracy). If all samples are predicted 'CC', all samples in the 'CC' category are correctly predicted and all samples in other categories are predicted wrong. Since there are 6 classes in the competition, the LB score will be 1/6 $\\approx$ 0.16."
  }
}