{
  "id": 153525,
  "title": "Optimize Thresholds for Kappa Score (Regression)",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/153525",
  "author_name": "Claudio Fanconi",
  "post_date": "2020-05-25T07:33:58.879000",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I was wondering:\nIf you guys are using a regression model, do you optimize the threshold boundaries for improved KAPPA on a validation set?</p>\n\n<p>This will obviously yield you better results on a validation set, but isn't it skewed for the distribution of only the single set?\nMathematically, using an MSE loss, I would believe that a model is optimized for having a natural threshold in the middle of every integer (0.5,1.5,2.5.3.5).</p>\n\n<p>Can you convince me otherwise, why optimizing the threshold will not lead to worse results in the test set?</p>\n\n<p>Thank you :)</p>",
  "messages": [
    {
      "id": 860312,
      "postDate": "2020-05-25T07:33:58.880Z",
      "content": "<p>I was wondering:\nIf you guys are using a regression model, do you optimize the threshold boundaries for improved KAPPA on a validation set?</p>\n\n<p>This will obviously yield you better results on a validation set, but isn't it skewed for the distribution of only the single set?\nMathematically, using an MSE loss, I would believe that a model is optimized for having a natural threshold in the middle of every integer (0.5,1.5,2.5.3.5).</p>\n\n<p>Can you convince me otherwise, why optimizing the threshold will not lead to worse results in the test set?</p>\n\n<p>Thank you :)</p>",
      "rawMarkdown": "I was wondering:\nIf you guys are using a regression model, do you optimize the threshold boundaries for improved KAPPA on a validation set?\n\nThis will obviously yield you better results on a validation set, but isn't it skewed for the distribution of only the single set?\nMathematically, using an MSE loss, I would believe that a model is optimized for having a natural threshold in the middle of every integer (0.5,1.5,2.5.3.5).\n\nCan you convince me otherwise, why optimizing the threshold will not lead to worse results in the test set?\n\nThank you :)",
      "votes": 5
    },
    {
      "id": 861049,
      "postDate": "2020-05-25T19:53:49.020Z",
      "content": "<p>I'm actually curious about this too. The easiest is probably to test it directly. Compute and save optimized threshold at your checkpoints then do 2 kaggle submissions.</p>",
      "rawMarkdown": "I'm actually curious about this too. The easiest is probably to test it directly. Compute and save optimized threshold at your checkpoints then do 2 kaggle submissions.",
      "votes": 1,
      "replies": [
        {
          "id": 861053,
          "postDate": "2020-05-25T19:58:43.163Z",
          "content": "<p>I did some digging in the APTOS 2019 challenge and came across the answer, that optimized threshold makes sense if the classes are imbalanced. Otherwise, if the classes are equally distributed, then .5 should do the job.</p>\n\n<p>I tried both approaches, and the one with optimized thresholds lead to slightly better results.</p>\n\n<p>Without optimization: 0.75\nWith optimization: 0.76</p>",
          "rawMarkdown": "I did some digging in the APTOS 2019 challenge and came across the answer, that optimized threshold makes sense if the classes are imbalanced. Otherwise, if the classes are equally distributed, then .5 should do the job.\n\nI tried both approaches, and the one with optimized thresholds lead to slightly better results.\n\nWithout optimization: 0.75\nWith optimization: 0.76",
          "votes": 2
        },
        {
          "id": 861083,
          "postDate": "2020-05-25T20:44:25.020Z",
          "content": "<p>You should optimize your threshold on the train data otherwise your CV score might be higher than it should be compared to LB. </p>",
          "rawMarkdown": "You should optimize your threshold on the train data otherwise your CV score might be higher than it should be compared to LB. ",
          "votes": 1
        },
        {
          "id": 861169,
          "postDate": "2020-05-25T23:25:25.070Z",
          "content": "<p>I did indeed have this experience: my CV yielded 0.8, but LB only 0.76...</p>\n\n<p>How exactly do you mean optimizer the thresholds on the training data? Won't I just receive an even more divergent score from LB, as I am using trained data? Anyway, thank you for the idea! </p>",
          "rawMarkdown": "I did indeed have this experience: my CV yielded 0.8, but LB only 0.76...\n\nHow exactly do you mean optimizer the thresholds on the training data? Won't I just receive an even more divergent score from LB, as I am using trained data? Anyway, thank you for the idea! "
        },
        {
          "id": 861250,
          "postDate": "2020-05-26T00:34:19.747Z",
          "content": "<p>here is an example in the train loop: </p>\n\n<p>```\nfor .... :\n      output=model(images)\n      preds.append(output.cpu().detach().numpy())\n      train_labels.append(target.cpu().detach().numpy())</p>\n\n<p>preds = np.concatenate(preds)\n train_labels = np.concatenate(train_labels)\n optR.fit(preds, train_labels)\n coefficients = optR.coefficients()\n```</p>\n\n<p>so you take all the predictions and the labels of the epoch to fit the coefficients and optimize the qwk score. Then use those coefficients in your validation phase like:</p>\n\n<p><code>final_preds = optR.predict(preds, coefficients)</code></p>",
          "rawMarkdown": "here is an example in the train loop: \n \n```\nfor .... :\n      output=model(images)\n      preds.append(output.cpu().detach().numpy())\n      train_labels.append(target.cpu().detach().numpy())\n\n preds = np.concatenate(preds)\n train_labels = np.concatenate(train_labels)\n optR.fit(preds, train_labels)\n coefficients = optR.coefficients()\n```\n\nso you take all the predictions and the labels of the epoch to fit the coefficients and optimize the qwk score. Then use those coefficients in your validation phase like:\n\n`final_preds = optR.predict(preds, coefficients)`",
          "votes": 1
        },
        {
          "id": 865776,
          "postDate": "2020-05-28T21:48:21.460Z",
          "content": "<p>I tested it, and it improved my score a tiny bit. \nCV from 0.8534 to 0.8550 and LB also improved (0.85 to 0.85 but higher).\nI ended up optimizing for each of the 4 folds and then averaging the thresholds to have the final thresholds to use.</p>\n\n<p>In short it improves, but probably not worth the effort.</p>",
          "rawMarkdown": "I tested it, and it improved my score a tiny bit. \nCV from 0.8534 to 0.8550 and LB also improved (0.85 to 0.85 but higher).\nI ended up optimizing for each of the 4 folds and then averaging the thresholds to have the final thresholds to use.\n\nIn short it improves, but probably not worth the effort.",
          "votes": 3
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 861049,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2020-05-25T19:53:49.020000",
      "content": "<p>I'm actually curious about this too. The easiest is probably to test it directly. Compute and save optimized threshold at your checkpoints then do 2 kaggle submissions.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 861053,
          "author_name": "Claudio Fanconi",
          "author_url": "",
          "post_date": "2020-05-25T19:58:43.163000",
          "content": "<p>I did some digging in the APTOS 2019 challenge and came across the answer, that optimized threshold makes sense if the classes are imbalanced. Otherwise, if the classes are equally distributed, then .5 should do the job.</p>\n\n<p>I tried both approaches, and the one with optimized thresholds lead to slightly better results.</p>\n\n<p>Without optimization: 0.75\nWith optimization: 0.76</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 861083,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-05-25T20:44:25.020000",
          "content": "<p>You should optimize your threshold on the train data otherwise your CV score might be higher than it should be compared to LB. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 861169,
          "author_name": "Claudio Fanconi",
          "author_url": "",
          "post_date": "2020-05-25T23:25:25.070000",
          "content": "<p>I did indeed have this experience: my CV yielded 0.8, but LB only 0.76...</p>\n\n<p>How exactly do you mean optimizer the thresholds on the training data? Won't I just receive an even more divergent score from LB, as I am using trained data? Anyway, thank you for the idea! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 861250,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-05-26T00:34:19.747000",
          "content": "<p>here is an example in the train loop: </p>\n\n<p>```\nfor .... :\n      output=model(images)\n      preds.append(output.cpu().detach().numpy())\n      train_labels.append(target.cpu().detach().numpy())</p>\n\n<p>preds = np.concatenate(preds)\n train_labels = np.concatenate(train_labels)\n optR.fit(preds, train_labels)\n coefficients = optR.coefficients()\n```</p>\n\n<p>so you take all the predictions and the labels of the epoch to fit the coefficients and optimize the qwk score. Then use those coefficients in your validation phase like:</p>\n\n<p><code>final_preds = optR.predict(preds, coefficients)</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 865776,
          "author_name": "Peter Cnudde",
          "author_url": "",
          "post_date": "2020-05-28T21:48:21.460000",
          "content": "<p>I tested it, and it improved my score a tiny bit. \nCV from 0.8534 to 0.8550 and LB also improved (0.85 to 0.85 but higher).\nI ended up optimizing for each of the 4 folds and then averaging the thresholds to have the final thresholds to use.</p>\n\n<p>In short it improves, but probably not worth the effort.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "860312": "I was wondering:\nIf you guys are using a regression model, do you optimize the threshold boundaries for improved KAPPA on a validation set?\n\nThis will obviously yield you better results on a validation set, but isn't it skewed for the distribution of only the single set?\nMathematically, using an MSE loss, I would believe that a model is optimized for having a natural threshold in the middle of every integer (0.5,1.5,2.5.3.5).\n\nCan you convince me otherwise, why optimizing the threshold will not lead to worse results in the test set?\n\nThank you :)",
    "861049": "I'm actually curious about this too. The easiest is probably to test it directly. Compute and save optimized threshold at your checkpoints then do 2 kaggle submissions."
  }
}