{
  "id": 384713,
  "title": "Best threshold is too low (0.04)",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/384713",
  "author_name": "Ju7on9",
  "post_date": "2023-02-09T06:47:16.932000",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello, </p>\n<p>I trained the model, but the best threshold is too low.</p>\n<p>I saw other people's result, but their threshold is almost 0.3~0.5.</p>\n<p>So, I think something is wrong in my code.</p>\n<p>Do you have any idea what and when causes the low threshold?</p>\n<p>At best threshold, my model  yields the following metric</p>\n<blockquote>\n  <p>val_pF1: 0.3170 - val_f1_score: 0.3370 - val_precision: 0.5897 - val_recall: 0.2359 - val_auc: 0.6855 - val_binary_accuracy: 0.9774</p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5286406%2F5e0ced3d5f1da65a979d86035515c60e%2F2.PNG?generation=1675925073648187&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2136198,
      "postDate": "2023-02-09T06:47:16.933Z",
      "content": "<p>Hello, </p>\n<p>I trained the model, but the best threshold is too low.</p>\n<p>I saw other people's result, but their threshold is almost 0.3~0.5.</p>\n<p>So, I think something is wrong in my code.</p>\n<p>Do you have any idea what and when causes the low threshold?</p>\n<p>At best threshold, my model  yields the following metric</p>\n<blockquote>\n  <p>val_pF1: 0.3170 - val_f1_score: 0.3370 - val_precision: 0.5897 - val_recall: 0.2359 - val_auc: 0.6855 - val_binary_accuracy: 0.9774</p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5286406%2F5e0ced3d5f1da65a979d86035515c60e%2F2.PNG?generation=1675925073648187&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hello, \n\nI trained the model, but the best threshold is too low.\n\nI saw other people's result, but their threshold is almost 0.3~0.5.\n\nSo, I think something is wrong in my code.\n\nDo you have any idea what and when causes the low threshold?\n\nAt best threshold, my model  yields the following metric\n\n>val_pF1: 0.3170 - val_f1_score: 0.3370 - val_precision: 0.5897 - val_recall: 0.2359 - val_auc: 0.6855 - val_binary_accuracy: 0.9774\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5286406%2F5e0ced3d5f1da65a979d86035515c60e%2F2.PNG?generation=1675925073648187&alt=media)",
      "votes": 7
    },
    {
      "id": 2136332,
      "postDate": "2023-02-09T08:52:39.607Z",
      "content": "<p>Keep in mind the validation set consists of just 20% of the training data with ~200 positive samples. The best threshold is likely to be overfitted on this small validation set with limited number of positive samples. Additionally, the F1 score by threshold line is almost flat with no clear peak, further indicating the best threshold is likely overfitting.</p>\n<p>Moreover, the validation predictions are overconfident, since all predictions are near 0/1. The threshold does not influence the binary predictions much, since the threshold will change the predicted value of just a few samples when changed (there are almost no samples with a predicted value between 0.20-0.80).</p>\n<p>The threshold would only influence the F1 score when many predictions are not near 0/1.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F0711dee7b2441c64c999f802ea92932e%2Ftemp.png?generation=1675932274603709&amp;alt=media\" alt=\"\"></p>\n<p>You have a few options:</p>\n<p>1) stick to a threshold of 0.50<br>\n2) Performs a full 5-Fold run and determine the best out-of-fold best threshold on the full dataset<br>\n3) Use a different training strategy (i.e. loss function) to have less confident predictions</p>",
      "rawMarkdown": "Keep in mind the validation set consists of just 20% of the training data with ~200 positive samples. The best threshold is likely to be overfitted on this small validation set with limited number of positive samples. Additionally, the F1 score by threshold line is almost flat with no clear peak, further indicating the best threshold is likely overfitting.\n\nMoreover, the validation predictions are overconfident, since all predictions are near 0/1. The threshold does not influence the binary predictions much, since the threshold will change the predicted value of just a few samples when changed (there are almost no samples with a predicted value between 0.20-0.80).\n\nThe threshold would only influence the F1 score when many predictions are not near 0/1.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F0711dee7b2441c64c999f802ea92932e%2Ftemp.png?generation=1675932274603709&alt=media)\n\nYou have a few options:\n\n1) stick to a threshold of 0.50\n2) Performs a full 5-Fold run and determine the best out-of-fold best threshold on the full dataset\n3) Use a different training strategy (i.e. loss function) to have less confident predictions",
      "votes": 5,
      "replies": [
        {
          "id": 2141876,
          "postDate": "2023-02-13T06:25:40.833Z",
          "content": "<p>Thank you for answer! Your answer is very helpful.</p>",
          "rawMarkdown": "Thank you for answer! Your answer is very helpful."
        },
        {
          "id": 2142384,
          "postDate": "2023-02-13T13:42:55.890Z",
          "content": "<p>I am new to this competition. Can you please help me? No matter what I do, my public score is always 0.04. I trained resnet model from scratch, with pre-trained weight, with / without validation set. No matter what I do the score is always 0.04. I trained those models in my PC and uploaded the dataset to kaggle and loaded those weight to models. What am I doing wrong, I cannot figure out.</p>",
          "rawMarkdown": "I am new to this competition. Can you please help me? No matter what I do, my public score is always 0.04. I trained resnet model from scratch, with pre-trained weight, with / without validation set. No matter what I do the score is always 0.04. I trained those models in my PC and uploaded the dataset to kaggle and loaded those weight to models. What am I doing wrong, I cannot figure out.",
          "replies": [
            {
              "id": 2142486,
              "postDate": "2023-02-13T15:05:16.737Z",
              "content": "<p>One tip I would have is to debug your inference notebook by generating a submission for your training data and comparing the results with your training notebook to validate your inference pipeline.</p>",
              "rawMarkdown": "One tip I would have is to debug your inference notebook by generating a submission for your training data and comparing the results with your training notebook to validate your inference pipeline."
            }
          ]
        }
      ]
    },
    {
      "id": 2136327,
      "postDate": "2023-02-09T08:47:55.910Z",
      "content": "<p>Did you transform the logits into probabilities using sigmoid/softmax?</p>",
      "rawMarkdown": "Did you transform the logits into probabilities using sigmoid/softmax?",
      "replies": [
        {
          "id": 2141877,
          "postDate": "2023-02-13T06:25:58.003Z",
          "content": "<p>Yes, I use sigmoid at the last layer of the output.</p>",
          "rawMarkdown": "Yes, I use sigmoid at the last layer of the output."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2136332,
      "author_name": "Mark Wijkhuizen",
      "author_url": "",
      "post_date": "2023-02-09T08:52:39.607000",
      "content": "<p>Keep in mind the validation set consists of just 20% of the training data with ~200 positive samples. The best threshold is likely to be overfitted on this small validation set with limited number of positive samples. Additionally, the F1 score by threshold line is almost flat with no clear peak, further indicating the best threshold is likely overfitting.</p>\n<p>Moreover, the validation predictions are overconfident, since all predictions are near 0/1. The threshold does not influence the binary predictions much, since the threshold will change the predicted value of just a few samples when changed (there are almost no samples with a predicted value between 0.20-0.80).</p>\n<p>The threshold would only influence the F1 score when many predictions are not near 0/1.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F0711dee7b2441c64c999f802ea92932e%2Ftemp.png?generation=1675932274603709&amp;alt=media\" alt=\"\"></p>\n<p>You have a few options:</p>\n<p>1) stick to a threshold of 0.50<br>\n2) Performs a full 5-Fold run and determine the best out-of-fold best threshold on the full dataset<br>\n3) Use a different training strategy (i.e. loss function) to have less confident predictions</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2141876,
          "author_name": "Ju7on9",
          "author_url": "",
          "post_date": "2023-02-13T06:25:40.833000",
          "content": "<p>Thank you for answer! Your answer is very helpful.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2142384,
          "author_name": "Rukesh Prajapati",
          "author_url": "",
          "post_date": "2023-02-13T13:42:55.890000",
          "content": "<p>I am new to this competition. Can you please help me? No matter what I do, my public score is always 0.04. I trained resnet model from scratch, with pre-trained weight, with / without validation set. No matter what I do the score is always 0.04. I trained those models in my PC and uploaded the dataset to kaggle and loaded those weight to models. What am I doing wrong, I cannot figure out.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2142486,
              "author_name": "Mark Wijkhuizen",
              "author_url": "",
              "post_date": "2023-02-13T15:05:16.737000",
              "content": "<p>One tip I would have is to debug your inference notebook by generating a submission for your training data and comparing the results with your training notebook to validate your inference pipeline.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2136327,
      "author_name": "Feng Qilong",
      "author_url": "",
      "post_date": "2023-02-09T08:47:55.910000",
      "content": "<p>Did you transform the logits into probabilities using sigmoid/softmax?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2141877,
          "author_name": "Ju7on9",
          "author_url": "",
          "post_date": "2023-02-13T06:25:58.003000",
          "content": "<p>Yes, I use sigmoid at the last layer of the output.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2136198": "Hello, \n\nI trained the model, but the best threshold is too low.\n\nI saw other people's result, but their threshold is almost 0.3~0.5.\n\nSo, I think something is wrong in my code.\n\nDo you have any idea what and when causes the low threshold?\n\nAt best threshold, my model  yields the following metric\n\n>val_pF1: 0.3170 - val_f1_score: 0.3370 - val_precision: 0.5897 - val_recall: 0.2359 - val_auc: 0.6855 - val_binary_accuracy: 0.9774\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5286406%2F5e0ced3d5f1da65a979d86035515c60e%2F2.PNG?generation=1675925073648187&alt=media)",
    "2136332": "Keep in mind the validation set consists of just 20% of the training data with ~200 positive samples. The best threshold is likely to be overfitted on this small validation set with limited number of positive samples. Additionally, the F1 score by threshold line is almost flat with no clear peak, further indicating the best threshold is likely overfitting.\n\nMoreover, the validation predictions are overconfident, since all predictions are near 0/1. The threshold does not influence the binary predictions much, since the threshold will change the predicted value of just a few samples when changed (there are almost no samples with a predicted value between 0.20-0.80).\n\nThe threshold would only influence the F1 score when many predictions are not near 0/1.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F0711dee7b2441c64c999f802ea92932e%2Ftemp.png?generation=1675932274603709&alt=media)\n\nYou have a few options:\n\n1) stick to a threshold of 0.50\n2) Performs a full 5-Fold run and determine the best out-of-fold best threshold on the full dataset\n3) Use a different training strategy (i.e. loss function) to have less confident predictions",
    "2136327": "Did you transform the logits into probabilities using sigmoid/softmax?"
  }
}