{
  "id": 374672,
  "title": "Very high probabilistic threshold",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/374672",
  "author_name": "Pandas Warrior",
  "post_date": "2022-12-28T09:45:25.419000",
  "votes": 6,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi, I have already done more than 100 experiments in this competition and I am worried about one thing, maybe you will be able to dispel my doubts.</p>\n<p>Namely, at the end of the training, I try to choose the optimal threshold to binarize the model response. I decide whether it is optimal using the simple heuristics below. </p>\n<pre><code>\n   ():\n        f1_thresholded = []\n        thresholds = np.linspace(, , )\n         threshold  thresholds:\n            predictions_thresholded = (predictions &gt; threshold).astype()\n            f1 = pfbeta(labels, predictions_thresholded)\n            f1_thresholded.append((f1))\n        max_f1 = np.(f1_thresholded)\n        best_threshold = thresholds[np.argmax(f1_thresholded)]\n         max_f1, best_threshold\n</code></pre>\n<p>Threshold I check both per image and per patient laterality, which I obtain in a way that I group the images by patient_id + laterality and then choose to perform aggregation using max(). To my surprise, the most common values of the optimal threshold are around 0.99, and the best result on LB (0.42) I just obtained with a threshold of 0.996.</p>\n<p>Has anyone experienced similar observations? It seems to me quite unnatural to use, in fact, only the predictions of the model of which one is almost 100% sure. In addition, I have the impression that with this the model is very sensitive to the fluctuation of such threshold, the difference on LB between using threshold 0.99 and 0.996 for the same model is as much as 0.03</p>",
  "messages": [
    {
      "id": 2078388,
      "postDate": "2022-12-28T09:45:25.420Z",
      "content": "<p>Hi, I have already done more than 100 experiments in this competition and I am worried about one thing, maybe you will be able to dispel my doubts.</p>\n<p>Namely, at the end of the training, I try to choose the optimal threshold to binarize the model response. I decide whether it is optimal using the simple heuristics below. </p>\n<pre><code>\n   ():\n        f1_thresholded = []\n        thresholds = np.linspace(, , )\n         threshold  thresholds:\n            predictions_thresholded = (predictions &gt; threshold).astype()\n            f1 = pfbeta(labels, predictions_thresholded)\n            f1_thresholded.append((f1))\n        max_f1 = np.(f1_thresholded)\n        best_threshold = thresholds[np.argmax(f1_thresholded)]\n         max_f1, best_threshold\n</code></pre>\n<p>Threshold I check both per image and per patient laterality, which I obtain in a way that I group the images by patient_id + laterality and then choose to perform aggregation using max(). To my surprise, the most common values of the optimal threshold are around 0.99, and the best result on LB (0.42) I just obtained with a threshold of 0.996.</p>\n<p>Has anyone experienced similar observations? It seems to me quite unnatural to use, in fact, only the predictions of the model of which one is almost 100% sure. In addition, I have the impression that with this the model is very sensitive to the fluctuation of such threshold, the difference on LB between using threshold 0.99 and 0.996 for the same model is as much as 0.03</p>",
      "rawMarkdown": "Hi, I have already done more than 100 experiments in this competition and I am worried about one thing, maybe you will be able to dispel my doubts.\n\nNamely, at the end of the training, I try to choose the optimal threshold to binarize the model response. I decide whether it is optimal using the simple heuristics below. \n\n```python\n@staticmethod\n  def optimize_thresholds(labels, predictions):\n        f1_thresholded = []\n        thresholds = np.linspace(0.001, 0.999, 999)\n        for threshold in thresholds:\n            predictions_thresholded = (predictions > threshold).astype(int)\n            f1 = pfbeta(labels, predictions_thresholded)\n            f1_thresholded.append(float(f1))\n        max_f1 = np.max(f1_thresholded)\n        best_threshold = thresholds[np.argmax(f1_thresholded)]\n        return max_f1, best_threshold\n```\n\nThreshold I check both per image and per patient laterality, which I obtain in a way that I group the images by patient_id + laterality and then choose to perform aggregation using max(). To my surprise, the most common values of the optimal threshold are around 0.99, and the best result on LB (0.42) I just obtained with a threshold of 0.996.\n\nHas anyone experienced similar observations? It seems to me quite unnatural to use, in fact, only the predictions of the model of which one is almost 100% sure. In addition, I have the impression that with this the model is very sensitive to the fluctuation of such threshold, the difference on LB between using threshold 0.99 and 0.996 for the same model is as much as 0.03",
      "votes": 6
    },
    {
      "id": 2079092,
      "postDate": "2022-12-29T01:32:01.960Z",
      "content": "<p>you have to imagine (and visualize) what happens in the data space to explain everything in data science/machine learning</p>\n<p><img src=\"https://i.ibb.co/P5K7swD/Selection-335.png\" alt=\"https://i.ibb.co/P5K7swD/Selection-335.png\"></p>",
      "rawMarkdown": "you have to imagine (and visualize) what happens in the data space to explain everything in data science/machine learning\n\n![https://i.ibb.co/P5K7swD/Selection-335.png](https://i.ibb.co/P5K7swD/Selection-335.png)",
      "votes": 3,
      "replies": [
        {
          "id": 2082174,
          "postDate": "2023-01-01T08:10:47.477Z",
          "content": "<p>probably heavy augmentation can improve the generalization? Also I had a thought of discarding difficult negative cases from the training set (not from validation) and use MixUp as one of augmentation along with it? What are your thoughts? I feel It would increase the generalization.</p>",
          "rawMarkdown": "probably heavy augmentation can improve the generalization? Also I had a thought of discarding difficult negative cases from the training set (not from validation) and use MixUp as one of augmentation along with it? What are your thoughts? I feel It would increase the generalization."
        }
      ]
    },
    {
      "id": 2082389,
      "postDate": "2023-01-01T14:18:32.877Z",
      "content": "<p>One of the reasons behind the high threshold is using .max() to aggregate the laterality predictions. If you calculate the standard deviation of the patient/laterality, most cases have a high std of ~0.5 which means, in many cases the same breast will get a score of 0.99 for view MLO and 0. for view CC (and vice versa). <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2Fc2b0ab4d487ef771b76017e95f5bab1d%2Flat_std.png?generation=1672581201405204&amp;alt=media\" alt=\"\"></p>\n<p>When you aggregate using .max(), the highest view prediction is selected, it tends to be +0.9 due to the model's overconfidence, which makes the threshold=0.9<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2Fb87715c8f377c5d19ba7cc833473414d%2Flat_max.png?generation=1672581970488688&amp;alt=media\" alt=\"\"></p>\n<p>When you aggregate using.mean(), the view predictions will even out and the threshold will be in a lower range, in this case, it's 0.35<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F57447ec41c69a8c13aad3f9d29008538%2Flat_mean.png?generation=1672582131619085&amp;alt=media\" alt=\"\"></p>\n<p>I am not sure though which threshold selection method is more sensitive to CV/LB fluctuation, but for now using both mean and max give more or less the same CV score (CV=0.35, LB=0.40). Did you try to submit using mean aggregations and noticed the same fluctuation?</p>",
      "rawMarkdown": "One of the reasons behind the high threshold is using .max() to aggregate the laterality predictions. If you calculate the standard deviation of the patient/laterality, most cases have a high std of ~0.5 which means, in many cases the same breast will get a score of 0.99 for view MLO and 0. for view CC (and vice versa). \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2Fc2b0ab4d487ef771b76017e95f5bab1d%2Flat_std.png?generation=1672581201405204&alt=media)\n\nWhen you aggregate using .max(), the highest view prediction is selected, it tends to be +0.9 due to the model's overconfidence, which makes the threshold=0.9\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2Fb87715c8f377c5d19ba7cc833473414d%2Flat_max.png?generation=1672581970488688&alt=media)\n\nWhen you aggregate using.mean(), the view predictions will even out and the threshold will be in a lower range, in this case, it's 0.35\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F57447ec41c69a8c13aad3f9d29008538%2Flat_mean.png?generation=1672582131619085&alt=media)\n\nI am not sure though which threshold selection method is more sensitive to CV/LB fluctuation, but for now using both mean and max give more or less the same CV score (CV=0.35, LB=0.40). Did you try to submit using mean aggregations and noticed the same fluctuation?",
      "votes": 4,
      "replies": [
        {
          "id": 2082416,
          "postDate": "2023-01-01T15:03:09.370Z",
          "content": "<p>Based on the CV results, the scores around the best threshold are stable but I think it all depends on the distribution of the private test set.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F2a43f1bc42230edb2784209e4240170b%2Fthresholds.png?generation=1672585278728567&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Based on the CV results, the scores around the best threshold are stable but I think it all depends on the distribution of the private test set.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F2a43f1bc42230edb2784209e4240170b%2Fthresholds.png?generation=1672585278728567&alt=media)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2079614,
      "postDate": "2022-12-29T13:36:30.520Z",
      "content": "<p>I think its based on type of loss you use, BCE outputs very confident scores hence my threshold is 0.95(for me) where for focal loss threshold is around 0.4-0.5.Trying focal loss with same pipeline</p>",
      "rawMarkdown": "I think its based on type of loss you use, BCE outputs very confident scores hence my threshold is 0.95(for me) where for focal loss threshold is around 0.4-0.5.Trying focal loss with same pipeline\n",
      "votes": 1,
      "replies": [
        {
          "id": 2079976,
          "postDate": "2022-12-29T18:54:26.777Z",
          "content": "<p>That's true, I also noticed that using focal loss \"optimal\" thresholds are much lower</p>",
          "rawMarkdown": "That's true, I also noticed that using focal loss \"optimal\" thresholds are much lower"
        }
      ]
    },
    {
      "id": 2079082,
      "postDate": "2022-12-29T00:54:43.683Z",
      "content": "<p>The \"probabilities\" are a byproduct of how you train the model (for instance, whether you use label smoothing) and of how long you train your model.</p>\n<p>The longer you train your model, the more \"overconfident\" it becomes, usually.</p>\n<p>So I guess possibly a more telling way of looking at what your model is doing might be a confusion matrix at the \"optimal\" threshold. How do the positives and false positives fluctuate as you change your threshold?</p>\n<p>Unfortunately, the competition metric is quite sensitive to this \"optimization\" which can possibly lead to a large shakeup… I believe this is something that <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> discussed in one of his posts!</p>",
      "rawMarkdown": "The \"probabilities\" are a byproduct of how you train the model (for instance, whether you use label smoothing) and of how long you train your model.\n\nThe longer you train your model, the more \"overconfident\" it becomes, usually.\n\nSo I guess possibly a more telling way of looking at what your model is doing might be a confusion matrix at the \"optimal\" threshold. How do the positives and false positives fluctuate as you change your threshold?\n\nUnfortunately, the competition metric is quite sensitive to this \"optimization\" which can possibly lead to a large shakeup... I believe this is something that @hengck23 discussed in one of his posts!",
      "votes": 1
    },
    {
      "id": 2078948,
      "postDate": "2022-12-28T19:28:02.847Z",
      "content": "<p>Maybe too high threshold means the model's generalization ability is not that good?<br>\nMy high evalution score model threshold is around 0.6-0.7.<br>\nI think if you can tune the model or do something to lower this threshold, you can get higher LB score. </p>",
      "rawMarkdown": "Maybe too high threshold means the model's generalization ability is not that good?\nMy high evalution score model threshold is around 0.6-0.7.\nI think if you can tune the model or do something to lower this threshold, you can get higher LB score. ",
      "votes": 2,
      "replies": [
        {
          "id": 2078970,
          "postDate": "2022-12-28T20:12:17.407Z",
          "content": "<p>When I use the mean function to aggregate the results, instead the maximum and then optimizie the threshold, I get values closer to 0.7, while the result on LB is then lower.</p>\n<p>What is your way to aggregate the predictions per patient-laterality?</p>",
          "rawMarkdown": "When I use the mean function to aggregate the results, instead the maximum and then optimizie the threshold, I get values closer to 0.7, while the result on LB is then lower.\n\nWhat is your way to aggregate the predictions per patient-laterality?",
          "votes": 1
        }
      ]
    },
    {
      "id": 2098282,
      "postDate": "2023-01-13T11:50:58.250Z",
      "content": "<p>Did you use sigmoid/softmax? Else the logit can be larger than 1, which results in high threshold</p>",
      "rawMarkdown": "Did you use sigmoid/softmax? Else the logit can be larger than 1, which results in high threshold"
    },
    {
      "id": 2078992,
      "postDate": "2022-12-28T21:09:07.477Z",
      "content": "<p>I have the same issue. My thresholds are between 0.8 - 0.9. This makes it really harder to determine a threshold for multiple model or fold ensembling. Do you have any recommadition??</p>",
      "rawMarkdown": "I have the same issue. My thresholds are between 0.8 - 0.9. This makes it really harder to determine a threshold for multiple model or fold ensembling. Do you have any recommadition??",
      "replies": [
        {
          "id": 2157963,
          "postDate": "2023-02-24T13:59:50.567Z",
          "content": "<p>I found out that it was caused by overweighting the loss function for positive cancer cases. When I punished the model too much for bad predictions on a positive sample, the distribution of its prediction shifted sharply to the right while at the same time tapering off. This resulted in a high optimal threshold and high sensitivity to fluctuations.</p>",
          "rawMarkdown": "I found out that it was caused by overweighting the loss function for positive cancer cases. When I punished the model too much for bad predictions on a positive sample, the distribution of its prediction shifted sharply to the right while at the same time tapering off. This resulted in a high optimal threshold and high sensitivity to fluctuations."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2079092,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-29T01:32:01.960000",
      "content": "<p>you have to imagine (and visualize) what happens in the data space to explain everything in data science/machine learning</p>\n<p><img src=\"https://i.ibb.co/P5K7swD/Selection-335.png\" alt=\"https://i.ibb.co/P5K7swD/Selection-335.png\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 2082174,
          "author_name": "Tanmay Mane",
          "author_url": "",
          "post_date": "2023-01-01T08:10:47.477000",
          "content": "<p>probably heavy augmentation can improve the generalization? Also I had a thought of discarding difficult negative cases from the training set (not from validation) and use MixUp as one of augmentation along with it? What are your thoughts? I feel It would increase the generalization.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2082389,
      "author_name": "Amin",
      "author_url": "",
      "post_date": "2023-01-01T14:18:32.877000",
      "content": "<p>One of the reasons behind the high threshold is using .max() to aggregate the laterality predictions. If you calculate the standard deviation of the patient/laterality, most cases have a high std of ~0.5 which means, in many cases the same breast will get a score of 0.99 for view MLO and 0. for view CC (and vice versa). <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2Fc2b0ab4d487ef771b76017e95f5bab1d%2Flat_std.png?generation=1672581201405204&amp;alt=media\" alt=\"\"></p>\n<p>When you aggregate using .max(), the highest view prediction is selected, it tends to be +0.9 due to the model's overconfidence, which makes the threshold=0.9<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2Fb87715c8f377c5d19ba7cc833473414d%2Flat_max.png?generation=1672581970488688&amp;alt=media\" alt=\"\"></p>\n<p>When you aggregate using.mean(), the view predictions will even out and the threshold will be in a lower range, in this case, it's 0.35<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F57447ec41c69a8c13aad3f9d29008538%2Flat_mean.png?generation=1672582131619085&amp;alt=media\" alt=\"\"></p>\n<p>I am not sure though which threshold selection method is more sensitive to CV/LB fluctuation, but for now using both mean and max give more or less the same CV score (CV=0.35, LB=0.40). Did you try to submit using mean aggregations and noticed the same fluctuation?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2082416,
          "author_name": "Amin",
          "author_url": "",
          "post_date": "2023-01-01T15:03:09.370000",
          "content": "<p>Based on the CV results, the scores around the best threshold are stable but I think it all depends on the distribution of the private test set.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F2a43f1bc42230edb2784209e4240170b%2Fthresholds.png?generation=1672585278728567&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2079614,
      "author_name": "pranav",
      "author_url": "",
      "post_date": "2022-12-29T13:36:30.520000",
      "content": "<p>I think its based on type of loss you use, BCE outputs very confident scores hence my threshold is 0.95(for me) where for focal loss threshold is around 0.4-0.5.Trying focal loss with same pipeline</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2079976,
          "author_name": "Pandas Warrior",
          "author_url": "",
          "post_date": "2022-12-29T18:54:26.777000",
          "content": "<p>That's true, I also noticed that using focal loss \"optimal\" thresholds are much lower</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2079082,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2022-12-29T00:54:43.683000",
      "content": "<p>The \"probabilities\" are a byproduct of how you train the model (for instance, whether you use label smoothing) and of how long you train your model.</p>\n<p>The longer you train your model, the more \"overconfident\" it becomes, usually.</p>\n<p>So I guess possibly a more telling way of looking at what your model is doing might be a confusion matrix at the \"optimal\" threshold. How do the positives and false positives fluctuate as you change your threshold?</p>\n<p>Unfortunately, the competition metric is quite sensitive to this \"optimization\" which can possibly lead to a large shakeup… I believe this is something that <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> discussed in one of his posts!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2078948,
      "author_name": "nicehzj",
      "author_url": "",
      "post_date": "2022-12-28T19:28:02.847000",
      "content": "<p>Maybe too high threshold means the model's generalization ability is not that good?<br>\nMy high evalution score model threshold is around 0.6-0.7.<br>\nI think if you can tune the model or do something to lower this threshold, you can get higher LB score. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2078970,
          "author_name": "Pandas Warrior",
          "author_url": "",
          "post_date": "2022-12-28T20:12:17.407000",
          "content": "<p>When I use the mean function to aggregate the results, instead the maximum and then optimizie the threshold, I get values closer to 0.7, while the result on LB is then lower.</p>\n<p>What is your way to aggregate the predictions per patient-laterality?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2098282,
      "author_name": "Feng Qilong",
      "author_url": "",
      "post_date": "2023-01-13T11:50:58.250000",
      "content": "<p>Did you use sigmoid/softmax? Else the logit can be larger than 1, which results in high threshold</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2078992,
      "author_name": "Erdi Kılıç",
      "author_url": "",
      "post_date": "2022-12-28T21:09:07.477000",
      "content": "<p>I have the same issue. My thresholds are between 0.8 - 0.9. This makes it really harder to determine a threshold for multiple model or fold ensembling. Do you have any recommadition??</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2157963,
          "author_name": "Pandas Warrior",
          "author_url": "",
          "post_date": "2023-02-24T13:59:50.567000",
          "content": "<p>I found out that it was caused by overweighting the loss function for positive cancer cases. When I punished the model too much for bad predictions on a positive sample, the distribution of its prediction shifted sharply to the right while at the same time tapering off. This resulted in a high optimal threshold and high sensitivity to fluctuations.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2078388": "Hi, I have already done more than 100 experiments in this competition and I am worried about one thing, maybe you will be able to dispel my doubts.\n\nNamely, at the end of the training, I try to choose the optimal threshold to binarize the model response. I decide whether it is optimal using the simple heuristics below. \n\n```python\n@staticmethod\n  def optimize_thresholds(labels, predictions):\n        f1_thresholded = []\n        thresholds = np.linspace(0.001, 0.999, 999)\n        for threshold in thresholds:\n            predictions_thresholded = (predictions > threshold).astype(int)\n            f1 = pfbeta(labels, predictions_thresholded)\n            f1_thresholded.append(float(f1))\n        max_f1 = np.max(f1_thresholded)\n        best_threshold = thresholds[np.argmax(f1_thresholded)]\n        return max_f1, best_threshold\n```\n\nThreshold I check both per image and per patient laterality, which I obtain in a way that I group the images by patient_id + laterality and then choose to perform aggregation using max(). To my surprise, the most common values of the optimal threshold are around 0.99, and the best result on LB (0.42) I just obtained with a threshold of 0.996.\n\nHas anyone experienced similar observations? It seems to me quite unnatural to use, in fact, only the predictions of the model of which one is almost 100% sure. In addition, I have the impression that with this the model is very sensitive to the fluctuation of such threshold, the difference on LB between using threshold 0.99 and 0.996 for the same model is as much as 0.03",
    "2079092": "you have to imagine (and visualize) what happens in the data space to explain everything in data science/machine learning\n\n![https://i.ibb.co/P5K7swD/Selection-335.png](https://i.ibb.co/P5K7swD/Selection-335.png)",
    "2082389": "One of the reasons behind the high threshold is using .max() to aggregate the laterality predictions. If you calculate the standard deviation of the patient/laterality, most cases have a high std of ~0.5 which means, in many cases the same breast will get a score of 0.99 for view MLO and 0. for view CC (and vice versa). \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2Fc2b0ab4d487ef771b76017e95f5bab1d%2Flat_std.png?generation=1672581201405204&alt=media)\n\nWhen you aggregate using .max(), the highest view prediction is selected, it tends to be +0.9 due to the model's overconfidence, which makes the threshold=0.9\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2Fb87715c8f377c5d19ba7cc833473414d%2Flat_max.png?generation=1672581970488688&alt=media)\n\nWhen you aggregate using.mean(), the view predictions will even out and the threshold will be in a lower range, in this case, it's 0.35\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3451735%2F57447ec41c69a8c13aad3f9d29008538%2Flat_mean.png?generation=1672582131619085&alt=media)\n\nI am not sure though which threshold selection method is more sensitive to CV/LB fluctuation, but for now using both mean and max give more or less the same CV score (CV=0.35, LB=0.40). Did you try to submit using mean aggregations and noticed the same fluctuation?",
    "2079614": "I think its based on type of loss you use, BCE outputs very confident scores hence my threshold is 0.95(for me) where for focal loss threshold is around 0.4-0.5.Trying focal loss with same pipeline\n",
    "2079082": "The \"probabilities\" are a byproduct of how you train the model (for instance, whether you use label smoothing) and of how long you train your model.\n\nThe longer you train your model, the more \"overconfident\" it becomes, usually.\n\nSo I guess possibly a more telling way of looking at what your model is doing might be a confusion matrix at the \"optimal\" threshold. How do the positives and false positives fluctuate as you change your threshold?\n\nUnfortunately, the competition metric is quite sensitive to this \"optimization\" which can possibly lead to a large shakeup... I believe this is something that @hengck23 discussed in one of his posts!",
    "2078948": "Maybe too high threshold means the model's generalization ability is not that good?\nMy high evalution score model threshold is around 0.6-0.7.\nI think if you can tune the model or do something to lower this threshold, you can get higher LB score. ",
    "2098282": "Did you use sigmoid/softmax? Else the logit can be larger than 1, which results in high threshold",
    "2078992": "I have the same issue. My thresholds are between 0.8 - 0.9. This makes it really harder to determine a threshold for multiple model or fold ensembling. Do you have any recommadition??"
  }
}