{
  "id": 377008,
  "title": "Is it against the rule to use a percentile of test set probabilities as threshold?",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/377008",
  "author_name": "Chenglu",
  "post_date": "2023-01-09T13:26:01.952000",
  "votes": 10,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I've already done some submissions with this, and it did bring a 0.01 boost to my LB score( 0.58 -&gt; 0.59 ), but it comes to me that this may violate the rules cause this was doing some \"training\" with test samples at some level.</p>\n<p>I'm not a Kaggle law expert so I just ask here for some helps, can I use this strategy? If not, I'll certainly not select the related submissions in the end.</p>\n<p>And to be more detailed i'm doing the following in the inference code(pseudo code):</p>\n<pre><code>probs = []\n sample  dataset:\n    prob = predict(sample)\n    probs.append(prob)\n\nfinal_predictions = probs &gt; percentile(probs, some_percentile_number)\n</code></pre>",
  "messages": [
    {
      "id": 2092651,
      "postDate": "2023-01-09T13:26:01.953Z",
      "content": "<p>I've already done some submissions with this, and it did bring a 0.01 boost to my LB score( 0.58 -&gt; 0.59 ), but it comes to me that this may violate the rules cause this was doing some \"training\" with test samples at some level.</p>\n<p>I'm not a Kaggle law expert so I just ask here for some helps, can I use this strategy? If not, I'll certainly not select the related submissions in the end.</p>\n<p>And to be more detailed i'm doing the following in the inference code(pseudo code):</p>\n<pre><code>probs = []\n sample  dataset:\n    prob = predict(sample)\n    probs.append(prob)\n\nfinal_predictions = probs &gt; percentile(probs, some_percentile_number)\n</code></pre>",
      "rawMarkdown": "I've already done some submissions with this, and it did bring a 0.01 boost to my LB score( 0.58 -> 0.59 ), but it comes to me that this may violate the rules cause this was doing some \"training\" with test samples at some level.\n\nI'm not a Kaggle law expert so I just ask here for some helps, can I use this strategy? If not, I'll certainly not select the related submissions in the end.\n\nAnd to be more detailed i'm doing the following in the inference code(pseudo code):\n\n```python\nprobs = []\nfor sample in dataset:\n    prob = predict(sample)\n    probs.append(prob)\n\nfinal_predictions = probs > percentile(probs, some_percentile_number)\n```\n",
      "votes": 10
    },
    {
      "id": 2092718,
      "postDate": "2023-01-09T14:29:14.437Z",
      "content": "<p>Making assumptions about the distribution of the test set and thresholding accordingly is a common practice. Sometimes it's helpful and sometimes it can lead to overfitting, but either way it's allowed.  </p>",
      "rawMarkdown": "Making assumptions about the distribution of the test set and thresholding accordingly is a common practice. Sometimes it's helpful and sometimes it can lead to overfitting, but either way it's allowed.  ",
      "votes": 3
    },
    {
      "id": 2094477,
      "postDate": "2023-01-10T19:19:53.090Z",
      "content": "<p>\"percentile trick\" can be used for speedup. </p>\n<hr>\n<p>in theory you can do this:</p>\n<ol>\n<li>use margin or weighted loss to modify your classifier to have very high recall rate while minimize false positive rate (fpr), i.e. modify the \"weakness\" or \"strongness\" of the classifier</li>\n<li>then you use cascade of classifiers.</li>\n</ol>\n<p>e.g. if each model has recall=0.99, fpr = 0.10, then cascade of 3 classifiers will have  recall = (0.99)^3, fpr = (0.10)^3</p>\n<p><img src=\"https://i.ibb.co/chj6gR2/Selection-492.png\" alt=\"https://i.ibb.co/chj6gR2/Selection-492.png\"></p>\n<p>The advantage is speedup. 90% of images uses one classifier, 9% uses 2 classifiers,  (1% uses 3 classifiers. Further since the function of the first classifier is only for rejection it is less complex and usually simple. </p>\n<p>(this is traditional trick in computer vision before deep learning era. google for \"cascade classifiers ensemble using adaboost\")</p>\n<p><img src=\"https://i.ibb.co/ZWK31TJ/Selection-491.png\" alt=\"https://i.ibb.co/ZWK31TJ/Selection-491.png\">.</p>\n<p>now you have a rough idea about the percentage of 2% of pos in hidden test (you can even probe it),<br>\nhence you can retain the correct percentile of test images at each stage of cascaded classifier.<br>\n(you can even probe the output at each stage to ensure at least say 99% of +ve test sampes pass through)</p>\n<hr>\n<p>on a side note, in modern deep learning era, some papers suggest:</p>\n<ol>\n<li>there is only one CNN network with many conv layers. You reject background images as you proceed from inital to final layers</li>\n<li>in transformer, there is token rejection. the length of token get less when the input passes through successive layers of trasnformer</li>\n<li>classic example is two stage rpn+ detection head for object detection. region proposal network is a background rejection network.</li>\n</ol>\n<p>these are meta algorithm trick for speedup </p>",
      "rawMarkdown": "\"percentile trick\" can be used for speedup. \n\n---\n\nin theory you can do this:\n\n1. use margin or weighted loss to modify your classifier to have very high recall rate while minimize false positive rate (fpr), i.e. modify the \"weakness\" or \"strongness\" of the classifier\n2. then you use cascade of classifiers.\n\ne.g. if each model has recall=0.99, fpr = 0.10, then cascade of 3 classifiers will have  recall = (0.99)^3, fpr = (0.10)^3\n\n![https://i.ibb.co/chj6gR2/Selection-492.png](https://i.ibb.co/chj6gR2/Selection-492.png)\n\n\nThe advantage is speedup. 90% of images uses one classifier, 9% uses 2 classifiers,  (1% uses 3 classifiers. Further since the function of the first classifier is only for rejection it is less complex and usually simple. \n\n(this is traditional trick in computer vision before deep learning era. google for \"cascade classifiers ensemble using adaboost\")\n\n![https://i.ibb.co/ZWK31TJ/Selection-491.png](https://i.ibb.co/ZWK31TJ/Selection-491.png).\n\nnow you have a rough idea about the percentage of 2% of pos in hidden test (you can even probe it),\nhence you can retain the correct percentile of test images at each stage of cascaded classifier.\n(you can even probe the output at each stage to ensure at least say 99% of +ve test sampes pass through)\n\n\n----\n\non a side note, in modern deep learning era, some papers suggest:\n1. there is only one CNN network with many conv layers. You reject background images as you proceed from inital to final layers\n2. in transformer, there is token rejection. the length of token get less when the input passes through successive layers of trasnformer\n3. classic example is two stage rpn+ detection head for object detection. region proposal network is a background rejection network.\n\nthese are meta algorithm trick for speedup \n\n",
      "votes": 4,
      "replies": [
        {
          "id": 2094959,
          "postDate": "2023-01-11T06:01:06.903Z",
          "content": "<p>as an example:</p>\n<p><img src=\"https://i.ibb.co/XWLjPJQ/Selection-501.png\" alt=\"https://i.ibb.co/XWLjPJQ/Selection-501.png\"></p>",
          "rawMarkdown": "as an example:\n\n![https://i.ibb.co/XWLjPJQ/Selection-501.png](https://i.ibb.co/XWLjPJQ/Selection-501.png)",
          "votes": 1,
          "replies": [
            {
              "id": 2095020,
              "postDate": "2023-01-11T06:45:58.070Z",
              "content": "<p><img src=\"https://i.ibb.co/7pS1xHM/Selection-506.png\" alt=\"https://i.ibb.co/7pS1xHM/Selection-506.png\"><br>\nand for those who want to write new paper. actually attention is a selection module (like little switches). hence transformers are very suitable for end-to-end rejection base models.</p>\n<p>the interesting thing is that the gradient needs not to be backprop from the end (because the signal doesn't reach the end) hence parallel block training is possible.</p>\n<p>update: it seems that nvidia already have a paper that has smiliar idea<br>\n<a href=\"https://github.com/NVlabs/A-ViT\" target=\"_blank\">https://github.com/NVlabs/A-ViT</a><br>\n\"In the new paper AdaViT: Adaptive Tokens for Efficient Vision Transformer, an Nvidia research team proposes AdaViT, an input-dependent mechanism that adaptively adjusts ViT inference cost by halting the compute of different tokens at different depths to reserve compute only for discriminative tokens.\"</p>",
              "rawMarkdown": "![https://i.ibb.co/7pS1xHM/Selection-506.png](https://i.ibb.co/7pS1xHM/Selection-506.png)\nand for those who want to write new paper. actually attention is a selection module (like little switches). hence transformers are very suitable for end-to-end rejection base models.\n\nthe interesting thing is that the gradient needs not to be backprop from the end (because the signal doesn't reach the end) hence parallel block training is possible.\n\n\nupdate: it seems that nvidia already have a paper that has smiliar idea\nhttps://github.com/NVlabs/A-ViT\n\"In the new paper AdaViT: Adaptive Tokens for Efficient Vision Transformer, an Nvidia research team proposes AdaViT, an input-dependent mechanism that adaptively adjusts ViT inference cost by halting the compute of different tokens at different depths to reserve compute only for discriminative tokens.\"",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2092727,
      "postDate": "2023-01-09T14:38:11.027Z",
      "content": "<p>I think it's allowed</p>",
      "rawMarkdown": "I think it's allowed",
      "votes": 1
    },
    {
      "id": 2093328,
      "postDate": "2023-01-10T00:20:17.303Z",
      "content": "<p>Thanks for everyone's replies. So to sum it up this is not against the rule but have a chance to very overfit the LB. Good to know that!</p>",
      "rawMarkdown": "Thanks for everyone's replies. So to sum it up this is not against the rule but have a chance to very overfit the LB. Good to know that!",
      "votes": 2
    },
    {
      "id": 2093010,
      "postDate": "2023-01-09T18:51:26.617Z",
      "content": "<p>doing some \"training\" with test samples at some level.<br>\ni thought it is allowed to do training with hidden test data?<br>\ne.g. online finetune,  few shot, etc</p>",
      "rawMarkdown": " doing some \"training\" with test samples at some level.\n\ni thought it is allowed to do training with hidden test data?\ne.g. online finetune,  few shot, etc",
      "votes": 2,
      "replies": [
        {
          "id": 2093019,
          "postDate": "2023-01-09T18:57:29.860Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2093021,
          "postDate": "2023-01-09T18:58:12.670Z",
          "content": "<p>I think that's incorrect analogy as we don't have hidden test labels.</p>",
          "rawMarkdown": "I think that's incorrect analogy as we don't have hidden test labels.",
          "replies": [
            {
              "id": 2093321,
              "postDate": "2023-01-10T00:15:51.087Z",
              "content": "<p>that should be right, never try this though but I think even with submission running we still can not get the labels.</p>",
              "rawMarkdown": "that should be right, never try this though but I think even with submission running we still can not get the labels."
            }
          ]
        }
      ]
    },
    {
      "id": 2092894,
      "postDate": "2023-01-09T17:18:59.520Z",
      "content": "<p>How do you pick the percentile ? :) expected cancer ratio? Or do you compute best percentile on CV and replicate on LB ?</p>",
      "rawMarkdown": "How do you pick the percentile ? :) expected cancer ratio? Or do you compute best percentile on CV and replicate on LB ?",
      "votes": 2,
      "replies": [
        {
          "id": 2093320,
          "postDate": "2023-01-10T00:14:19.597Z",
          "content": "<p>I get all the possibilities and instead of a hard threshold or a threshold picking by CV, I chose to get a percentile of the probabilities as the threshold.</p>",
          "rawMarkdown": "I get all the possibilities and instead of a hard threshold or a threshold picking by CV, I chose to get a percentile of the probabilities as the threshold."
        }
      ]
    },
    {
      "id": 2092834,
      "postDate": "2023-01-09T16:38:13.287Z",
      "content": "<p><a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a> You may find it has the opposite effect at the end of the competition, as the private leaderboard data could be different than  the public leaderboard data.</p>\n<p>This is a very common problem on kaggle and sometimes leads to very significant shakeup of results at the end because people do this sort of thresholding without good reason.</p>\n<p>Check out this notebook here for visualization of competition shakeups that have occured due to things like this  :  <a href=\"https://www.kaggle.com/code/jtrotman/meta-kaggle-scatter-plot-competition-shake-up\" target=\"_blank\">https://www.kaggle.com/code/jtrotman/meta-kaggle-scatter-plot-competition-shake-up</a></p>",
      "rawMarkdown": "@snaker You may find it has the opposite effect at the end of the competition, as the private leaderboard data could be different than  the public leaderboard data.\n\nThis is a very common problem on kaggle and sometimes leads to very significant shakeup of results at the end because people do this sort of thresholding without good reason.\n\nCheck out this notebook here for visualization of competition shakeups that have occured due to things like this  :  https://www.kaggle.com/code/jtrotman/meta-kaggle-scatter-plot-competition-shake-up",
      "votes": 2
    },
    {
      "id": 2092774,
      "postDate": "2023-01-09T15:33:09.623Z",
      "content": "<p>Of course this is allowed.</p>",
      "rawMarkdown": "Of course this is allowed.",
      "votes": 2
    },
    {
      "id": 2095113,
      "postDate": "2023-01-11T07:54:41.023Z",
      "content": "<p>Since both of the post processing(hard threshold/ percentage as a soft threshold) will cause overfitting, which is more severe in your opinion? <a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a> </p>",
      "rawMarkdown": "Since both of the post processing(hard threshold/ percentage as a soft threshold) will cause overfitting, which is more severe in your opinion? @snaker ",
      "replies": [
        {
          "id": 2095209,
          "postDate": "2023-01-11T08:37:31.883Z",
          "content": "<p>Honestly, I have no ideas, I think the only way not to overfit LB is to never look at the LB 😂 .</p>\n<p>If the private test set has the same distribution(the ratio of positive samples) as the training set, then the percentile trick  should beat hard threshold, but if not, the percentile will be shaken down a lot I think.</p>\n<p>A practical recommendation from me is to use the 2 final submission selections, one for percentile and one for hard threshold.</p>",
          "rawMarkdown": "Honestly, I have no ideas, I think the only way not to overfit LB is to never look at the LB 😂 .\n\nIf the private test set has the same distribution(the ratio of positive samples) as the training set, then the percentile trick  should beat hard threshold, but if not, the percentile will be shaken down a lot I think.\n\nA practical recommendation from me is to use the 2 final submission selections, one for percentile and one for hard threshold.",
          "votes": 1,
          "replies": [
            {
              "id": 2095501,
              "postDate": "2023-01-11T12:04:20.207Z",
              "content": "<p>You are right 👍</p>",
              "rawMarkdown": "You are right 👍"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2092718,
      "author_name": "David Austin",
      "author_url": "",
      "post_date": "2023-01-09T14:29:14.437000",
      "content": "<p>Making assumptions about the distribution of the test set and thresholding accordingly is a common practice. Sometimes it's helpful and sometimes it can lead to overfitting, but either way it's allowed.  </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2094477,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-01-10T19:19:53.090000",
      "content": "<p>\"percentile trick\" can be used for speedup. </p>\n<hr>\n<p>in theory you can do this:</p>\n<ol>\n<li>use margin or weighted loss to modify your classifier to have very high recall rate while minimize false positive rate (fpr), i.e. modify the \"weakness\" or \"strongness\" of the classifier</li>\n<li>then you use cascade of classifiers.</li>\n</ol>\n<p>e.g. if each model has recall=0.99, fpr = 0.10, then cascade of 3 classifiers will have  recall = (0.99)^3, fpr = (0.10)^3</p>\n<p><img src=\"https://i.ibb.co/chj6gR2/Selection-492.png\" alt=\"https://i.ibb.co/chj6gR2/Selection-492.png\"></p>\n<p>The advantage is speedup. 90% of images uses one classifier, 9% uses 2 classifiers,  (1% uses 3 classifiers. Further since the function of the first classifier is only for rejection it is less complex and usually simple. </p>\n<p>(this is traditional trick in computer vision before deep learning era. google for \"cascade classifiers ensemble using adaboost\")</p>\n<p><img src=\"https://i.ibb.co/ZWK31TJ/Selection-491.png\" alt=\"https://i.ibb.co/ZWK31TJ/Selection-491.png\">.</p>\n<p>now you have a rough idea about the percentage of 2% of pos in hidden test (you can even probe it),<br>\nhence you can retain the correct percentile of test images at each stage of cascaded classifier.<br>\n(you can even probe the output at each stage to ensure at least say 99% of +ve test sampes pass through)</p>\n<hr>\n<p>on a side note, in modern deep learning era, some papers suggest:</p>\n<ol>\n<li>there is only one CNN network with many conv layers. You reject background images as you proceed from inital to final layers</li>\n<li>in transformer, there is token rejection. the length of token get less when the input passes through successive layers of trasnformer</li>\n<li>classic example is two stage rpn+ detection head for object detection. region proposal network is a background rejection network.</li>\n</ol>\n<p>these are meta algorithm trick for speedup </p>",
      "votes": 4,
      "replies": [
        {
          "id": 2094959,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-01-11T06:01:06.903000",
          "content": "<p>as an example:</p>\n<p><img src=\"https://i.ibb.co/XWLjPJQ/Selection-501.png\" alt=\"https://i.ibb.co/XWLjPJQ/Selection-501.png\"></p>",
          "votes": 1,
          "replies": [
            {
              "id": 2095020,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-01-11T06:45:58.070000",
              "content": "<p><img src=\"https://i.ibb.co/7pS1xHM/Selection-506.png\" alt=\"https://i.ibb.co/7pS1xHM/Selection-506.png\"><br>\nand for those who want to write new paper. actually attention is a selection module (like little switches). hence transformers are very suitable for end-to-end rejection base models.</p>\n<p>the interesting thing is that the gradient needs not to be backprop from the end (because the signal doesn't reach the end) hence parallel block training is possible.</p>\n<p>update: it seems that nvidia already have a paper that has smiliar idea<br>\n<a href=\"https://github.com/NVlabs/A-ViT\" target=\"_blank\">https://github.com/NVlabs/A-ViT</a><br>\n\"In the new paper AdaViT: Adaptive Tokens for Efficient Vision Transformer, an Nvidia research team proposes AdaViT, an input-dependent mechanism that adaptively adjusts ViT inference cost by halting the compute of different tokens at different depths to reserve compute only for discriminative tokens.\"</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2092727,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2023-01-09T14:38:11.027000",
      "content": "<p>I think it's allowed</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2093328,
      "author_name": "Chenglu",
      "author_url": "",
      "post_date": "2023-01-10T00:20:17.303000",
      "content": "<p>Thanks for everyone's replies. So to sum it up this is not against the rule but have a chance to very overfit the LB. Good to know that!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2093010,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-01-09T18:51:26.617000",
      "content": "<p>doing some \"training\" with test samples at some level.<br>\ni thought it is allowed to do training with hidden test data?<br>\ne.g. online finetune,  few shot, etc</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2093019,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-01-09T18:57:29.860000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2093021,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2023-01-09T18:58:12.670000",
          "content": "<p>I think that's incorrect analogy as we don't have hidden test labels.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2093321,
              "author_name": "Chenglu",
              "author_url": "",
              "post_date": "2023-01-10T00:15:51.087000",
              "content": "<p>that should be right, never try this though but I think even with submission running we still can not get the labels.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2092894,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2023-01-09T17:18:59.520000",
      "content": "<p>How do you pick the percentile ? :) expected cancer ratio? Or do you compute best percentile on CV and replicate on LB ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2093320,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2023-01-10T00:14:19.597000",
          "content": "<p>I get all the possibilities and instead of a hard threshold or a threshold picking by CV, I chose to get a percentile of the probabilities as the threshold.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2092834,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2023-01-09T16:38:13.287000",
      "content": "<p><a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a> You may find it has the opposite effect at the end of the competition, as the private leaderboard data could be different than  the public leaderboard data.</p>\n<p>This is a very common problem on kaggle and sometimes leads to very significant shakeup of results at the end because people do this sort of thresholding without good reason.</p>\n<p>Check out this notebook here for visualization of competition shakeups that have occured due to things like this  :  <a href=\"https://www.kaggle.com/code/jtrotman/meta-kaggle-scatter-plot-competition-shake-up\" target=\"_blank\">https://www.kaggle.com/code/jtrotman/meta-kaggle-scatter-plot-competition-shake-up</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2092774,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2023-01-09T15:33:09.623000",
      "content": "<p>Of course this is allowed.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2095113,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2023-01-11T07:54:41.023000",
      "content": "<p>Since both of the post processing(hard threshold/ percentage as a soft threshold) will cause overfitting, which is more severe in your opinion? <a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2095209,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2023-01-11T08:37:31.883000",
          "content": "<p>Honestly, I have no ideas, I think the only way not to overfit LB is to never look at the LB 😂 .</p>\n<p>If the private test set has the same distribution(the ratio of positive samples) as the training set, then the percentile trick  should beat hard threshold, but if not, the percentile will be shaken down a lot I think.</p>\n<p>A practical recommendation from me is to use the 2 final submission selections, one for percentile and one for hard threshold.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2095501,
              "author_name": "ForcewithMe",
              "author_url": "",
              "post_date": "2023-01-11T12:04:20.207000",
              "content": "<p>You are right 👍</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2092651": "I've already done some submissions with this, and it did bring a 0.01 boost to my LB score( 0.58 -> 0.59 ), but it comes to me that this may violate the rules cause this was doing some \"training\" with test samples at some level.\n\nI'm not a Kaggle law expert so I just ask here for some helps, can I use this strategy? If not, I'll certainly not select the related submissions in the end.\n\nAnd to be more detailed i'm doing the following in the inference code(pseudo code):\n\n```python\nprobs = []\nfor sample in dataset:\n    prob = predict(sample)\n    probs.append(prob)\n\nfinal_predictions = probs > percentile(probs, some_percentile_number)\n```\n",
    "2092718": "Making assumptions about the distribution of the test set and thresholding accordingly is a common practice. Sometimes it's helpful and sometimes it can lead to overfitting, but either way it's allowed.  ",
    "2094477": "\"percentile trick\" can be used for speedup. \n\n---\n\nin theory you can do this:\n\n1. use margin or weighted loss to modify your classifier to have very high recall rate while minimize false positive rate (fpr), i.e. modify the \"weakness\" or \"strongness\" of the classifier\n2. then you use cascade of classifiers.\n\ne.g. if each model has recall=0.99, fpr = 0.10, then cascade of 3 classifiers will have  recall = (0.99)^3, fpr = (0.10)^3\n\n![https://i.ibb.co/chj6gR2/Selection-492.png](https://i.ibb.co/chj6gR2/Selection-492.png)\n\n\nThe advantage is speedup. 90% of images uses one classifier, 9% uses 2 classifiers,  (1% uses 3 classifiers. Further since the function of the first classifier is only for rejection it is less complex and usually simple. \n\n(this is traditional trick in computer vision before deep learning era. google for \"cascade classifiers ensemble using adaboost\")\n\n![https://i.ibb.co/ZWK31TJ/Selection-491.png](https://i.ibb.co/ZWK31TJ/Selection-491.png).\n\nnow you have a rough idea about the percentage of 2% of pos in hidden test (you can even probe it),\nhence you can retain the correct percentile of test images at each stage of cascaded classifier.\n(you can even probe the output at each stage to ensure at least say 99% of +ve test sampes pass through)\n\n\n----\n\non a side note, in modern deep learning era, some papers suggest:\n1. there is only one CNN network with many conv layers. You reject background images as you proceed from inital to final layers\n2. in transformer, there is token rejection. the length of token get less when the input passes through successive layers of trasnformer\n3. classic example is two stage rpn+ detection head for object detection. region proposal network is a background rejection network.\n\nthese are meta algorithm trick for speedup \n\n",
    "2092727": "I think it's allowed",
    "2093328": "Thanks for everyone's replies. So to sum it up this is not against the rule but have a chance to very overfit the LB. Good to know that!",
    "2093010": " doing some \"training\" with test samples at some level.\n\ni thought it is allowed to do training with hidden test data?\ne.g. online finetune,  few shot, etc",
    "2092894": "How do you pick the percentile ? :) expected cancer ratio? Or do you compute best percentile on CV and replicate on LB ?",
    "2092834": "@snaker You may find it has the opposite effect at the end of the competition, as the private leaderboard data could be different than  the public leaderboard data.\n\nThis is a very common problem on kaggle and sometimes leads to very significant shakeup of results at the end because people do this sort of thresholding without good reason.\n\nCheck out this notebook here for visualization of competition shakeups that have occured due to things like this  :  https://www.kaggle.com/code/jtrotman/meta-kaggle-scatter-plot-competition-shake-up",
    "2092774": "Of course this is allowed.",
    "2095113": "Since both of the post processing(hard threshold/ percentage as a soft threshold) will cause overfitting, which is more severe in your opinion? @snaker "
  }
}