{
  "id": 371874,
  "title": "ensemble not working?",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/371874",
  "author_name": "hengck23",
  "post_date": "2022-12-12T20:52:12.743000",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<p><img src=\"https://i.ibb.co/n1nMm0j/Selection-155.png\" alt=\"https://i.ibb.co/n1nMm0j/Selection-155.png\"></p>\n<p>you observe that LB-CV does correlate but variance is high. ensemble not necessarily improve LB (like what you would expected in other competitions)</p>\n<p>The problem occurs because we want to find a single threshold to binarized. The main issue is not imbalanced but lack of pos samples. When there is lack of dataset, everything goes haywire.</p>\n<p>You may need to realigned your predicted probability (calibration) if:</p>\n<ol>\n<li><p>you are using different train parameters (e.g. different weights in class weighted loss, different over sampling) for different fold</p></li>\n<li><p>even if you are using the same train parameters,  the validation sample size is too small and biased for each fold</p></li>\n</ol>\n<p>Even if you aligned and prove to work on public test set, it may not work on private set. </p>\n<hr>\n<p>You probably needs to think of how to</p>\n<ul>\n<li>use external data (then there is another problem of possibly domain shift)</li>\n<li>better fold stratification and compute more reliable statistics  </li>\n<li>probe your hidden test data for magic characteristics (you probably didn't solve the problem but employ competition tricks to stabilized your score)</li>\n</ul>",
  "messages": [
    {
      "id": 2063384,
      "postDate": "2022-12-12T20:52:12.743Z",
      "content": "<p><img src=\"https://i.ibb.co/n1nMm0j/Selection-155.png\" alt=\"https://i.ibb.co/n1nMm0j/Selection-155.png\"></p>\n<p>you observe that LB-CV does correlate but variance is high. ensemble not necessarily improve LB (like what you would expected in other competitions)</p>\n<p>The problem occurs because we want to find a single threshold to binarized. The main issue is not imbalanced but lack of pos samples. When there is lack of dataset, everything goes haywire.</p>\n<p>You may need to realigned your predicted probability (calibration) if:</p>\n<ol>\n<li><p>you are using different train parameters (e.g. different weights in class weighted loss, different over sampling) for different fold</p></li>\n<li><p>even if you are using the same train parameters,  the validation sample size is too small and biased for each fold</p></li>\n</ol>\n<p>Even if you aligned and prove to work on public test set, it may not work on private set. </p>\n<hr>\n<p>You probably needs to think of how to</p>\n<ul>\n<li>use external data (then there is another problem of possibly domain shift)</li>\n<li>better fold stratification and compute more reliable statistics  </li>\n<li>probe your hidden test data for magic characteristics (you probably didn't solve the problem but employ competition tricks to stabilized your score)</li>\n</ul>",
      "rawMarkdown": "![https://i.ibb.co/n1nMm0j/Selection-155.png](https://i.ibb.co/n1nMm0j/Selection-155.png)\n\n\nyou observe that LB-CV does correlate but variance is high. ensemble not necessarily improve LB (like what you would expected in other competitions)\n\nThe problem occurs because we want to find a single threshold to binarized. The main issue is not imbalanced but lack of pos samples. When there is lack of dataset, everything goes haywire.\n\nYou may need to realigned your predicted probability (calibration) if:\n1. you are using different train parameters (e.g. different weights in class weighted loss, different over sampling) for different fold\n\n2. even if you are using the same train parameters,  the validation sample size is too small and biased for each fold\n\nEven if you aligned and prove to work on public test set, it may not work on private set. \n\n---\n\nYou probably needs to think of how to\n- use external data (then there is another problem of possibly domain shift)\n- better fold stratification and compute more reliable statistics  \n- probe your hidden test data for magic characteristics (you probably didn't solve the problem but employ competition tricks to stabilized your score)\n",
      "votes": 11
    },
    {
      "id": 2066160,
      "postDate": "2022-12-15T12:53:55.060Z",
      "content": "<p>totally agree</p>",
      "rawMarkdown": "totally agree"
    },
    {
      "id": 2064725,
      "postDate": "2022-12-14T04:34:57.903Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2066160,
      "author_name": "DataManyo",
      "author_url": "",
      "post_date": "2022-12-15T12:53:55.060000",
      "content": "<p>totally agree</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2064725,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-14T04:34:57.903000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2063384": "![https://i.ibb.co/n1nMm0j/Selection-155.png](https://i.ibb.co/n1nMm0j/Selection-155.png)\n\n\nyou observe that LB-CV does correlate but variance is high. ensemble not necessarily improve LB (like what you would expected in other competitions)\n\nThe problem occurs because we want to find a single threshold to binarized. The main issue is not imbalanced but lack of pos samples. When there is lack of dataset, everything goes haywire.\n\nYou may need to realigned your predicted probability (calibration) if:\n1. you are using different train parameters (e.g. different weights in class weighted loss, different over sampling) for different fold\n\n2. even if you are using the same train parameters,  the validation sample size is too small and biased for each fold\n\nEven if you aligned and prove to work on public test set, it may not work on private set. \n\n---\n\nYou probably needs to think of how to\n- use external data (then there is another problem of possibly domain shift)\n- better fold stratification and compute more reliable statistics  \n- probe your hidden test data for magic characteristics (you probably didn't solve the problem but employ competition tricks to stabilized your score)\n",
    "2066160": "totally agree",
    "2064725": ""
  }
}