{
  "id": 117168,
  "title": "A plausible insight from stage 2 public leaderboard",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/117168",
  "author_name": "arctic",
  "post_date": "2019-11-13T18:29:08.412000",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<h3>The facts we (should) know:</h3>\n\n<ul>\n<li>the <strong>stage 2 public leaderboard</strong> is only based on <strong>&lt; 1% of the test data</strong>, and <strong>most of the data</strong> (if not all of them) are <strong>positive cases</strong> (ANY = 1). </li>\n<li>in the <strong>stage 1 test data</strong>, the <strong>ratio</strong> between <strong>positive cases</strong> (<strong>ANY = 1</strong>) and <strong>negative cases</strong> (<strong>ANY = 0</strong>) is around <strong>1 : 6</strong>. So, <strong>there are more negative cases</strong> (ANY = 0). </li>\n</ul>\n\n<h3>What I tried:</h3>\n\n<ol>\n<li><p>I <strong>took</strong> several <strong>public notebooks with good score</strong> in <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/notebooks\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/notebooks</a>. </p></li>\n<li><p>I evaluated their generated CSV based on the competition metric on 4 settings:\na) <strong><em>evaluate</em></strong> them <strong><em>on</em></strong> the <strong><em>whole stage 1 test data</em></strong>\nb) <strong><em>evaluate</em></strong> them <strong><em>only on</em></strong> the <strong><em>negative cases</em></strong> (ANY=0) of <strong>stage 1 test data</strong>\nc) <strong><em>evaluate</em></strong> them <strong><em>only on</em></strong> the <strong><em>positive cases</em></strong> (ANY=1) of <strong>stage 1 test data</strong>\nd) <strong><em>evaluate</em></strong> them on the fragment of the stage 1 test data where the <strong><em>number of positive and negative cases is balanced</em></strong> (ratio 1:1).</p></li>\n</ol>\n\n<h3>What I found and what can be deduced?</h3>\n\n<ul>\n<li><p>evaluation over <strong>the whole stage 1 test data</strong> (step 2a) --&gt; they have <strong>good score</strong> (low error)</p></li>\n<li><p>evaluation over <strong>the negative cases only</strong> (step 2b) --&gt; they have <strong>good score</strong> (low error)</p></li>\n<li><p>evaluation over <strong>the positive cases only</strong> (step 2c) --&gt; the <strong>score is getting worse</strong> (error increasing)</p></li>\n<li><p>evaluation over <strong>the balanced dataset</strong> (step 2d) --&gt; the <strong>score is getting worse</strong> (error increasing)</p></li>\n<li><p>For the last two, <strong><em>most likely the score is getting worse because there are less negative cases</em></strong>. </p></li>\n<li><p><em><strong>Since the ratio between positive and negative cases in stage 1 test set is 1 : 6</strong></em>, we might conclude that <strong><em>a model might get a good score in the stage 1 public leaderboard if they perform very well on negative cases</em></strong>. I guess this is an important aspect for getting a high score in stage 1 leaderboard.</p></li>\n<li><p>In the data for <strong><em>stage 2 public leaderboard (&lt; 1% of the whole data)</em></strong>, there are <strong><em>mostly only positive cases</em></strong>. Thus, <strong><em>this might be a plausible explanation why many people who got a good score in stage 1 leaderboard get a lower score in the stage 2 public leaderboard</em></strong>. <strong>Most likely their model perform very well on normal cases, but not that good on positive cases</strong>.    </p></li>\n<li><p><em><strong>Ideally, a good model should perform well on both positive and negative cases</strong></em>. </p></li>\n<li><p>Anyway, <strong>1% of the data is not enough to predict the final result</strong>. We still <strong>don't know exactly</strong>  about <strong>the composition of the other 99%</strong>. However, based on the experience of stage 1 test set, the <strong><em>ratio between positive and negative cases in the final test set might affect the final leaderboard if the model is not good in both cases</em></strong>. </p></li>\n<li><p>If the composition of the stage 2 test set is similar to the stage 1 test set, then I guess the final leaderboard should be similar to the stage 1 (unless people re-tune their model). If the composition of the stage 2 test set is so much different from stage 1 test set, then we might see some surprise. Thus, select your final submission wisely. </p></li>\n</ul>\n\n<p>You may repeat what I did and you might observe similar phenomena :)</p>\n\n<p>Anyway, all the best for everyone in this competition. Although it was a bit chaotic at some point, it was a nice competition.</p>\n\n<p>PS: <strong>One thing that most likely I could guarantee, none of you will be at the last spot in the final leaderboard</strong> ;) </p>",
  "messages": [
    {
      "id": 672298,
      "postDate": "2019-11-13T18:29:08.413Z",
      "content": "<h3>The facts we (should) know:</h3>\n\n<ul>\n<li>the <strong>stage 2 public leaderboard</strong> is only based on <strong>&lt; 1% of the test data</strong>, and <strong>most of the data</strong> (if not all of them) are <strong>positive cases</strong> (ANY = 1). </li>\n<li>in the <strong>stage 1 test data</strong>, the <strong>ratio</strong> between <strong>positive cases</strong> (<strong>ANY = 1</strong>) and <strong>negative cases</strong> (<strong>ANY = 0</strong>) is around <strong>1 : 6</strong>. So, <strong>there are more negative cases</strong> (ANY = 0). </li>\n</ul>\n\n<h3>What I tried:</h3>\n\n<ol>\n<li><p>I <strong>took</strong> several <strong>public notebooks with good score</strong> in <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/notebooks\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/notebooks</a>. </p></li>\n<li><p>I evaluated their generated CSV based on the competition metric on 4 settings:\na) <strong><em>evaluate</em></strong> them <strong><em>on</em></strong> the <strong><em>whole stage 1 test data</em></strong>\nb) <strong><em>evaluate</em></strong> them <strong><em>only on</em></strong> the <strong><em>negative cases</em></strong> (ANY=0) of <strong>stage 1 test data</strong>\nc) <strong><em>evaluate</em></strong> them <strong><em>only on</em></strong> the <strong><em>positive cases</em></strong> (ANY=1) of <strong>stage 1 test data</strong>\nd) <strong><em>evaluate</em></strong> them on the fragment of the stage 1 test data where the <strong><em>number of positive and negative cases is balanced</em></strong> (ratio 1:1).</p></li>\n</ol>\n\n<h3>What I found and what can be deduced?</h3>\n\n<ul>\n<li><p>evaluation over <strong>the whole stage 1 test data</strong> (step 2a) --&gt; they have <strong>good score</strong> (low error)</p></li>\n<li><p>evaluation over <strong>the negative cases only</strong> (step 2b) --&gt; they have <strong>good score</strong> (low error)</p></li>\n<li><p>evaluation over <strong>the positive cases only</strong> (step 2c) --&gt; the <strong>score is getting worse</strong> (error increasing)</p></li>\n<li><p>evaluation over <strong>the balanced dataset</strong> (step 2d) --&gt; the <strong>score is getting worse</strong> (error increasing)</p></li>\n<li><p>For the last two, <strong><em>most likely the score is getting worse because there are less negative cases</em></strong>. </p></li>\n<li><p><em><strong>Since the ratio between positive and negative cases in stage 1 test set is 1 : 6</strong></em>, we might conclude that <strong><em>a model might get a good score in the stage 1 public leaderboard if they perform very well on negative cases</em></strong>. I guess this is an important aspect for getting a high score in stage 1 leaderboard.</p></li>\n<li><p>In the data for <strong><em>stage 2 public leaderboard (&lt; 1% of the whole data)</em></strong>, there are <strong><em>mostly only positive cases</em></strong>. Thus, <strong><em>this might be a plausible explanation why many people who got a good score in stage 1 leaderboard get a lower score in the stage 2 public leaderboard</em></strong>. <strong>Most likely their model perform very well on normal cases, but not that good on positive cases</strong>.    </p></li>\n<li><p><em><strong>Ideally, a good model should perform well on both positive and negative cases</strong></em>. </p></li>\n<li><p>Anyway, <strong>1% of the data is not enough to predict the final result</strong>. We still <strong>don't know exactly</strong>  about <strong>the composition of the other 99%</strong>. However, based on the experience of stage 1 test set, the <strong><em>ratio between positive and negative cases in the final test set might affect the final leaderboard if the model is not good in both cases</em></strong>. </p></li>\n<li><p>If the composition of the stage 2 test set is similar to the stage 1 test set, then I guess the final leaderboard should be similar to the stage 1 (unless people re-tune their model). If the composition of the stage 2 test set is so much different from stage 1 test set, then we might see some surprise. Thus, select your final submission wisely. </p></li>\n</ul>\n\n<p>You may repeat what I did and you might observe similar phenomena :)</p>\n\n<p>Anyway, all the best for everyone in this competition. Although it was a bit chaotic at some point, it was a nice competition.</p>\n\n<p>PS: <strong>One thing that most likely I could guarantee, none of you will be at the last spot in the final leaderboard</strong> ;) </p>",
      "rawMarkdown": "### The facts we (should) know:\n- the **stage 2 public leaderboard** is only based on **&lt; 1% of the test data**, and **most of the data** (if not all of them) are **positive cases** (ANY = 1). \n- in the **stage 1 test data**, the **ratio** between **positive cases** (**ANY = 1**) and **negative cases** (**ANY = 0**) is around **1 : 6**. So, **there are more negative cases** (ANY = 0). \n\n### What I tried:\n1.  I **took** several **public notebooks with good score** in https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/notebooks. \n\n2.  I evaluated their generated CSV based on the competition metric on 4 settings:\n\ta) ***evaluate*** them ***on*** the ***whole stage 1 test data***\n\tb) ***evaluate*** them ***only on*** the ***negative cases*** (ANY=0) of **stage 1 test data**\n\tc) ***evaluate*** them ***only on*** the ***positive cases*** (ANY=1) of **stage 1 test data**\n\td) ***evaluate*** them on the fragment of the stage 1 test data where the ***number of positive and negative cases is balanced*** (ratio 1:1).\n\n### What I found and what can be deduced?\n- evaluation over **the whole stage 1 test data** (step 2a) --&gt; they have **good score** (low error)\n\n- evaluation over **the negative cases only** (step 2b) --&gt; they have **good score** (low error)\n\n- evaluation over **the positive cases only** (step 2c) --&gt; the **score is getting worse** (error increasing)\n\n- evaluation over **the balanced dataset** (step 2d) --&gt; the **score is getting worse** (error increasing)\n\n- For the last two, ***most likely the score is getting worse because there are less negative cases***. \n\n- ***Since the ratio between positive and negative cases in stage 1 test set is 1 : 6***, we might conclude that ***a model might get a good score in the stage 1 public leaderboard if they perform very well on negative cases***. I guess this is an important aspect for getting a high score in stage 1 leaderboard.\n\n- In the data for ***stage 2 public leaderboard (&lt; 1% of the whole data)***, there are ***mostly only positive cases***. Thus, ***this might be a plausible explanation why many people who got a good score in stage 1 leaderboard get a lower score in the stage 2 public leaderboard***. **Most likely their model perform very well on normal cases, but not that good on positive cases**.    \n\n- ***Ideally, a good model should perform well on both positive and negative cases***. \n\n- Anyway, **1% of the data is not enough to predict the final result**. We still **don't know exactly**  about **the composition of the other 99%**. However, based on the experience of stage 1 test set, the ***ratio between positive and negative cases in the final test set might affect the final leaderboard if the model is not good in both cases***. \n\n- If the composition of the stage 2 test set is similar to the stage 1 test set, then I guess the final leaderboard should be similar to the stage 1 (unless people re-tune their model). If the composition of the stage 2 test set is so much different from stage 1 test set, then we might see some surprise. Thus, select your final submission wisely. \n\nYou may repeat what I did and you might observe similar phenomena :)\n\nAnyway, all the best for everyone in this competition. Although it was a bit chaotic at some point, it was a nice competition.\n\nPS: **One thing that most likely I could guarantee, none of you will be at the last spot in the final leaderboard** ;) \n",
      "votes": 3
    },
    {
      "id": 672332,
      "postDate": "2019-11-13T19:22:02.280Z",
      "content": "<p>Hi <a href=\"/anginscribe\">@anginscribe</a> interresting analysis. Looking forward to see more details after the competition has ended and the stage2 test set is also known.</p>\n\n<p>To add to this...not only select your final submission wisely...but also make sure that it is inline with the procedure that you followed during stage 1 and what has been 'locked' into the model upload.\nIf there is a shakeup and some of us might end up in the top regions it would be dissapointing if it is not approved because someone did not follow the procedures as described in there model upload.</p>",
      "rawMarkdown": "Hi @anginscribe interresting analysis. Looking forward to see more details after the competition has ended and the stage2 test set is also known.\n\nTo add to this...not only select your final submission wisely...but also make sure that it is inline with the procedure that you followed during stage 1 and what has been 'locked' into the model upload.\nIf there is a shakeup and some of us might end up in the top regions it would be dissapointing if it is not approved because someone did not follow the procedures as described in there model upload.",
      "votes": 1,
      "replies": [
        {
          "id": 672399,
          "postDate": "2019-11-13T20:51:54.917Z",
          "content": "<p>That's absolutely true. I couldn't agree more. </p>",
          "rawMarkdown": "That's absolutely true. I couldn't agree more. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 672330,
      "postDate": "2019-11-13T19:16:58.220Z",
      "content": "<p><a href=\"/arctic\">@arctic</a> thanks\nlet me tell one more observation.. in stage1 more aug was leading to the poor score.. in stage2 more aug are leading to huge jump in the score ( rotations ,hz flip). Aug same as train set. </p>\n\n<p>From this it appears positive cases are well detected using aug compared to the negative cases.</p>",
      "rawMarkdown": "@arctic thanks\nlet me tell one more observation.. in stage1 more aug was leading to the poor score.. in stage2 more aug are leading to huge jump in the score ( rotations ,hz flip). Aug same as train set. \n\nFrom this it appears positive cases are well detected using aug compared to the negative cases.\n\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 672401,
          "postDate": "2019-11-13T20:53:05.647Z",
          "content": "<p>Interesting! I didn't try that.</p>",
          "rawMarkdown": "Interesting! I didn't try that."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 672332,
      "author_name": "Robin Smits",
      "author_url": "",
      "post_date": "2019-11-13T19:22:02.280000",
      "content": "<p>Hi <a href=\"/anginscribe\">@anginscribe</a> interresting analysis. Looking forward to see more details after the competition has ended and the stage2 test set is also known.</p>\n\n<p>To add to this...not only select your final submission wisely...but also make sure that it is inline with the procedure that you followed during stage 1 and what has been 'locked' into the model upload.\nIf there is a shakeup and some of us might end up in the top regions it would be dissapointing if it is not approved because someone did not follow the procedures as described in there model upload.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 672399,
          "author_name": "arctic",
          "author_url": "",
          "post_date": "2019-11-13T20:51:54.917000",
          "content": "<p>That's absolutely true. I couldn't agree more. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 672330,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2019-11-13T19:16:58.220000",
      "content": "<p><a href=\"/arctic\">@arctic</a> thanks\nlet me tell one more observation.. in stage1 more aug was leading to the poor score.. in stage2 more aug are leading to huge jump in the score ( rotations ,hz flip). Aug same as train set. </p>\n\n<p>From this it appears positive cases are well detected using aug compared to the negative cases.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 672401,
          "author_name": "arctic",
          "author_url": "",
          "post_date": "2019-11-13T20:53:05.647000",
          "content": "<p>Interesting! I didn't try that.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "672298": "### The facts we (should) know:\n- the **stage 2 public leaderboard** is only based on **&lt; 1% of the test data**, and **most of the data** (if not all of them) are **positive cases** (ANY = 1). \n- in the **stage 1 test data**, the **ratio** between **positive cases** (**ANY = 1**) and **negative cases** (**ANY = 0**) is around **1 : 6**. So, **there are more negative cases** (ANY = 0). \n\n### What I tried:\n1.  I **took** several **public notebooks with good score** in https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/notebooks. \n\n2.  I evaluated their generated CSV based on the competition metric on 4 settings:\n\ta) ***evaluate*** them ***on*** the ***whole stage 1 test data***\n\tb) ***evaluate*** them ***only on*** the ***negative cases*** (ANY=0) of **stage 1 test data**\n\tc) ***evaluate*** them ***only on*** the ***positive cases*** (ANY=1) of **stage 1 test data**\n\td) ***evaluate*** them on the fragment of the stage 1 test data where the ***number of positive and negative cases is balanced*** (ratio 1:1).\n\n### What I found and what can be deduced?\n- evaluation over **the whole stage 1 test data** (step 2a) --&gt; they have **good score** (low error)\n\n- evaluation over **the negative cases only** (step 2b) --&gt; they have **good score** (low error)\n\n- evaluation over **the positive cases only** (step 2c) --&gt; the **score is getting worse** (error increasing)\n\n- evaluation over **the balanced dataset** (step 2d) --&gt; the **score is getting worse** (error increasing)\n\n- For the last two, ***most likely the score is getting worse because there are less negative cases***. \n\n- ***Since the ratio between positive and negative cases in stage 1 test set is 1 : 6***, we might conclude that ***a model might get a good score in the stage 1 public leaderboard if they perform very well on negative cases***. I guess this is an important aspect for getting a high score in stage 1 leaderboard.\n\n- In the data for ***stage 2 public leaderboard (&lt; 1% of the whole data)***, there are ***mostly only positive cases***. Thus, ***this might be a plausible explanation why many people who got a good score in stage 1 leaderboard get a lower score in the stage 2 public leaderboard***. **Most likely their model perform very well on normal cases, but not that good on positive cases**.    \n\n- ***Ideally, a good model should perform well on both positive and negative cases***. \n\n- Anyway, **1% of the data is not enough to predict the final result**. We still **don't know exactly**  about **the composition of the other 99%**. However, based on the experience of stage 1 test set, the ***ratio between positive and negative cases in the final test set might affect the final leaderboard if the model is not good in both cases***. \n\n- If the composition of the stage 2 test set is similar to the stage 1 test set, then I guess the final leaderboard should be similar to the stage 1 (unless people re-tune their model). If the composition of the stage 2 test set is so much different from stage 1 test set, then we might see some surprise. Thus, select your final submission wisely. \n\nYou may repeat what I did and you might observe similar phenomena :)\n\nAnyway, all the best for everyone in this competition. Although it was a bit chaotic at some point, it was a nice competition.\n\nPS: **One thing that most likely I could guarantee, none of you will be at the last spot in the final leaderboard** ;) \n",
    "672332": "Hi @anginscribe interresting analysis. Looking forward to see more details after the competition has ended and the stage2 test set is also known.\n\nTo add to this...not only select your final submission wisely...but also make sure that it is inline with the procedure that you followed during stage 1 and what has been 'locked' into the model upload.\nIf there is a shakeup and some of us might end up in the top regions it would be dissapointing if it is not approved because someone did not follow the procedures as described in there model upload.",
    "672330": "@arctic thanks\nlet me tell one more observation.. in stage1 more aug was leading to the poor score.. in stage2 more aug are leading to huge jump in the score ( rotations ,hz flip). Aug same as train set. \n\nFrom this it appears positive cases are well detected using aug compared to the negative cases.\n\n\n"
  }
}