{
  "id": 115901,
  "title": "public  leaderboard is calculated with <1% of the test data",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/115901",
  "author_name": "Mobassir",
  "post_date": "2019-11-06T02:34:18.745000",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>what is  the purpose of using less than 1% data for public leaderboard? i saw similar issue in pneumothorax competition,i am curiously wanting to know \"what are the barriers of not using &gt;10% and &lt;25% data for public  leaderboard in stage-2?</p>",
  "messages": [
    {
      "id": 666476,
      "postDate": "2019-11-06T05:59:26.283Z",
      "content": "<p>I think the purpose of using less than 1% of data for stage2 public leaderboard is to prevent pseudo labelling and only let competitors know whether they successfully submit the stage-2 result file.</p>",
      "rawMarkdown": "I think the purpose of using less than 1% of data for stage2 public leaderboard is to prevent pseudo labelling and only let competitors know whether they successfully submit the stage-2 result file.",
      "votes": 1,
      "replies": [
        {
          "id": 666483,
          "postDate": "2019-11-06T06:04:34.200Z",
          "content": "<p>maybe you are right,thanks for the confirmation</p>",
          "rawMarkdown": "maybe you are right,thanks for the confirmation"
        }
      ]
    },
    {
      "id": 666337,
      "postDate": "2019-11-06T02:34:18.747Z",
      "content": "<p>what is  the purpose of using less than 1% data for public leaderboard? i saw similar issue in pneumothorax competition,i am curiously wanting to know \"what are the barriers of not using &gt;10% and &lt;25% data for public  leaderboard in stage-2?</p>",
      "rawMarkdown": "what is  the purpose of using less than 1% data for public leaderboard? i saw similar issue in pneumothorax competition,i am curiously wanting to know \"what are the barriers of not using &gt;10% and &lt;25% data for public  leaderboard in stage-2?",
      "votes": 1
    },
    {
      "id": 666355,
      "postDate": "2019-11-06T02:55:20.043Z",
      "content": "<p>There should be absolutely no scientific changes between your stage 1 predictions and stage two predictions. </p>\n\n<p>So why even use 1%, instead of 0%? I guess so that people have some feedback that their predictions actually work.</p>\n\n<p>I will turn it around on you and ask: What would be the purpose of using 10% or 25%? What would be the motivation there? It is not like you are allowed to improve the model any more using the leader board results, so for those that are not in the money and will probably not have their submissions checked against their submitted code, providing more data in the public leader board just gives them more opportunity to get away with over fitting to stage 2.</p>",
      "rawMarkdown": "There should be absolutely no scientific changes between your stage 1 predictions and stage two predictions. \n\nSo why even use 1%, instead of 0%? I guess so that people have some feedback that their predictions actually work.\n\nI will turn it around on you and ask: What would be the purpose of using 10% or 25%? What would be the motivation there? It is not like you are allowed to improve the model any more using the leader board results, so for those that are not in the money and will probably not have their submissions checked against their submitted code, providing more data in the public leader board just gives them more opportunity to get away with over fitting to stage 2.",
      "replies": [
        {
          "id": 666365,
          "postDate": "2019-11-06T03:22:32.647Z",
          "content": "<p>then there is also no reason to set 7  days for stage 2?</p>",
          "rawMarkdown": "then there is also no reason to set 7  days for stage 2?"
        },
        {
          "id": 666488,
          "postDate": "2019-11-06T06:13:00.257Z",
          "content": "<p>You are allowed to retrain your model using the stage 1 labels, however it must be done with the code that was uploaded and no hyperparameter tuning is allowed unless it is automated in the code, so I guess the 7 days is to give people time to do that.. </p>",
          "rawMarkdown": "You are allowed to retrain your model using the stage 1 labels, however it must be done with the code that was uploaded and no hyperparameter tuning is allowed unless it is automated in the code, so I guess the 7 days is to give people time to do that.. "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 666476,
      "author_name": "Zehao Zhao",
      "author_url": "",
      "post_date": "2019-11-06T05:59:26.283000",
      "content": "<p>I think the purpose of using less than 1% of data for stage2 public leaderboard is to prevent pseudo labelling and only let competitors know whether they successfully submit the stage-2 result file.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 666483,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2019-11-06T06:04:34.200000",
          "content": "<p>maybe you are right,thanks for the confirmation</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 666355,
      "author_name": "cherring",
      "author_url": "",
      "post_date": "2019-11-06T02:55:20.043000",
      "content": "<p>There should be absolutely no scientific changes between your stage 1 predictions and stage two predictions. </p>\n\n<p>So why even use 1%, instead of 0%? I guess so that people have some feedback that their predictions actually work.</p>\n\n<p>I will turn it around on you and ask: What would be the purpose of using 10% or 25%? What would be the motivation there? It is not like you are allowed to improve the model any more using the leader board results, so for those that are not in the money and will probably not have their submissions checked against their submitted code, providing more data in the public leader board just gives them more opportunity to get away with over fitting to stage 2.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 666365,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2019-11-06T03:22:32.647000",
          "content": "<p>then there is also no reason to set 7  days for stage 2?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 666488,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "2019-11-06T06:13:00.257000",
          "content": "<p>You are allowed to retrain your model using the stage 1 labels, however it must be done with the code that was uploaded and no hyperparameter tuning is allowed unless it is automated in the code, so I guess the 7 days is to give people time to do that.. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "666476": "I think the purpose of using less than 1% of data for stage2 public leaderboard is to prevent pseudo labelling and only let competitors know whether they successfully submit the stage-2 result file.",
    "666337": "what is  the purpose of using less than 1% data for public leaderboard? i saw similar issue in pneumothorax competition,i am curiously wanting to know \"what are the barriers of not using &gt;10% and &lt;25% data for public  leaderboard in stage-2?",
    "666355": "There should be absolutely no scientific changes between your stage 1 predictions and stage two predictions. \n\nSo why even use 1%, instead of 0%? I guess so that people have some feedback that their predictions actually work.\n\nI will turn it around on you and ask: What would be the purpose of using 10% or 25%? What would be the motivation there? It is not like you are allowed to improve the model any more using the leader board results, so for those that are not in the money and will probably not have their submissions checked against their submitted code, providing more data in the public leader board just gives them more opportunity to get away with over fitting to stage 2."
  }
}