{
  "id": 168623,
  "title": "Private Test Set (timeouts in final submissions?)",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/168623",
  "author_name": "fergusoci",
  "post_date": "2020-07-21T09:52:11.180000",
  "votes": 2,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Wondering a bit about the private test set… Can anyone clarify/correct my understanding of things:</p>\n<ul>\n<li>The public test set has close to 1000 samples. When we make a submission it runs over these samples and its score is published on the public LB.</li>\n<li>The private test set is around 1000 samples also, but only our final selected submissions will be run on this, and these will generate our private LB score.</li>\n</ul>\n<p>If this is the case, how can we be certain that our submissions will not time out on the private set? If we're getting close to the 6 hr limit on the public set, should we be concerned?</p>",
  "messages": [
    {
      "id": 938125,
      "postDate": "2020-07-21T11:05:06.337Z",
      "content": "<p>I think when you submit the notebook it makes the predictions on the whole test set (with roughly 1000 samples). Then from the predictions, it computes both the public LB and the private LB scores using 42% and the remaining 58% of the data respectively. The private score is already there, it's just hidden.</p>",
      "rawMarkdown": "I think when you submit the notebook it makes the predictions on the whole test set (with roughly 1000 samples). Then from the predictions, it computes both the public LB and the private LB scores using 42% and the remaining 58% of the data respectively. The private score is already there, it's just hidden.",
      "votes": 4,
      "replies": [
        {
          "id": 938130,
          "postDate": "2020-07-21T11:09:59.877Z",
          "content": "<p>Exactly. If I understand this correctly, whenever you make a submission, it performs inference and gets evaluated on both public and private parts of the test set. As long as you see your public score after submission, you should not worry about getting a timeout since your code has already been executed on the full test set. </p>",
          "rawMarkdown": "Exactly. If I understand this correctly, whenever you make a submission, it performs inference and gets evaluated on both public and private parts of the test set. As long as you see your public score after submission, you should not worry about getting a timeout since your code has already been executed on the full test set. "
        },
        {
          "id": 938171,
          "postDate": "2020-07-21T11:38:08.953Z",
          "content": "<p>Are you sure? The rules <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/data\" target=\"_blank\">here</a> seem to suggest the private test set consists of 1000 samples. </p>\n<p>\"You can expect roughly 1,000 images in the hidden test set. Note that slightly different procedures were in place for the images used in the test set than the training set. Some of the training set images have stray pen marks on them, but the test set slides are free of pen marks.\"</p>\n<p>I can see that the public test set contains about 1000 samples (based on runtimes).</p>",
          "rawMarkdown": "Are you sure? The rules [here](https://www.kaggle.com/c/prostate-cancer-grade-assessment/data) seem to suggest the private test set consists of 1000 samples. \n\n\"You can expect roughly 1,000 images in the hidden test set. Note that slightly different procedures were in place for the images used in the test set than the training set. Some of the training set images have stray pen marks on them, but the test set slides are free of pen marks.\"\n\nI can see that the public test set contains about 1000 samples (based on runtimes)."
        },
        {
          "id": 938181,
          "postDate": "2020-07-21T11:47:27.317Z",
          "content": "<p><a href=\"/fergusoci\">@fergusoci</a> the rules state that there are about 1000 images in the “hidden” test set. I interpret “hidden” as public + private, which would mean that the code is run on the full test set. But it would be great to get confirmation from the organizers to be sure 🙂 </p>",
          "rawMarkdown": "@fergusoci the rules state that there are about 1000 images in the “hidden” test set. I interpret “hidden” as public + private, which would mean that the code is run on the full test set. But it would be great to get confirmation from the organizers to be sure 🙂 ",
          "votes": 2
        },
        {
          "id": 938183,
          "postDate": "2020-07-21T11:48:07.390Z",
          "content": "<p>They call it the \"hidden\" test set because you can't access it (while in some other competitions you can), not because it's the one determining the private LB score. Anyway, I'm a Kaggle newbie so I may be wrong :D</p>",
          "rawMarkdown": "They call it the \"hidden\" test set because you can't access it (while in some other competitions you can), not because it's the one determining the private LB score. Anyway, I'm a Kaggle newbie so I may be wrong :D"
        },
        {
          "id": 938190,
          "postDate": "2020-07-21T11:50:24.987Z",
          "content": "<p>Out of interest, <a href=\"https://www.kaggle.com/pasqualed\" target=\"_blank\">@pasqualed</a>, where did you get the info on 42% public vs 58% private? A private set using QWK with only 580 samples… This is going to get rocky!!</p>",
          "rawMarkdown": "Out of interest, @pasqualed, where did you get the info on 42% public vs 58% private? A private set using QWK with only 580 samples... This is going to get rocky!!"
        },
        {
          "id": 938212,
          "postDate": "2020-07-21T12:03:47.483Z",
          "content": "<p><a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> from the Leaderboard page:  \"This leaderboard is calculated with approximately 42% of the test data. The final results will be based on the other 58%, so the final standings may be different.\"</p>",
          "rawMarkdown": "@fergusoci from the Leaderboard page:  \"This leaderboard is calculated with approximately 42% of the test data. The final results will be based on the other 58%, so the final standings may be different.\""
        },
        {
          "id": 938298,
          "postDate": "2020-07-21T12:45:42.110Z",
          "content": "<p>Ha! Missed that somehow…🤔</p>",
          "rawMarkdown": "Ha! Missed that somehow...🤔"
        },
        {
          "id": 938740,
          "postDate": "2020-07-21T18:05:24.450Z",
          "content": "<p>I didnt know the sample size is so small. This will be indeed rocky with QWK…</p>",
          "rawMarkdown": "I didnt know the sample size is so small. This will be indeed rocky with QWK..."
        }
      ]
    },
    {
      "id": 938015,
      "postDate": "2020-07-21T09:52:11.180Z",
      "content": "<p>Wondering a bit about the private test set… Can anyone clarify/correct my understanding of things:</p>\n<ul>\n<li>The public test set has close to 1000 samples. When we make a submission it runs over these samples and its score is published on the public LB.</li>\n<li>The private test set is around 1000 samples also, but only our final selected submissions will be run on this, and these will generate our private LB score.</li>\n</ul>\n<p>If this is the case, how can we be certain that our submissions will not time out on the private set? If we're getting close to the 6 hr limit on the public set, should we be concerned?</p>",
      "rawMarkdown": "Wondering a bit about the private test set... Can anyone clarify/correct my understanding of things:\n\n- The public test set has close to 1000 samples. When we make a submission it runs over these samples and its score is published on the public LB.\n- The private test set is around 1000 samples also, but only our final selected submissions will be run on this, and these will generate our private LB score.\n\nIf this is the case, how can we be certain that our submissions will not time out on the private set? If we're getting close to the 6 hr limit on the public set, should we be concerned?",
      "votes": 2
    },
    {
      "id": 938117,
      "postDate": "2020-07-21T11:01:57.420Z",
      "content": "<p>If you run a model in your kernel you could seperate this into a new kernel, comitt it and then take the result as input to your submission kernel. Then as I have experienced it is very quickly submitted.</p>",
      "rawMarkdown": "If you run a model in your kernel you could seperate this into a new kernel, comitt it and then take the result as input to your submission kernel. Then as I have experienced it is very quickly submitted."
    }
  ],
  "comments": [
    {
      "id": 938125,
      "author_name": "Pasquale",
      "author_url": "",
      "post_date": "2020-07-21T11:05:06.337000",
      "content": "<p>I think when you submit the notebook it makes the predictions on the whole test set (with roughly 1000 samples). Then from the predictions, it computes both the public LB and the private LB scores using 42% and the remaining 58% of the data respectively. The private score is already there, it's just hidden.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 938130,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2020-07-21T11:09:59.877000",
          "content": "<p>Exactly. If I understand this correctly, whenever you make a submission, it performs inference and gets evaluated on both public and private parts of the test set. As long as you see your public score after submission, you should not worry about getting a timeout since your code has already been executed on the full test set. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 938171,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-07-21T11:38:08.953000",
          "content": "<p>Are you sure? The rules <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/data\" target=\"_blank\">here</a> seem to suggest the private test set consists of 1000 samples. </p>\n<p>\"You can expect roughly 1,000 images in the hidden test set. Note that slightly different procedures were in place for the images used in the test set than the training set. Some of the training set images have stray pen marks on them, but the test set slides are free of pen marks.\"</p>\n<p>I can see that the public test set contains about 1000 samples (based on runtimes).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 938181,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2020-07-21T11:47:27.317000",
          "content": "<p><a href=\"/fergusoci\">@fergusoci</a> the rules state that there are about 1000 images in the “hidden” test set. I interpret “hidden” as public + private, which would mean that the code is run on the full test set. But it would be great to get confirmation from the organizers to be sure 🙂 </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 938183,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-07-21T11:48:07.390000",
          "content": "<p>They call it the \"hidden\" test set because you can't access it (while in some other competitions you can), not because it's the one determining the private LB score. Anyway, I'm a Kaggle newbie so I may be wrong :D</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 938190,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-07-21T11:50:24.987000",
          "content": "<p>Out of interest, <a href=\"https://www.kaggle.com/pasqualed\" target=\"_blank\">@pasqualed</a>, where did you get the info on 42% public vs 58% private? A private set using QWK with only 580 samples… This is going to get rocky!!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 938212,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-07-21T12:03:47.483000",
          "content": "<p><a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> from the Leaderboard page:  \"This leaderboard is calculated with approximately 42% of the test data. The final results will be based on the other 58%, so the final standings may be different.\"</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 938298,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-07-21T12:45:42.110000",
          "content": "<p>Ha! Missed that somehow…🤔</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 938740,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-21T18:05:24.450000",
          "content": "<p>I didnt know the sample size is so small. This will be indeed rocky with QWK…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 938117,
      "author_name": "Andreas Horlbeck",
      "author_url": "",
      "post_date": "2020-07-21T11:01:57.420000",
      "content": "<p>If you run a model in your kernel you could seperate this into a new kernel, comitt it and then take the result as input to your submission kernel. Then as I have experienced it is very quickly submitted.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "938125": "I think when you submit the notebook it makes the predictions on the whole test set (with roughly 1000 samples). Then from the predictions, it computes both the public LB and the private LB scores using 42% and the remaining 58% of the data respectively. The private score is already there, it's just hidden.",
    "938015": "Wondering a bit about the private test set... Can anyone clarify/correct my understanding of things:\n\n- The public test set has close to 1000 samples. When we make a submission it runs over these samples and its score is published on the public LB.\n- The private test set is around 1000 samples also, but only our final selected submissions will be run on this, and these will generate our private LB score.\n\nIf this is the case, how can we be certain that our submissions will not time out on the private set? If we're getting close to the 6 hr limit on the public set, should we be concerned?",
    "938117": "If you run a model in your kernel you could seperate this into a new kernel, comitt it and then take the result as input to your submission kernel. Then as I have experienced it is very quickly submitted."
  }
}