{
  "id": 158687,
  "title": "An efficient way to probe LB",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/158687",
  "author_name": "Bo",
  "post_date": "2020-06-15T05:55:06.039000",
  "votes": 37,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I don't know if anyone already posted this, but I discovered a more efficient way to probe LB in notebook competitions than binary probe (success vs fail).</p>\n\n<p>Suppose I already had <code>n</code> submissions with different scores. They could be my <code>n</code> models, or even <code>n</code> public notebooks.</p>\n\n<p>Then I can do n-way search in one submission, i.e. divide search space into <code>n</code> conditions:\n<code>\nif test_data_size &lt; 500: run model 1\nelif test_data_size &lt; 600: run model 2\nelif test_data_size &lt; 700: run model 3\n...\n</code>\nThen I will know which condition was hit from the score of the probe submission. For example if the probe sub scored the same as model 2 in above example, then <code>test_data_size</code> would be within [500, 600)</p>\n\n<p>I found out the test data size in this competition is in <strong>[933,940]</strong> range, to save you a few submissions. I also found out Karolinska ratio is in range [0.4, 0.45)</p>\n\n<p>Please share if you further reduce the ranges.</p>",
  "messages": [
    {
      "id": 886548,
      "postDate": "2020-06-15T05:55:06.040Z",
      "content": "<p>I don't know if anyone already posted this, but I discovered a more efficient way to probe LB in notebook competitions than binary probe (success vs fail).</p>\n\n<p>Suppose I already had <code>n</code> submissions with different scores. They could be my <code>n</code> models, or even <code>n</code> public notebooks.</p>\n\n<p>Then I can do n-way search in one submission, i.e. divide search space into <code>n</code> conditions:\n<code>\nif test_data_size &lt; 500: run model 1\nelif test_data_size &lt; 600: run model 2\nelif test_data_size &lt; 700: run model 3\n...\n</code>\nThen I will know which condition was hit from the score of the probe submission. For example if the probe sub scored the same as model 2 in above example, then <code>test_data_size</code> would be within [500, 600)</p>\n\n<p>I found out the test data size in this competition is in <strong>[933,940]</strong> range, to save you a few submissions. I also found out Karolinska ratio is in range [0.4, 0.45)</p>\n\n<p>Please share if you further reduce the ranges.</p>",
      "rawMarkdown": "I don't know if anyone already posted this, but I discovered a more efficient way to probe LB in notebook competitions than binary probe (success vs fail).\n\nSuppose I already had `n` submissions with different scores. They could be my `n` models, or even `n` public notebooks.\n\nThen I can do n-way search in one submission, i.e. divide search space into `n` conditions:\n```\nif test_data_size &lt; 500: run model 1\nelif test_data_size &lt; 600: run model 2\nelif test_data_size &lt; 700: run model 3\n...\n```\nThen I will know which condition was hit from the score of the probe submission. For example if the probe sub scored the same as model 2 in above example, then `test_data_size` would be within [500, 600)\n\nI found out the test data size in this competition is in **[933,940]** range, to save you a few submissions. I also found out Karolinska ratio is in range [0.4, 0.45)\n\nPlease share if you further reduce the ranges.",
      "votes": 36
    },
    {
      "id": 891422,
      "postDate": "2020-06-18T07:07:51.977Z",
      "content": "<p>Nice!</p>",
      "rawMarkdown": "Nice!"
    },
    {
      "id": 889481,
      "postDate": "2020-06-17T01:22:24.630Z",
      "content": "<p>Haha, nice injection.</p>",
      "rawMarkdown": "Haha, nice injection."
    },
    {
      "id": 887168,
      "postDate": "2020-06-15T14:12:28.850Z",
      "content": "<p>Hi ! Nice approach, I was trying stuff with \"sleep()\", but then you have to stay around the computer to know exactly how long it took to end, so not optimal in terms of personal time use :D</p>\n\n<p>What do you mean by test data \"size\" ? size in Mb ? I think the general shapes of the images would be of more interest, wouldn't it ? </p>",
      "rawMarkdown": "Hi ! Nice approach, I was trying stuff with \"sleep()\", but then you have to stay around the computer to know exactly how long it took to end, so not optimal in terms of personal time use :D\n\nWhat do you mean by test data \"size\" ? size in Mb ? I think the general shapes of the images would be of more interest, wouldn't it ? ",
      "replies": [
        {
          "id": 887190,
          "postDate": "2020-06-15T14:26:18.320Z",
          "content": "<p>I think he refers to the number of images in the test set, since the data description says we \"can expect roughly 1,000 images in the hidden test set\"</p>",
          "rawMarkdown": "I think he refers to the number of images in the test set, since the data description says we \"can expect roughly 1,000 images in the hidden test set\""
        },
        {
          "id": 887242,
          "postDate": "2020-06-15T15:07:35.747Z",
          "content": "<p>Yes I meant the number of rows in <code>test.csv</code>, but you can use this general approach to probe other test data information such as image shapes</p>",
          "rawMarkdown": "Yes I meant the number of rows in `test.csv`, but you can use this general approach to probe other test data information such as image shapes",
          "votes": 2
        }
      ]
    },
    {
      "id": 888498,
      "postDate": "2020-06-16T11:58:21.697Z",
      "content": "<p>nice bro thanks foe sharing</p>",
      "rawMarkdown": "nice bro thanks foe sharing"
    },
    {
      "id": 888081,
      "postDate": "2020-06-16T05:39:43.077Z",
      "content": "<p>nice approach!! thanks for sharing!!</p>",
      "rawMarkdown": "nice approach!! thanks for sharing!!"
    },
    {
      "id": 886793,
      "postDate": "2020-06-15T09:21:13.243Z",
      "content": "<p>Thanks for sharing, nice approach</p>",
      "rawMarkdown": "Thanks for sharing, nice approach",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 891422,
      "author_name": "White_Walker",
      "author_url": "",
      "post_date": "2020-06-18T07:07:51.977000",
      "content": "<p>Nice!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 889481,
      "author_name": "Chenglu",
      "author_url": "",
      "post_date": "2020-06-17T01:22:24.630000",
      "content": "<p>Haha, nice injection.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 887168,
      "author_name": "Benjamin Dubreu",
      "author_url": "",
      "post_date": "2020-06-15T14:12:28.850000",
      "content": "<p>Hi ! Nice approach, I was trying stuff with \"sleep()\", but then you have to stay around the computer to know exactly how long it took to end, so not optimal in terms of personal time use :D</p>\n\n<p>What do you mean by test data \"size\" ? size in Mb ? I think the general shapes of the images would be of more interest, wouldn't it ? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 887190,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-06-15T14:26:18.320000",
          "content": "<p>I think he refers to the number of images in the test set, since the data description says we \"can expect roughly 1,000 images in the hidden test set\"</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 887242,
          "author_name": "Bo",
          "author_url": "",
          "post_date": "2020-06-15T15:07:35.747000",
          "content": "<p>Yes I meant the number of rows in <code>test.csv</code>, but you can use this general approach to probe other test data information such as image shapes</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 888498,
      "author_name": "Akash kumar",
      "author_url": "",
      "post_date": "2020-06-16T11:58:21.697000",
      "content": "<p>nice bro thanks foe sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 888081,
      "author_name": "Anjali",
      "author_url": "",
      "post_date": "2020-06-16T05:39:43.077000",
      "content": "<p>nice approach!! thanks for sharing!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 886793,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-15T09:21:13.243000",
      "content": "<p>Thanks for sharing, nice approach</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "886548": "I don't know if anyone already posted this, but I discovered a more efficient way to probe LB in notebook competitions than binary probe (success vs fail).\n\nSuppose I already had `n` submissions with different scores. They could be my `n` models, or even `n` public notebooks.\n\nThen I can do n-way search in one submission, i.e. divide search space into `n` conditions:\n```\nif test_data_size &lt; 500: run model 1\nelif test_data_size &lt; 600: run model 2\nelif test_data_size &lt; 700: run model 3\n...\n```\nThen I will know which condition was hit from the score of the probe submission. For example if the probe sub scored the same as model 2 in above example, then `test_data_size` would be within [500, 600)\n\nI found out the test data size in this competition is in **[933,940]** range, to save you a few submissions. I also found out Karolinska ratio is in range [0.4, 0.45)\n\nPlease share if you further reduce the ranges.",
    "891422": "Nice!",
    "889481": "Haha, nice injection.",
    "887168": "Hi ! Nice approach, I was trying stuff with \"sleep()\", but then you have to stay around the computer to know exactly how long it took to end, so not optimal in terms of personal time use :D\n\nWhat do you mean by test data \"size\" ? size in Mb ? I think the general shapes of the images would be of more interest, wouldn't it ? ",
    "888498": "nice bro thanks foe sharing",
    "888081": "nice approach!! thanks for sharing!!",
    "886793": "Thanks for sharing, nice approach"
  }
}