{
  "id": 336136,
  "title": "Hidden test set already working as intended?",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/336136",
  "author_name": "Chris X",
  "post_date": "2022-07-09T14:05:20.499000",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am a little confused. My understanding is that when submitting a prediction, the dummy test set containing only 4 rows (which actually are the first 4 rows in the training set!?) would be replaced by the actual test set (ca. 280 observations) and the code rerun on that test set. <br>\nHow can it then work (you can see that in a few notebooks, including mine) to submit just the (static) dummy test set and nevertheless get a working submission (even if the score is low)?<br>\nFurthermore I am wondering that no one has yet beaten the trivial benchmark of submitting only 0.5 probs…<br>\nAny comment highly appreciated!<br>\nGood luck and have fun with the competiton!</p>",
  "messages": [
    {
      "id": 1849431,
      "postDate": "2022-07-09T14:05:20.500Z",
      "content": "<p>I am a little confused. My understanding is that when submitting a prediction, the dummy test set containing only 4 rows (which actually are the first 4 rows in the training set!?) would be replaced by the actual test set (ca. 280 observations) and the code rerun on that test set. <br>\nHow can it then work (you can see that in a few notebooks, including mine) to submit just the (static) dummy test set and nevertheless get a working submission (even if the score is low)?<br>\nFurthermore I am wondering that no one has yet beaten the trivial benchmark of submitting only 0.5 probs…<br>\nAny comment highly appreciated!<br>\nGood luck and have fun with the competiton!</p>",
      "rawMarkdown": "I am a little confused. My understanding is that when submitting a prediction, the dummy test set containing only 4 rows (which actually are the first 4 rows in the training set!?) would be replaced by the actual test set (ca. 280 observations) and the code rerun on that test set. \nHow can it then work (you can see that in a few notebooks, including mine) to submit just the (static) dummy test set and nevertheless get a working submission (even if the score is low)?\nFurthermore I am wondering that no one has yet beaten the trivial benchmark of submitting only 0.5 probs...\nAny comment highly appreciated!\nGood luck and have fun with the competiton!",
      "votes": 3
    },
    {
      "id": 1849502,
      "postDate": "2022-07-09T15:34:42.103Z",
      "content": "<p>The (static) test set (submission file too!) as you call it is not static. </p>\n<p>If you submit the dummy <strong><code>sample_submission.csv</code></strong> file, the backend rerun will submit a different (hidden) sample_submission.csv that aligns with the different (hidden) test.csv and test images. This will submit all 0.5 for all hidden test images for both categories.</p>\n<p>The reason why no-one has beaten this benchmark yet is that there is no simple way (NOT YET) to 'trick' the LB metric as it is already weighted. If it was NOT weighted, we could simply change the likelihoods to reflect the probability distribution of etiology found in the training data.</p>\n<p>I say NOT YET, because there still may be simple unequal distributions that might allow us to trivially achieve better than 0.5.</p>\n<ul>\n<li>An example would be if larger images are more likely to be CE. We would then change probabilities to reflect this (0.75 and 0.25 respectively) for larger images… this might result in a slightly higher score as it would be slightly better than the coin-toss probabilities that exist now.</li>\n</ul>",
      "rawMarkdown": "The (static) test set (submission file too!) as you call it is not static. \n\nIf you submit the dummy **`sample_submission.csv`** file, the backend rerun will submit a different (hidden) sample_submission.csv that aligns with the different (hidden) test.csv and test images. This will submit all 0.5 for all hidden test images for both categories.\n\nThe reason why no-one has beaten this benchmark yet is that there is no simple way (NOT YET) to 'trick' the LB metric as it is already weighted. If it was NOT weighted, we could simply change the likelihoods to reflect the probability distribution of etiology found in the training data.\n\nI say NOT YET, because there still may be simple unequal distributions that might allow us to trivially achieve better than 0.5.\n* An example would be if larger images are more likely to be CE. We would then change probabilities to reflect this (0.75 and 0.25 respectively) for larger images... this might result in a slightly higher score as it would be slightly better than the coin-toss probabilities that exist now.",
      "votes": 2,
      "replies": [
        {
          "id": 1849607,
          "postDate": "2022-07-09T16:54:00.513Z",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>! That really helps!</p>",
          "rawMarkdown": "Thanks a lot @dschettler8845! That really helps!",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1849502,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2022-07-09T15:34:42.103000",
      "content": "<p>The (static) test set (submission file too!) as you call it is not static. </p>\n<p>If you submit the dummy <strong><code>sample_submission.csv</code></strong> file, the backend rerun will submit a different (hidden) sample_submission.csv that aligns with the different (hidden) test.csv and test images. This will submit all 0.5 for all hidden test images for both categories.</p>\n<p>The reason why no-one has beaten this benchmark yet is that there is no simple way (NOT YET) to 'trick' the LB metric as it is already weighted. If it was NOT weighted, we could simply change the likelihoods to reflect the probability distribution of etiology found in the training data.</p>\n<p>I say NOT YET, because there still may be simple unequal distributions that might allow us to trivially achieve better than 0.5.</p>\n<ul>\n<li>An example would be if larger images are more likely to be CE. We would then change probabilities to reflect this (0.75 and 0.25 respectively) for larger images… this might result in a slightly higher score as it would be slightly better than the coin-toss probabilities that exist now.</li>\n</ul>",
      "votes": 2,
      "replies": [
        {
          "id": 1849607,
          "author_name": "Chris X",
          "author_url": "",
          "post_date": "2022-07-09T16:54:00.513000",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>! That really helps!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1849431": "I am a little confused. My understanding is that when submitting a prediction, the dummy test set containing only 4 rows (which actually are the first 4 rows in the training set!?) would be replaced by the actual test set (ca. 280 observations) and the code rerun on that test set. \nHow can it then work (you can see that in a few notebooks, including mine) to submit just the (static) dummy test set and nevertheless get a working submission (even if the score is low)?\nFurthermore I am wondering that no one has yet beaten the trivial benchmark of submitting only 0.5 probs...\nAny comment highly appreciated!\nGood luck and have fun with the competiton!",
    "1849502": "The (static) test set (submission file too!) as you call it is not static. \n\nIf you submit the dummy **`sample_submission.csv`** file, the backend rerun will submit a different (hidden) sample_submission.csv that aligns with the different (hidden) test.csv and test images. This will submit all 0.5 for all hidden test images for both categories.\n\nThe reason why no-one has beaten this benchmark yet is that there is no simple way (NOT YET) to 'trick' the LB metric as it is already weighted. If it was NOT weighted, we could simply change the likelihoods to reflect the probability distribution of etiology found in the training data.\n\nI say NOT YET, because there still may be simple unequal distributions that might allow us to trivially achieve better than 0.5.\n* An example would be if larger images are more likely to be CE. We would then change probabilities to reflect this (0.75 and 0.25 respectively) for larger images... this might result in a slightly higher score as it would be slightly better than the coin-toss probabilities that exist now."
  }
}