{
  "competition": "cafa-5-protein-function-prediction",
  "topic_id": "421887",
  "comments": [],
  "messages": [],
  "raw_show": {
    "topic": {
      "id": 421887,
      "title": "Still Confused about Final Test Set",
      "authorName": "AbaoJiang",
      "commentCount": 4,
      "votes": 15,
      "postDate": "2023-07-07T10:23:55.357000"
    },
    "comments": [
      {
        "id": 2352162,
        "authorName": "John Mitchell",
        "votes": 2,
        "postDate": "2023-07-20T17:11:11.867000",
        "content": "<p>Simple answer is that \"1. Section Evaluation Metrics in evaluation page\" refers to the private leaderboard on which the real competition results will eventually be evaluated. \"2. Section Leaderboard in evaluation page\" refers to the public leaderboard which is used to calculate the provisional scores we see now. The wording \"These proteins will not be included in the test set for the subontologies used for the leaderboard evaluation\" means simply that they are mutually exclusive. Both public and private LB sets of proteins are subsets of the given test set, the exact composition of the private test set will depend on which proteins acquire new annotations in the relevant time period.</p>\n<p>You've set my head spinning with all the mathematical notation, but IIUC then condition 4 is the only relevant one since</p>\n<blockquote>\n  <p>These proteins will not be included in the test set for the subontologies used for the leaderboard evaluation.</p>\n</blockquote>"
      },
      {
        "id": 2352525,
        "authorName": "AbaoJiang",
        "votes": 1,
        "postDate": "2023-07-21T04:40:31.060000",
        "content": "<p>Hi <a href=\"https://www.kaggle.com/jbomitchell\" target=\"_blank\">@jbomitchell</a>,</p>\n<p>There are three test sets constructed for the final evaluation, each per namespace. Doesn't it mean that the <em>condition 3</em> still holds?</p>\n<blockquote>\n  <p>These proteins will not be included in the test set <strong>for the subontologies</strong> used for the leaderboard evaluation.</p>\n</blockquote>\n<p>If a protein is included in the public leaderboard for subontology <code>MFO</code>, then this protein may still be included in the final test set of <code>BPO</code> (if it obtains experimental annotation in <code>BPO</code> after submission deadline). Is it correct?</p>\n<p>Thanks for the comment!</p>"
      },
      {
        "id": 2352645,
        "authorName": "John Mitchell",
        "votes": 2,
        "postDate": "2023-07-21T06:47:59.577000",
        "content": "<p>I assume that the organisers say what they mean and mean what they say.</p>\n<blockquote>\n  <p>These proteins will not be included in the test set for the subontologies used for the leaderboard evaluation.</p>\n</blockquote>\n<p>[EDIT] But if they were to mean \"that particular subontology\" rather than \"the subontologies\", I'm not sure it would make much practical difference. Even if a user had LB-probed the content of the public set and had their own independent validation protocol, would being able to exclude a small fraction of proteins give a worthwhile advantage? Almost all methods suitable for this competition are those that can be applied at scale to ~140k sequences/structures. So most people use methods which cover as many proteins as possible. Ultimately, none of us know exactly what proteins will be annotated in the coming weeks or months and included in the actual test set.</p>"
      },
      {
        "id": 2351938,
        "authorName": "Correlation",
        "votes": 1,
        "postDate": "2023-07-20T13:46:17.140000",
        "content": "<p>Thanks for your share.</p>"
      }
    ]
  },
  "topic": null,
  "index": {
    "id": "421887",
    "title": "Still Confused about Final Test Set",
    "authorName": "",
    "commentCount": "4",
    "votes": "15",
    "postDate": "2023-07-07 10:23:55.357000"
  }
}