{
  "topic": {
    "id": 614668,
    "title": "Creating a time-shifted validation environment (CAFA5)",
    "authorName": "xaxipiruli",
    "commentCount": 16,
    "votes": 16,
    "postDate": "2025-11-05T14:49:56.603000"
  },
  "comments": [
    {
      "id": 3311788,
      "authorName": "An Phan",
      "votes": 0,
      "postDate": "2025-11-05T19:08:16.470000",
      "content": "<p><a href=\"https://www.kaggle.com/claradepaolis\" target=\"_blank\">@claradepaolis</a> Can you help answer these questions, especially the second bullet point?</p>"
    },
    {
      "id": 3312004,
      "authorName": "common-people",
      "votes": 0,
      "postDate": "2025-11-06T07:31:53.640000",
      "content": "<p>Also, I have a question, because in the evaluator provided in the github produced a score which is much higher than online platform.How to adjust it? By setting topK to the evaluator or?</p>"
    },
    {
      "id": 3312155,
      "authorName": "An Phan",
      "votes": 4,
      "postDate": "2025-11-06T15:17:10.640000",
      "content": "<p>What are you using as the ground truth file? That is likely the reason why the score differs. Other arguments to reproduce the score on the leaderboard are: \n<code>-max_terms 500 -prop fill -norm cafa -no_orphans</code></p>"
    },
    {
      "id": 3312293,
      "authorName": "common-people",
      "votes": 0,
      "postDate": "2025-11-06T23:42:48.863000",
      "content": "<p>Oh, it is helpful, I think!</p>"
    },
    {
      "id": 3312420,
      "authorName": "Diogo R. Ferreira",
      "votes": 1,
      "postDate": "2025-11-07T06:43:47.053000",
      "content": "<p>I thought <code>max_terms</code> would be <code>1500</code>:<br>\n<em>Finally, to limit prediction file sizes, one target cannot be associated with more than 1500 terms for MF, BP, and CC subontologies combined.</em><br>\n(Overview &gt; Evaluation &gt; Submission File)</p>\n<p>EDIT: Wait, <code>-max_terms</code> is the <em>target limit for every namespace</em> (where \"namespace\" refers to subontology/aspect). So <code>-max_terms 500</code> means 3x500=1500, but the limit is 500 terms per subontology (i.e. 500 terms per protein and subontology).</p>"
    },
    {
      "id": 3312626,
      "authorName": "An Phan",
      "votes": 1,
      "postDate": "2025-11-07T15:18:12.280000",
      "content": "<blockquote>\n  <p>but the limit is 500 terms per subontology (i.e. 500 terms per protein and subontology) </p>\n</blockquote>\n<p>This 500 terms/subontology limit is imposed by <code>-max_terms</code> argument in CAFA-evaluator package</p>\n<blockquote>\n  <p>one target cannot be associated with more than 1500 terms for MF, BP, and CC subontologies combined. </p>\n</blockquote>\n<p>This 1500 terms/target limit is imposed during file parsing on Kaggle platform, before passing in the input into CAFA-evaluator</p>"
    },
    {
      "id": 3312634,
      "authorName": "common-people",
      "votes": 0,
      "postDate": "2025-11-07T15:35:17.803000",
      "content": "<p>Hi，sir, I saw that all the training data is in the test set, so, if it will be used in in the score of the public lb?</p>"
    },
    {
      "id": 3312640,
      "authorName": "An Phan",
      "votes": 1,
      "postDate": "2025-11-07T16:01:09.293000",
      "content": "<p>Proteins that appear in both training data and test superset could either be limited-knowledge or partial-knowledge. If proteins had 1 or 2 subontologies annotated (in training data), and gain the remaining subontology later, they are limited-knowledge. If proteins had all 3 subontologies annotated (in training data), and gain more new annotations in any subontology, they are partial-knowledge. </p>\n<blockquote>\n  <p>Note that in this competition, we also include the evaluation of proteins that already had experimental terms in all three subontologies, and have accumulated more experimental terms after the submission deadline, this is known as partial-knowledge protein targets. </p>\n</blockquote>\n<p>This is why all proteins in the training data are in the test superset, because they can still qualify as limited-knowledge or partial-knowledge proteins.</p>"
    },
    {
      "id": 3312889,
      "authorName": "common-people",
      "votes": 0,
      "postDate": "2025-11-08T06:56:07.447000",
      "content": "<p>And to this point, I also have a question, because of the training set is included in the test set, if the annotation in the public lb is more than it in the training set. Only in this way, we could mimic the private score using the public lb. And also the public score of cafa5 was covered and could not be used to see how the shake is</p>"
    },
    {
      "id": 3313944,
      "authorName": "An Phan",
      "votes": 0,
      "postDate": "2025-11-10T02:13:14.457000",
      "content": "<p>I'm not sure I understand what your question is. The public leaderboard currently has a small number of proteins that have hold-out experimental annotations that are not included in UniProt. If these proteins appeared in the training set, they were included in the public leaderboard as either limited-knowledge or partial-knowledge proteins. If these proteins didn't appear in the training set, they were included as no-knowledge proteins.</p>"
    },
    {
      "id": 3314622,
      "authorName": "sroger",
      "votes": 0,
      "postDate": "2025-11-10T09:01:20.653000",
      "content": "<p>Does the 500/1500 term limit apply to ancestral terms not explicitly listed in train_terms?</p>\n<p>Follow-up: If we choose to submit ancestral terms instead of letting the evaluation process infer them, does that reduce our term limit?</p>"
    },
    {
      "id": 3315258,
      "authorName": "An Phan",
      "votes": 0,
      "postDate": "2025-11-10T14:33:31.137000",
      "content": "<p>The term limit applies before any propagation occurs (in the evaluation process), so it does not affect ancestral terms if you don't submit ancestral terms.</p>\n<p>Yes, if you submit ancestral terms, those will count towards the term limit.</p>"
    },
    {
      "id": 3318631,
      "authorName": "common-people",
      "votes": 0,
      "postDate": "2025-11-11T09:14:36.923000",
      "content": "<p>Hi, I use the train_terms.tsv as the ground truth file, is this the reason that this script does not take the test set into account?</p>"
    },
    {
      "id": 3319374,
      "authorName": "An Phan",
      "votes": 1,
      "postDate": "2025-11-11T17:54:57.963000",
      "content": "<p>Ground truth file should be annotations that are made in the future, after an annotation accumulation period. If you want to mimic the evaluation process and use train_terms.tsv (collected June 2025) as ground truth file, let's say the annotation accumulation period is Feb 2025 until June 2025, then your model must predict GO terms for the proteins using only knowledge publicly available before Feb 2025 to prevent leakage.</p>"
    },
    {
      "id": 3322220,
      "authorName": "xaxipiruli",
      "votes": 0,
      "postDate": "2025-11-13T11:50:24.250000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ahphan\" target=\"_blank\">@ahphan</a>,</p>\n<p>regarding what you said about using <strong>cafaeval</strong>:</p>\n<pre><code>-max_terms 500 -prop fill -norm cafa -no_orphans\n</code></pre>\n<p>what about the options <code>-known</code> and <code>-toi</code> — will they also be used?</p>"
    },
    {
      "id": 3322926,
      "authorName": "An Phan",
      "votes": 1,
      "postDate": "2025-11-13T20:50:10.127000",
      "content": "<p><code>-toi</code> is not currently used for the public leaderboard. <code>-known</code> was used for partial-knowledge proteins, where we passed in the training GO terms of those proteins into the cafaevaluator to exclude these known GO terms from the evaluation process.</p>"
    }
  ],
  "index": {
    "id": "614668",
    "title": "Creating a time-shifted validation environment (CAFA5)",
    "authorName": "xaxipiruli",
    "commentCount": "16",
    "votes": "16",
    "postDate": "2025-11-05 14:49:56.603000"
  }
}