{
  "topic": {
    "id": 402696,
    "title": "Maybe we're not actually predicting protein function? ",
    "authorName": "Darek Kłeczek",
    "commentCount": 14,
    "votes": 30,
    "postDate": "2023-04-19T12:55:30.335000"
  },
  "comments": [
    {
      "id": 2227168,
      "authorName": "Iddo Friedberg",
      "votes": 12,
      "postDate": "2023-04-19T14:46:23.677000",
      "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> You raise an interesting point, also known as the \"incomplete knowledge problem\" or \"the open world problem\". It is true that we only know what a protein is doing at any given point in time, and that in the future we may know more. The problem as applied to CAFA was addressed in <a href=\"https://www.cell.com/trends/genetics/fulltext/S0168-9525(13)00166-2\" target=\"_blank\">this paper</a>. However, we found that the effect on CAFA evaluations, within a range of several years, <a href=\"https://academic.oup.com/bioinformatics/article/30/17/i609/201287\" target=\"_blank\">is not large</a>. We have been running CAFA since 2010, and we revisit top methods from time to time using newer corpora of proteins, to see the difference in performance over time.  However, it is a valid question, and it applies to other problems in which knowledge is perpetually incomplete, and knowledge is acquired over time. </p>"
    },
    {
      "id": 2227447,
      "authorName": "Clara De Paolis",
      "votes": 8,
      "postDate": "2023-04-19T19:02:19.677000",
      "content": "<p>Exactly. Specifically, we are interested in predicting function for exactly those proteins for which we <em>do not currently know the function</em>. The only ones of those we can evaluate against are the ones what acquire functional annotations in the future.  To determine a winner, we need to cap \"the future\" to some date (Dec 21 in this case) and evaluate with the knowledge at that time. </p>\n<p>Lastly I will point out that if a protein is not annotated with some function, this is not a negative label but an absence of label. This is known as <a href=\"https://en.wikipedia.org/wiki/One-class_classification\" target=\"_blank\">Positive-Unlabeled learning</a>, a challenging and actually ubiquitous setting in many domains. Evaluation and prediction in this domain is challenging and fascinating. Anyone interested is encouraged to check out <a href=\"https://www.ccs.neu.edu/home/radivojac/publications.html\" target=\"_blank\">papers from our lab</a> and others to learn more.</p>"
    },
    {
      "id": 2228001,
      "authorName": "Dan Ofer",
      "votes": 1,
      "postDate": "2023-04-20T07:36:29.507000",
      "content": "<p>I do love me PU learning :D <br>\n(Also done in our NeuroPID paper)</p>"
    },
    {
      "id": 2227999,
      "authorName": "Dan Ofer",
      "votes": 1,
      "postDate": "2023-04-20T07:34:40.863000",
      "content": "<p>That is correct. That said, in previous ways, we didn't find a way to optimize for this, the amount of annotations added are few, scarce and depend on the labs. <br>\n(And date diffing uniprot annotations is a bitch)</p>"
    },
    {
      "id": 2227656,
      "authorName": "Tilii",
      "votes": 2,
      "postDate": "2023-04-19T22:56:00.163000",
      "content": "<p>You bring up a good point, and hopefully it is clear from the hosts' answers how this will be addressed during evaluation.</p>\n<p>Part of the difficulty in correctly annotating protein functions is that not all proteins have singular functions, as many of them \"moonlight\" by doing something else on the side. Kind of like people who are lawyers during the week but strippers during the weekend. I kid …</p>\n<p>Here is an example that may be relatable. If a brand new protein was required to neutralize every single medication or toxic chemical we get exposed to during our lifetimes, we'd quickly run out of options. As a species we haven't been exposed to a bunch of novel drugs and chemicals for any significant period of time, yet our livers still may have a way of eliminating them. That is because many of our detoxifying enzymes have broad specificities, and can inactivate many chemicals as long as they are at least somewhat similar to each other. They are not equally efficient in dealing with all problematic compounds, but even a poor enzyme activity on a given toxin is better than no activity at all. It can be argued that such enzymes have many functions, yet we are unlikely to know all of them for the foreseeable future as it is impossible to test. Similarly, some proteins work preferentially with RNA, yet they may be able to do similar things with DNA given general similarity between nucleic acids.</p>"
    },
    {
      "id": 2227441,
      "authorName": "serangu",
      "votes": 2,
      "postDate": "2023-04-19T18:56:13.933000",
      "content": "<p>GO aims to represent the current state of knowledge in biology, hence it is constantly revised and expanded as biological knowledge accumulates. <strong>Changes are made on a weekly basis</strong> (most relatively minor).</p>\n<p>I see this <a href=\"http://geneontology.org/docs/ontology-documentation/\" target=\"_blank\">here</a></p>"
    },
    {
      "id": 2227042,
      "authorName": "Marília Prata",
      "votes": 2,
      "postDate": "2023-04-19T13:21:42.817000",
      "content": "<p>That was A. Chervov comment on my topic. I hope he come here to provide his skilled, professional point-of-view.</p>\n<p>\"It is a regular challenge in bioinformatics and huge literature exists:\"  Chervov tip.</p>\n<p><a href=\"https://scholar.google.fr/scholar?start=0&amp;q=cafa+protein+function+prediction&amp;hl=en&amp;as_sdt=0,5&amp;as_vis=1\" target=\"_blank\">https://scholar.google.fr/scholar?start=0&amp;q=cafa+protein+function+prediction&amp;hl=en&amp;as_sdt=0,5&amp;as_vis=1</a></p>\n<p>And the ChatGPT answer (for me it doesn't say much): </p>\n<p>\"It's worth noting that the best approach for a specific problem in the CAFA challenge can depend on the specific features of the data being analyzed. Therefore, it's common to use a combination of multiple methods and to continuously refine them based on the results obtained in the challenge. \"</p>"
    },
    {
      "id": 2227417,
      "authorName": "serangu",
      "votes": 2,
      "postDate": "2023-04-19T18:43:22.390000",
      "content": "<p><a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> But how do you know that you're moving in the right direction if the criteria for evaluating a model change as you go along?</p>\n<p>I understand correctly that if we know only one thing about some protein right now. Then the evaluation of the model will be made precisely from the current knowledge.</p>\n<p>If the contest organizers learn something new about this protein in a couple of months, then the model's score will change?</p>"
    },
    {
      "id": 2227512,
      "authorName": "Marília Prata",
      "votes": 0,
      "postDate": "2023-04-19T20:06:54.507000",
      "content": "<p>Impermanence is the essence of life Serangu.<br>\nWhenever we find out/learn something new, things change. Mostly in Health field.</p>"
    },
    {
      "id": 2228476,
      "authorName": "Iddo Friedberg",
      "votes": 2,
      "postDate": "2023-04-20T15:17:57.007000",
      "content": "<p>From the <a href=\"https://www.kaggle.com/competitions/cafa-5-protein-function-prediction/data\" target=\"_blank\">Data</a> tab:</p>\n<p>'''<br>\nThe participants are cautioned that the leaderboard was designed to display method performance on a relatively small selection of proteins from the test superset (see Data), provided to us by the UniProtKB team but not available in UniProtKB or other public databases. These proteins will not be included in the test set for the subontologies used for the leaderboard evaluation. The <strong>final test set will</strong> consist of proteins that will have accumulated functional terms <strong>after</strong> the submission deadline and therefore, some distribution shift between the sample of proteins used for the leaderboard and the final evaluation sample is to be expected. Overall, the participants are encouraged to maximize the generalization performance and use the leaderboard only as a rough indicator of their model's performance.<br>\n'''</p>\n<p>We are not changing the Gene Ontology we are using for assessment. We will always use the  version  from 2023-01 in this challenge (the .obo file)</p>\n<p>We will be using a different (expanded) test set over the leaderboard test set. Does that make sense? </p>"
    },
    {
      "id": 2307089,
      "authorName": "Maryam Mohammadi",
      "votes": 0,
      "postDate": "2023-06-17T19:48:33.437000",
      "content": "<p>It was a challenging question, of course, nothing is impossible in the world of artificial intelligence, and I think we can achieve this… for example by using NLP or CNN or ensemble methods …the procedure will be very complex but it can detect more aspect of the problem</p>"
    },
    {
      "id": 2227414,
      "authorName": "serangu",
      "votes": 0,
      "postDate": "2023-04-19T18:39:08.273000",
      "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> I am not an expert. But I agree with you. That the criteria for assessing the correctness of a prediction are not just blurred but will change several times. This is explicitly stated in the rules.</p>\n<p>Is this field of science evolving so fast that there is no established dataset for proteins?</p>"
    },
    {
      "id": 2229608,
      "authorName": "Andrew",
      "votes": 2,
      "postDate": "2023-04-21T14:12:44.110000",
      "content": "<blockquote>\n  <p>Is this field of science evolving so fast that there is no established dataset for proteins?</p>\n</blockquote>\n<p>I think the \"problem\" (from a Kaggle competition point of view) is that this research tends to be done in public. Therefore, there's no easy way to hold back a test set for the leaderboard(s) - we could just look up the right answers by consulting the public databases.</p>"
    },
    {
      "id": 2229923,
      "authorName": "Iddo Friedberg",
      "votes": 4,
      "postDate": "2023-04-21T20:19:22.657000",
      "content": "<blockquote>\n  <p>Therefore, there's no easy way to hold back a test set for the leaderboard(s) - we could just look up the right answers by consulting the public databases.</p>\n</blockquote>\n<p>Exactly. Although we do have some friendly people in Uniprot that are holding back on annotations for this competition.  And, like March Madness competitions, we have the future giving us more data. </p>"
    }
  ],
  "index": {
    "id": "402696",
    "title": "Maybe we're not actually predicting protein function? ",
    "authorName": "",
    "commentCount": "14",
    "votes": "30",
    "postDate": "2023-04-19 12:55:30.335000"
  },
  "competition": "cafa-5-protein-function-prediction"
}