{
  "topic": {
    "id": 466917,
    "title": "1st Place Solution for the CAFA5",
    "authorName": "王少钧",
    "commentCount": 7,
    "votes": 34,
    "postDate": "2024-01-10T11:55:37.361000"
  },
  "comments": [
    {
      "id": 2638950,
      "authorName": "hermit.mb",
      "votes": 0,
      "postDate": "2024-02-06T15:23:17.467000",
      "content": "<p>Is it possible to get your source code of the entire solution?</p>"
    },
    {
      "id": 2639149,
      "authorName": "Moori Varun",
      "votes": 0,
      "postDate": "2024-02-06T17:34:30.887000",
      "content": "<p>how can i download the dataset?</p>"
    },
    {
      "id": 2600363,
      "authorName": "Dan Ofer",
      "votes": 0,
      "postDate": "2024-01-13T15:11:06.523000",
      "content": "<p>Neat! (I'm presenting about Cafa5/your approach in our lab). </p>\n<p>Some questions:</p>\n<ol>\n<li>How did you get predictions for the first level of the ensemble, simple CV?.</li>\n<li>Blast-KNN = Simple PSI-blast? </li>\n<li>LR-Interpro = ? </li>\n<li>Why extract the features from interProScan and not the uniProt records?</li>\n</ol>"
    },
    {
      "id": 2601162,
      "authorName": "王少钧",
      "votes": 2,
      "postDate": "2024-01-14T07:29:14.857000",
      "content": "<ol>\n<li>We splited the dataset into training and validation sets, and used the training set to train component methods. Then we used the trained component methods to predict function for validation proteins. Finally, we trained the ensemble model based on the prediction on validation sets. For more details, you can find in \"GOLabeler: improving sequence-based large-scale protein function prediction by learning to rank\".</li>\n<li>Yep, we used psi-blast to find similar proteins for target proteins.</li>\n<li>LR-Interpro is also a component method derived from GOLabeler. It used InterProScan to extract protein families, domains, and motifs. Then we built a binary feature vector for each proteins based on the InterProScan results, where the '1' in j-th element indicates that the j-th feature belongs to target protein. Finally, we used these vectors to train logistic regression classifiers and make prediction.</li>\n<li>InterProScan is updated irregularly and collects more functional information. So, we believe that using InterProScan directly is more helpful.</li>\n</ol>"
    },
    {
      "id": 2632074,
      "authorName": "",
      "votes": 0,
      "postDate": "2024-02-02T06:01:07.263000",
      "content": "<p>Thanks for sharing your work. I am also working on a similar project where I have created three embeddings from three language models (LLMs) and structural features. My doubt is about how to create my output vector. How can I predict the GO terms for each protein? Can you elaborate on that part?</p>"
    },
    {
      "id": 2632303,
      "authorName": "",
      "votes": 0,
      "postDate": "2024-02-02T08:56:56.350000",
      "content": ""
    },
    {
      "id": 2632304,
      "authorName": "王少钧",
      "votes": 0,
      "postDate": "2024-02-02T08:57:51.817000",
      "content": "<ol>\n<li>For the first question, are you referring to how to obtain the representation of proteins from LLMs? Protein LLMs are multi-layered and can generate representations for each amino acid. Typically, we perform mean pooling on the output of the last layer for each amino acid, and generate the feature representation of the protein.</li>\n<li>To predict the GO terms for proteins, we collected large-scale functional annotation information, where each GO term is associated with numerous proteins. We treated these proteins as positive instances and considered the remaining proteins as negative instances. Then we trained a classifier for each GO term, aiming to predict the association scores between proteins and GO term.</li>\n</ol>"
    }
  ],
  "index": {
    "id": "466917",
    "title": "1st Place Solution for the CAFA5",
    "authorName": "",
    "commentCount": "7",
    "votes": "34",
    "postDate": "2024-01-10 13:44:04.713000"
  },
  "competition": "cafa-5-protein-function-prediction"
}