{
  "topic": {
    "id": 402648,
    "title": "ProteinBert - Protein Function language Model",
    "authorName": "Dan Ofer",
    "commentCount": 15,
    "votes": 45,
    "postDate": "2023-04-19T07:00:25.335000"
  },
  "comments": [
    {
      "id": 2239458,
      "authorName": "snowdrop613",
      "votes": 1,
      "postDate": "2023-04-29T14:08:29.590000",
      "content": "<p>Thanks for sending the link!</p>\n<p>Has anyone given BioBERT a go on Kaggle notebooks yet?</p>\n<p>I tried the <a href=\"https://github.com/nadavbra/protein_bert/blob/master/ProteinBERT%20demo.ipynb\" target=\"_blank\">protein-bert demo </a>after installing it with pip, but it threw an error when running the finetune func.</p>\n<blockquote>\n  <p>AttributeError: 'Adam' object has no attribute 'get_weights'</p>\n</blockquote>\n<p>I heard the error has <a href=\"https://github.com/nadavbra/protein_bert/pull/44\" target=\"_blank\">already been fixed</a>, but I can't seem to install it with pip.<br>\nI also attempted to clone it from GitHub, but it didn't work out.</p>\n<p>Do you have any ideas on what I should do next?</p>"
    },
    {
      "id": 2229759,
      "authorName": "Alexander Chervov",
      "votes": 1,
      "postDate": "2023-04-21T16:49:01.297000",
      "content": "<p>Cool ! <br>\nWould it be possible to generate the embeddings for the train/test here and share them with the community in the Kaggle datasets ? </p>\n<p>What is the average time for inference on sequences like we have here  ? </p>"
    },
    {
      "id": 2231515,
      "authorName": "Dan Ofer",
      "votes": 2,
      "postDate": "2023-04-23T11:42:55.750000",
      "content": "<p>It's very fast. I'd love to share such embeddings! If you have the capacity to run them and to share as a dataset, i'd be delighted to link to them from the <a href=\"https://github.com/nadavbra/protein_bert\" target=\"_blank\">protein_bert</a> repo as well!</p>"
    },
    {
      "id": 2231527,
      "authorName": "Alexander Chervov",
      "votes": 1,
      "postDate": "2023-04-23T11:57:02.310000",
      "content": "<p>Thanks for the comment ! <br>\nKaggle itself gives unlimited capacity to share the data - just create public dataset and put &lt;50G there. <br>\nIf it is more than 50G - just create several datasets. Notebook can access any number of datasets.<br>\nThe only technical remark is that huge Kaggle datasets - with more than say 5G on  - leads to quite long notebook initialization - when Kaggle loads the dataset to the notebook environment, so sometimes it is better to have several smaller datasets rather than one big.</p>"
    },
    {
      "id": 2229695,
      "authorName": "Andrew",
      "votes": 2,
      "postDate": "2023-04-21T15:56:39.923000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/danofer\" target=\"_blank\">@danofer</a>,</p>\n<p>Thanks for the link.  I wonder if you have comparative data on this vs other protein embeddings.  I've found <a href=\"https://github.com/Rostlab/goPredSim#performance-assessment\" target=\"_blank\">one comparison</a> that puts ProtT5 ahead of ProteinBert by ~2 percentage points for GO prediction across all of BP, MF, CC.  But that comparison was done by the author of ProtT5 so perhaps it uses ProteinBert in a sub-optimal way or has other problems that make it misleading.</p>\n<p>Is there anything you can point to that should lead me to question those findings or does it look right to you?  (It sounds like ProtT5 is much bigger than ProteinBert, which probably makes it harder to work with.  So perhaps the key advantage of ProteinBert is its small size?)</p>\n<p>Thanks,<br>\nAndrew</p>"
    },
    {
      "id": 2231387,
      "authorName": "Dan Ofer",
      "votes": 1,
      "postDate": "2023-04-23T09:01:29.510000",
      "content": "<p>They didn't compare ProteinBert there, they compared the author's own ProtBERT (which has a different architecture from us, and didn't do a GO pretraining task). </p>\n<p>PS - We were very stringent in filtering out annotations in the pretraining, but that's a seperate thing :)</p>"
    },
    {
      "id": 2236192,
      "authorName": "Andrew",
      "votes": 1,
      "postDate": "2023-04-26T15:52:38.763000",
      "content": "<p>Sorry, my mistake.  I had assumed ProtBert = ProteinBert.  Thanks for the clarification.</p>"
    },
    {
      "id": 2236882,
      "authorName": "Dan Ofer",
      "votes": 0,
      "postDate": "2023-04-27T07:00:12.283000",
      "content": "<p>Haha. It's a legit mistake. There's little in common between the two though, surprisingly ;)</p>"
    },
    {
      "id": 2227164,
      "authorName": "nhgiang",
      "votes": 1,
      "postDate": "2023-04-19T14:45:14.057000",
      "content": "<p>I'm also trying this approach and did some minimal work to prepare data for finetuning ProteinBERT. Kaggle-provided environment doesn't have enough resources, though, so you'll have to run it elsewhere.</p>\n<p>See my shared notebook in this competition (I've posted the link too many times for Kaggle's spam-limit)</p>"
    },
    {
      "id": 2229232,
      "authorName": "Dan Ofer",
      "votes": 1,
      "postDate": "2023-04-21T07:14:49.247000",
      "content": "<p>Kaggle's env should easily have enough resources, it's a 16M model! People ran it in other protein competitions. Where's the bottleneck for you?</p>"
    },
    {
      "id": 2398409,
      "authorName": "Levi Junyan Zhang",
      "votes": 0,
      "postDate": "2023-08-19T17:02:37.890000",
      "content": "<p>really cooooool</p>"
    },
    {
      "id": 2327818,
      "authorName": "",
      "votes": 0,
      "postDate": "2023-07-03T07:33:32.713000",
      "content": "<p><a href=\"https://www.kaggle.com/danofer\" target=\"_blank\">@danofer</a> , hello</p>\n<p>I looking for CAFA 3 4 and cannot find them..<br>\nCan you help me find link to CAFA 3 CAFA 4 ?</p>"
    },
    {
      "id": 2328015,
      "authorName": "Alexander Chervov",
      "votes": 1,
      "postDate": "2023-07-03T09:50:26.670000",
      "content": "<p><a href=\"https://www.kaggle.com/datasets/alexandervc/cafa-protein-function-annotation-challenges\" target=\"_blank\">https://www.kaggle.com/datasets/alexandervc/cafa-protein-function-annotation-challenges</a></p>"
    },
    {
      "id": 2307093,
      "authorName": "Maryam Mohammadi",
      "votes": 0,
      "postDate": "2023-06-17T19:50:36.550000",
      "content": "<p>oh really thanks for sharing the github link of this great project</p>"
    },
    {
      "id": 2259380,
      "authorName": "Abish Pius",
      "votes": 0,
      "postDate": "2023-05-15T00:21:05.937000",
      "content": "<p>Could I get some help with running the inference using the pre-trained model, not sure where to start to be hones.</p>"
    }
  ],
  "index": {
    "id": "402648",
    "title": "ProteinBert - Protein Function language Model",
    "authorName": "",
    "commentCount": "15",
    "votes": "45",
    "postDate": "2023-04-19 07:00:25.335000"
  },
  "competition": "cafa-5-protein-function-prediction"
}