{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"colab":{"provenance":[],"gpuType":"T4"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":41875,"databundleVersionId":5521661,"sourceType":"competition"},{"sourceId":5499219,"sourceType":"datasetVersion","datasetId":3167603}],"dockerImageVersionId":30674,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# Protein-Sequences: train_sequences.fasta\n# train_terms.tsv: Ground truth for the proteins in train_sequences.fasta.\n    # (1)unique protein id\n    # (2)GO Term id\n    # (3)In which ontology the term appears\n\n# Predict all terms/functions/GO Term IDs of protein sequences.\n# One Protein Sequence can have many functions/terms\n# This means that the task at hand is a multi-label classification problem.\n\n\n# To train a machine learning model we cannot use the alphabetical protein sequences intrain_sequences.fasta directly.\n# First Convert them into a vector format.\n# In this notebook, we will use embeddings of the protein sequences to train the model.\n# You can think of protein embeddings to be similar to word embeddings used to train NLP models.\n\n\n# Protein embeddings Captures protein's structural and functional characteristics.\n# Either train custom ML model to learn the protein embeddings of the protein sequences.\n# Dataset represents proteins using amino-acid sequences which is a standard approach, Use publicly available pre-trained protein embedding models to generate the embeddings.\n\n# used the precalculated protein embeddings created by Sergei Fironov using the Rost Lab's T5 protein language model.\n\n# The protein embeddings to be used for training are recorded in train_embeds.npy and the corresponding protein ids are available in train_ids.npy.","metadata":{"execution":{"iopub.status.busy":"2024-04-05T16:11:19.123219Z","iopub.execute_input":"2024-04-05T16:11:19.123923Z","iopub.status.idle":"2024-04-05T16:11:19.129453Z","shell.execute_reply.started":"2024-04-05T16:11:19.123892Z","shell.execute_reply":"2024-04-05T16:11:19.12842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# CAFA 5 protein function Prediction with TensorFlow\n\nThis notebook walks you through how to train a DNN model using TensorFlow on the CAFA 5 protein function Prediction dataset made available for this competition.\n\nThe objective of the model is to predict the function(aka **GO term ID**) of a set of proteins based on their amino acid sequences and other data.\n\n\n**Note** : This notebook runs without any GPU. This is because enabling GPUs leaves less RAM memory on the VM and the submission step needs a lot of memory. One point where this would impact is when training the model. With CPU it will take around 2 minutes while on GPU it would take around 30 seconds.","metadata":{}},{"cell_type":"markdown","source":"## About the Data\n\n### Protein Sequence\n\nEach protein is composed of dozens or hundreds of amino acids that are linked sequentially. Each amino acid in the sequence may be represented by a one-letter or three-letter code. Thus the sequence of a protein is often notated as a string of letters.\n\n<img src=\"https://cityu-bioinformatics.netlify.app/img/tools/protein/pro_seq.png\" alt =\"Sequence.png\" style='width: 800px;' >\n\nImage source - [https://cityu-bioinformatics.netlify.app/](https://cityu-bioinformatics.netlify.app/too2/new_proteo/pro_seq/)\n\nThe `train_sequences.fasta` made available for this competitions, contains the sequences for proteins with annotations (labelled proteins).","metadata":{}},{"cell_type":"markdown","source":"# Gene Ontology\n\nWe can define the functional properties of a proteins using Gene Ontology(GO). Gene Ontology (GO) describes our understanding of the biological domain with respect to three aspects:\n1. Molecular Function (MF)\n2. Biological Process (BP)\n3. Cellular Component (CC)\n\nRead more about Gene Ontology [here](http://geneontology.org/docs/ontology-documentation).\n\nFile `train_terms.tsv` contains the list of annotated terms (ground truth) for the proteins in `train_sequences.fasta`. In `train_terms.tsv` the first column indicates the protein's UniProt accession ID (unique protein id), the second is the `GO Term ID`, and the third indicates in which ontology the term appears.","metadata":{}},{"cell_type":"markdown","source":"# Labels of the dataset\n\nThe objective of our model is to predict the terms (functions) of a protein sequence. One protein sequence can have many functions and can thus be classified into any number of terms. Each term is uniquely identified by a `GO Term ID`. Thus our model has to predict all the `GO Term ID`s for a protein sequence. This means that the task at hand is a multi-label classification problem.","metadata":{}},{"cell_type":"markdown","source":"# Protein embeddings for train and test data\n\nTo train a machine learning model we cannot use the alphabetical protein sequences in`train_sequences.fasta` directly. They have to be converted into a vector format. In this notebook, we will use embeddings of the protein sequences to train the model. You can think of protein embeddings to be similar to word embeddings used to train NLP models.\n<!-- Instead, to make calculations and data preparation easier we will use precalculated protein embeddings.\n -->\nProtein embeddings are a machine-friendly method of capturing the protein's structural and functional characteristics, mainly through its sequence. One approach is to train a custom ML model to learn the protein embeddings of the protein sequences in the dataset being used in this notebook. Since this dataset represents proteins using amino-acid sequences which is a standard approach, we can use any publicly available pre-trained protein embedding models to generate the embeddings.\n\nThere are a variety of protein embedding models. To make data preparation easier, we have used the precalculated protein embeddings created by [Sergei Fironov](https://www.kaggle.com/sergeifironov) using the Rost Lab's T5 protein language model in this notebook. The precalculated protein embeddings can be found [here](https://www.kaggle.com/datasets/sergeifironov/t5embeds). We have added this dataset to the notebook along with the dataset made available for the competition.\n\nTo add this to your enviroment, on the right side panel, click on `Add Data` and search for `t5embeds` (make sure that it's the correct [one](https://www.kaggle.com/datasets/sergeifironov/t5embeds)) and then click on the `+` beside it.\n\n","metadata":{}},{"cell_type":"markdown","source":"# Import the Required Libraries","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport progressbar","metadata":{"execution":{"iopub.status.busy":"2024-04-05T16:11:26.354102Z","iopub.execute_input":"2024-04-05T16:11:26.354905Z","iopub.status.idle":"2024-04-05T16:11:49.016228Z","shell.execute_reply.started":"2024-04-05T16:11:26.354874Z","shell.execute_reply":"2024-04-05T16:11:49.015294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# from google.colab import drive\n# drive.mount('/content/drive')","metadata":{"execution":{"iopub.status.busy":"2024-04-05T16:11:57.317932Z","iopub.execute_input":"2024-04-05T16:11:57.318543Z","iopub.status.idle":"2024-04-05T16:11:57.322608Z","shell.execute_reply.started":"2024-04-05T16:11:57.318511Z","shell.execute_reply":"2024-04-05T16:11:57.321642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_terms=pd.read_csv('/kaggle/input/cafa-5-protein-function-prediction/Train/train_terms.tsv', sep=\"\\t\")\ntrain_terms.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-05T16:41:03.916722Z","iopub.execute_input":"2024-04-05T16:41:03.917583Z","iopub.status.idle":"2024-04-05T16:41:06.730472Z","shell.execute_reply.started":"2024-04-05T16:41:03.91755Z","shell.execute_reply":"2024-04-05T16:41:06.729472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_terms.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-05T16:41:08.330448Z","iopub.execute_input":"2024-04-05T16:41:08.331326Z","iopub.status.idle":"2024-04-05T16:41:08.337094Z","shell.execute_reply.started":"2024-04-05T16:41:08.331287Z","shell.execute_reply":"2024-04-05T16:41:08.336132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loading the protein embeddings\n\n\nWe will now load the pre calculated protein embeddings created by [Sergei Fironov](https://www.kaggle.com/sergeifironov) using the Rost Lab's T5 protein language model.\n\nIf the `tfembeds` is not yet on the input data of the notebook, you can add it to your enviromentby clicking on `Add Data` and search for `t5embeds` (make sure that it's the correct [one](https://www.kaggle.com/datasets/sergeifironov/t5embeds) ) and then click on the `+` beside it.\n\nThe protein embeddings to be used for training are recorded in `train_embeds.npy` and the corresponding protein ids are available in `train_ids.npy`.","metadata":{}},{"cell_type":"markdown","source":"First, we will load the protein ids of the protein embeddings in the train dataset contained in `train_ids.npy` into a numpy array.","metadata":{}},{"cell_type":"code","source":"train_protein_ids = np.load('/kaggle/input/t5embeds/train_ids.npy')\ntrain_protein_ids_pd=pd.DataFrame(train_protein_ids)\ntrain_protein_ids_pd.head(100)","metadata":{"execution":{"iopub.status.busy":"2024-04-05T16:12:42.243466Z","iopub.execute_input":"2024-04-05T16:12:42.243829Z","iopub.status.idle":"2024-04-05T16:12:42.343702Z","shell.execute_reply.started":"2024-04-05T16:12:42.243802Z","shell.execute_reply":"2024-04-05T16:12:42.342734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<!-- Now, we will load`train_embeds.py` which contains the pre-calculated embeddings of the proteins in the train dataset. with protein_ids (`id`s we loaded previously from the **train_ids.npy**) into a numpy array. This array now contains the precalculated embeddings for the protein_ids( Ids we loaded above from **train_ids.npy**) needed for training. -->\n\nAfter loading the files as numpy arrays, we will convert them into Pandas dataframe.\n\nEach protein embedding is a vector of length 1024. We create the resulting dataframe such that there are 1024 columns to represent the values in each of the 1024 places in the vector.","metadata":{}},{"cell_type":"code","source":"train_embeddings = np.load('/kaggle/input/t5embeds/train_embeds.npy')\ncolumn_num = train_embeddings.shape[1]\ntrain_df  = pd.DataFrame(train_embeddings, columns = [\"Column_\" + str(i) for i in range(1, column_num+1)])\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-05T16:12:48.410969Z","iopub.execute_input":"2024-04-05T16:12:48.411443Z","iopub.status.idle":"2024-04-05T16:13:01.377428Z","shell.execute_reply.started":"2024-04-05T16:12:48.411408Z","shell.execute_reply":"2024-04-05T16:13:01.376519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prepare the dataset\n\nReference: https://www.kaggle.com/code/alexandervc/baseline-multilabel-to-multitarget-binary","metadata":{}},{"cell_type":"markdown","source":"First we will extract all the needed labels(`GO term ID`) from `train_terms.tsv` file. There are more than 40,000 labels. In order to simplify our model, we will choose the most frequent 1500 `GO term ID`s as labels.","metadata":{}},{"cell_type":"markdown","source":"Let's plot the most frequent 100 `GO Term ID`s in `train_terms.tsv`.","metadata":{}},{"cell_type":"code","source":"# Extracting Go IDs\nnum_of_labels = 1500\nlabels = train_terms['term'].value_counts().index\nlabels=labels[:num_of_labels]\nlabels","metadata":{"execution":{"iopub.status.busy":"2024-04-05T16:13:20.631151Z","iopub.execute_input":"2024-04-05T16:13:20.632479Z","iopub.status.idle":"2024-04-05T16:13:21.663166Z","shell.execute_reply.started":"2024-04-05T16:13:20.632434Z","shell.execute_reply":"2024-04-05T16:13:21.662371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, we will create a new dataframe by filtering the train terms with the selected `GO Term ID`s.","metadata":{}},{"cell_type":"code","source":"# Extracting the dataset which contains the top 1500(labels) GO terms\ntrain_terms_updated = train_terms.loc[train_terms['term'].isin(labels)]\ntrain_terms_updated=train_terms_updated.reset_index(drop=True)\ntrain_terms_updated","metadata":{"execution":{"iopub.status.busy":"2024-04-05T16:13:27.309875Z","iopub.execute_input":"2024-04-05T16:13:27.310225Z","iopub.status.idle":"2024-04-05T16:13:28.135816Z","shell.execute_reply.started":"2024-04-05T16:13:27.310199Z","shell.execute_reply":"2024-04-05T16:13:28.135014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pie_df = train_terms_updated['aspect'].value_counts()\npalette_color = sns.color_palette('bright')\nplt.pie(pie_df.values, labels=np.array(pie_df.index), colors=palette_color, autopct='%.0f%%')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-05T16:13:30.821474Z","iopub.execute_input":"2024-04-05T16:13:30.821828Z","iopub.status.idle":"2024-04-05T16:13:31.785323Z","shell.execute_reply.started":"2024-04-05T16:13:30.821802Z","shell.execute_reply":"2024-04-05T16:13:31.784019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see, majority of the `GO term Id`s have BPO(Biological Process Ontology) as their aspect. Since this is a multi label classification problem, in the labels array we will denote the presence or absence of each Go Term Id for a protein id using a 1 or 0. First, we will create a numpy array train_labels of required size for the labels. To update the train_labels array with the appropriate values, we will loop through the label list.","metadata":{}},{"cell_type":"markdown","source":"The final labels dataframe (`label_df`) is composed of 1500 columns and 142246 entries. We can see all 1500 dimensions(results will be truncated since the number of columns is big) of our dataset by printing out the first 5 entries using the following code:","metadata":{}},{"cell_type":"code","source":"train_size = train_protein_ids.shape[0]\ntrain_labels = np.zeros((train_size ,num_of_labels))\nseries_train_protein_ids = pd.Series(train_protein_ids)\n\nbar = progressbar.ProgressBar(maxval=num_of_labels, widgets=[progressbar.Bar('=', '[', ']'), ' ', progressbar.Percentage()])\nfor i in range(num_of_labels):\n    n_train_terms = train_terms_updated[train_terms_updated['term'] ==  labels[i]]\n    label_related_proteins = n_train_terms['EntryID'].unique()\n    train_labels[:,i] =  series_train_protein_ids.isin(label_related_proteins).astype(float)\n    bar.update(i+1)\nbar.finish()\n\nlabels_df = pd.DataFrame(data = train_labels, columns = labels)\nprint(labels_df.shape)","metadata":{"execution":{"iopub.status.busy":"2024-04-05T16:13:38.233927Z","iopub.execute_input":"2024-04-05T16:13:38.234317Z","iopub.status.idle":"2024-04-05T16:33:01.270232Z","shell.execute_reply.started":"2024-04-05T16:13:38.234285Z","shell.execute_reply":"2024-04-05T16:33:01.26929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_size = train_protein_ids.shape[0]\ntrain_labels = np.zeros((train_size ,num_of_labels))\ntrain_labels.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-05T16:58:22.967647Z","iopub.execute_input":"2024-04-05T16:58:22.968685Z","iopub.status.idle":"2024-04-05T16:58:22.975982Z","shell.execute_reply.started":"2024-04-05T16:58:22.968636Z","shell.execute_reply":"2024-04-05T16:58:22.975039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"series_train_protein_ids = pd.Series(train_protein_ids)\nseries_train_protein_ids","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:01:57.504412Z","iopub.execute_input":"2024-04-05T17:01:57.505188Z","iopub.status.idle":"2024-04-05T17:01:57.530222Z","shell.execute_reply.started":"2024-04-05T17:01:57.505157Z","shell.execute_reply":"2024-04-05T17:01:57.529257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bar = progressbar.ProgressBar(maxval=num_of_labels, widgets=[progressbar.Bar('=', '[', ']'), ' ', progressbar.Percentage()])\nfor i in range(num_of_labels):\n    n_train_terms = train_terms_updated[train_terms_updated['term'] ==  labels[i]]\n    label_related_proteins = n_train_terms['EntryID'].unique()\n    train_labels[:,i] =  series_train_protein_ids.isin(label_related_proteins).astype(float)\n    bar.update(i+1)\nbar.finish()","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:02:28.013472Z","iopub.execute_input":"2024-04-05T17:02:28.013828Z","iopub.status.idle":"2024-04-05T17:22:31.147303Z","shell.execute_reply.started":"2024-04-05T17:02:28.013802Z","shell.execute_reply":"2024-04-05T17:22:31.146313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:22:58.613028Z","iopub.execute_input":"2024-04-05T17:22:58.613412Z","iopub.status.idle":"2024-04-05T17:22:58.619982Z","shell.execute_reply.started":"2024-04-05T17:22:58.613382Z","shell.execute_reply":"2024-04-05T17:22:58.619088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_terms_updated[train_terms_updated['term'] ==  labels[i]]","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:23:15.281873Z","iopub.execute_input":"2024-04-05T17:23:15.282485Z","iopub.status.idle":"2024-04-05T17:23:16.078384Z","shell.execute_reply.started":"2024-04-05T17:23:15.282453Z","shell.execute_reply":"2024-04-05T17:23:16.07705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"label_related_proteins","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:23:26.871983Z","iopub.execute_input":"2024-04-05T17:23:26.872517Z","iopub.status.idle":"2024-04-05T17:23:26.880455Z","shell.execute_reply.started":"2024-04-05T17:23:26.87247Z","shell.execute_reply":"2024-04-05T17:23:26.87946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"series_train_protein_ids","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:23:40.819752Z","iopub.execute_input":"2024-04-05T17:23:40.820114Z","iopub.status.idle":"2024-04-05T17:23:40.828252Z","shell.execute_reply.started":"2024-04-05T17:23:40.820086Z","shell.execute_reply":"2024-04-05T17:23:40.827297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"series_train_protein_ids.isin(label_related_proteins)","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:23:44.458557Z","iopub.execute_input":"2024-04-05T17:23:44.459384Z","iopub.status.idle":"2024-04-05T17:23:44.477549Z","shell.execute_reply.started":"2024-04-05T17:23:44.459355Z","shell.execute_reply":"2024-04-05T17:23:44.476614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%mkdir /kaggle/working/final_dataset","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:24:24.345979Z","iopub.execute_input":"2024-04-05T17:24:24.346459Z","iopub.status.idle":"2024-04-05T17:24:25.383341Z","shell.execute_reply.started":"2024-04-05T17:24:24.34643Z","shell.execute_reply":"2024-04-05T17:24:25.382089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels_df = pd.DataFrame(data = train_labels, columns = labels)\nlabels_df.to_csv('/kaggle/working/final_dataset/data.csv')","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:24:35.273908Z","iopub.execute_input":"2024-04-05T17:24:35.274311Z","iopub.status.idle":"2024-04-05T17:27:12.582481Z","shell.execute_reply.started":"2024-04-05T17:24:35.274275Z","shell.execute_reply":"2024-04-05T17:27:12.581429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels_df=pd.read_csv('/kaggle/working/final_dataset/data.csv')\nlabels_df=labels_df.drop(columns='Unnamed: 0')","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:28:47.359034Z","iopub.execute_input":"2024-04-05T17:28:47.359454Z","iopub.status.idle":"2024-04-05T17:29:18.162865Z","shell.execute_reply.started":"2024-04-05T17:28:47.359423Z","shell.execute_reply":"2024-04-05T17:29:18.161996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels_df.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:29:18.164637Z","iopub.execute_input":"2024-04-05T17:29:18.164962Z","iopub.status.idle":"2024-04-05T17:29:18.17228Z","shell.execute_reply.started":"2024-04-05T17:29:18.164935Z","shell.execute_reply":"2024-04-05T17:29:18.171313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels_df.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:29:18.173575Z","iopub.execute_input":"2024-04-05T17:29:18.173934Z","iopub.status.idle":"2024-04-05T17:29:18.21498Z","shell.execute_reply.started":"2024-04-05T17:29:18.1739Z","shell.execute_reply":"2024-04-05T17:29:18.214094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training\n\nNext, we will use Tensorflow to train a Deep Neural Network with the protein embeddings.","metadata":{}},{"cell_type":"code","source":"train_df.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:30:20.384379Z","iopub.execute_input":"2024-04-05T17:30:20.384753Z","iopub.status.idle":"2024-04-05T17:30:20.391375Z","shell.execute_reply.started":"2024-04-05T17:30:20.384728Z","shell.execute_reply":"2024-04-05T17:30:20.390286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels_df.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:30:25.197885Z","iopub.execute_input":"2024-04-05T17:30:25.198258Z","iopub.status.idle":"2024-04-05T17:30:25.204457Z","shell.execute_reply.started":"2024-04-05T17:30:25.19822Z","shell.execute_reply":"2024-04-05T17:30:25.203439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"INPUT_SHAPE = [train_df.shape[1]]\nBATCH_SIZE = 5120\n\nmodel = tf.keras.Sequential([\n    tf.keras.layers.BatchNormalization(input_shape=INPUT_SHAPE),\n    tf.keras.layers.Dense(units=512, activation='relu'),\n    tf.keras.layers.Dense(units=512, activation='relu'),\n    tf.keras.layers.Dense(units=512, activation='relu'),\n    tf.keras.layers.Dense(units=num_of_labels,activation='sigmoid')\n])\n\nmodel.compile(\n    optimizer=tf.keras.optimizers.Adam(learning_rate=0.001),\n    loss='binary_crossentropy',\n    metrics=['binary_accuracy', tf.keras.metrics.AUC()],\n)\n\nhistory = model.fit(\n    train_df,\n    labels_df,\n    batch_size=BATCH_SIZE,\n    epochs=10\n)","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:30:31.847518Z","iopub.execute_input":"2024-04-05T17:30:31.848631Z","iopub.status.idle":"2024-04-05T17:30:57.119842Z","shell.execute_reply.started":"2024-04-05T17:30:31.848599Z","shell.execute_reply":"2024-04-05T17:30:57.118566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!mkdir model\nmodel.save(\"/kaggle/working/model/model_1.h5\")","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:31:03.216908Z","iopub.execute_input":"2024-04-05T17:31:03.223049Z","iopub.status.idle":"2024-04-05T17:31:04.467029Z","shell.execute_reply.started":"2024-04-05T17:31:03.223017Z","shell.execute_reply":"2024-04-05T17:31:04.465746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plot the model's loss and accuracy for each epoch","metadata":{}},{"cell_type":"code","source":"history_df = pd.DataFrame(history.history)\nhistory_df.loc[:, ['loss']].plot(title=\"Cross-entropy\")\nhistory_df.loc[:, ['binary_accuracy']].plot(title=\"Accuracy\")","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:31:26.436207Z","iopub.execute_input":"2024-04-05T17:31:26.43693Z","iopub.status.idle":"2024-04-05T17:31:27.149689Z","shell.execute_reply.started":"2024-04-05T17:31:26.436889Z","shell.execute_reply":"2024-04-05T17:31:27.148672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"markdown","source":"For submission we will use the protein embeddings of the test data created by \n\n*   List item\n*   List item\n\n[Sergei Fironov](https://www.kaggle.com/sergeifironov) using the Rost Lab's T5 protein language model.","metadata":{}},{"cell_type":"code","source":"test_embeddings = np.load('/kaggle/input/t5embeds/test_embeds.npy')\ncolumn_num = test_embeddings.shape[1]\ntest_df = pd.DataFrame(test_embeddings, columns = [\"Column_\" + str(i) for i in range(1, column_num+1)])\ntest_df","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:31:43.096119Z","iopub.execute_input":"2024-04-05T17:31:43.096835Z","iopub.status.idle":"2024-04-05T17:31:55.542911Z","shell.execute_reply.started":"2024-04-05T17:31:43.096799Z","shell.execute_reply":"2024-04-05T17:31:55.541877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.models import load_model\nmodel=load_model('/kaggle/working/model/model_1.h5')\npredictions =  model.predict(test_df)","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:32:02.681554Z","iopub.execute_input":"2024-04-05T17:32:02.681908Z","iopub.status.idle":"2024-04-05T17:32:14.317456Z","shell.execute_reply.started":"2024-04-05T17:32:02.68188Z","shell.execute_reply":"2024-04-05T17:32:14.316229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pickle\nfile_path = '/kaggle/working/predictions'\nwith open(file_path, 'wb') as file:\n    pickle.dump(predictions, file)","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:32:14.3194Z","iopub.execute_input":"2024-04-05T17:32:14.319691Z","iopub.status.idle":"2024-04-05T17:32:15.711757Z","shell.execute_reply.started":"2024-04-05T17:32:14.319667Z","shell.execute_reply":"2024-04-05T17:32:15.710924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport progressbar","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:34:38.131047Z","iopub.execute_input":"2024-04-05T17:34:38.131493Z","iopub.status.idle":"2024-04-05T17:34:38.136684Z","shell.execute_reply.started":"2024-04-05T17:34:38.131461Z","shell.execute_reply":"2024-04-05T17:34:38.135718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# from google.colab import drive\n# drive.mount('/content/drive')","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:34:41.372876Z","iopub.execute_input":"2024-04-05T17:34:41.373272Z","iopub.status.idle":"2024-04-05T17:34:41.377721Z","shell.execute_reply.started":"2024-04-05T17:34:41.373229Z","shell.execute_reply":"2024-04-05T17:34:41.376758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_terms=pd.read_csv('/kaggle/input/cafa-5-protein-function-prediction/Train/train_terms.tsv', sep=\"\\t\")\nnum_of_labels = 1500\nlabels = train_terms['term'].value_counts().index\nlabels=labels[:num_of_labels]\nlabels","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:38:16.018186Z","iopub.execute_input":"2024-04-05T17:38:16.018595Z","iopub.status.idle":"2024-04-05T17:38:21.115373Z","shell.execute_reply.started":"2024-04-05T17:38:16.018562Z","shell.execute_reply":"2024-04-05T17:38:21.114383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pickle\nfile_path = '/kaggle/working/predictions'\nwith open(file_path, 'rb') as file:\n    predictions = pickle.load(file)","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:36:17.390073Z","iopub.execute_input":"2024-04-05T17:36:17.390796Z","iopub.status.idle":"2024-04-05T17:36:20.984184Z","shell.execute_reply.started":"2024-04-05T17:36:17.390764Z","shell.execute_reply":"2024-04-05T17:36:20.983284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions.shape","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:36:20.985739Z","iopub.execute_input":"2024-04-05T17:36:20.986074Z","iopub.status.idle":"2024-04-05T17:36:20.993646Z","shell.execute_reply.started":"2024-04-05T17:36:20.986045Z","shell.execute_reply":"2024-04-05T17:36:20.992476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\ntest_protein_ids = np.load('/kaggle/input/t5embeds/test_ids.npy')\nnum_predictions = predictions.shape[1]\nnum_labels = predictions.shape[0]\n\ndata = {\n    'Protein Id': pd.Series(np.repeat(test_protein_ids, num_predictions)),\n    'GO Term Id': pd.Series(np.tile(labels, num_predictions)),\n    'Prediction': pd.Series(predictions.ravel())\n}\n\ndf_submission = pd.DataFrame(data)\ndf_submission.to_csv(\"submission.tsv\",sep=\"\\t\")\n","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:38:27.140116Z","iopub.execute_input":"2024-04-05T17:38:27.140495Z","iopub.status.idle":"2024-04-05T17:43:58.736306Z","shell.execute_reply.started":"2024-04-05T17:38:27.140466Z","shell.execute_reply":"2024-04-05T17:43:58.734218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the predictions we will create the submission data frame.\n\n**Note**: This will take atleast **15 to 20** minutes to finish.","metadata":{}},{"cell_type":"code","source":"predictions.ravel()","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:44:06.163916Z","iopub.execute_input":"2024-04-05T17:44:06.164339Z","iopub.status.idle":"2024-04-05T17:44:06.17193Z","shell.execute_reply.started":"2024-04-05T17:44:06.164304Z","shell.execute_reply":"2024-04-05T17:44:06.170895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_submission","metadata":{"execution":{"iopub.status.busy":"2024-04-05T17:43:58.739354Z","iopub.status.idle":"2024-04-05T17:43:58.739831Z","shell.execute_reply.started":"2024-04-05T17:43:58.739588Z","shell.execute_reply":"2024-04-05T17:43:58.739607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}