{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# My EDA\n\nDo EDA on the CAFA 5 data","metadata":{"papermill":{"duration":0.008875,"end_time":"2023-05-09T08:30:09.052999","exception":false,"start_time":"2023-05-09T08:30:09.044124","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# Import Libraries","metadata":{"papermill":{"duration":0.009086,"end_time":"2023-05-09T08:30:09.140473","exception":false,"start_time":"2023-05-09T08:30:09.131387","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt","metadata":{"papermill":{"duration":9.85331,"end_time":"2023-05-09T08:30:19.002985","exception":false,"start_time":"2023-05-09T08:30:09.149675","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-28T20:31:42.738064Z","iopub.execute_input":"2023-06-28T20:31:42.738432Z","iopub.status.idle":"2023-06-28T20:31:43.368119Z","shell.execute_reply.started":"2023-06-28T20:31:42.738402Z","shell.execute_reply":"2023-06-28T20:31:43.367333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load the train_terms Dataset","metadata":{"papermill":{"duration":0.008429,"end_time":"2023-05-09T08:30:19.047756","exception":false,"start_time":"2023-05-09T08:30:19.039327","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"First we will load the file `train_terms.tsv` which contains the list of annotated terms (functions) for the proteins. We will extract the labels aka `GO term ID` and create a label dataframe for the protein embeddings.","metadata":{"papermill":{"duration":0.008388,"end_time":"2023-05-09T08:30:19.065367","exception":false,"start_time":"2023-05-09T08:30:19.056979","status":"completed"},"tags":[]}},{"cell_type":"code","source":"train_terms = pd.read_csv(\"/kaggle/input/cafa-5-protein-function-prediction/Train/train_terms.tsv\",sep=\"\\t\")\nprint(train_terms.shape)","metadata":{"papermill":{"duration":3.69155,"end_time":"2023-05-09T08:30:22.766144","exception":false,"start_time":"2023-05-09T08:30:19.074594","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-28T20:32:03.42942Z","iopub.execute_input":"2023-06-28T20:32:03.429796Z","iopub.status.idle":"2023-06-28T20:32:06.805176Z","shell.execute_reply.started":"2023-06-28T20:32:03.429764Z","shell.execute_reply":"2023-06-28T20:32:06.804071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"`train_terms` dataframe is composed of 3 columns and 5363863 entries. We can see all 3 dimensions of our dataset by printing out the first 5 entries using the following code:","metadata":{"papermill":{"duration":0.008358,"end_time":"2023-05-09T08:30:22.783293","exception":false,"start_time":"2023-05-09T08:30:22.774935","status":"completed"},"tags":[]}},{"cell_type":"code","source":"train_terms.head()","metadata":{"papermill":{"duration":0.038607,"end_time":"2023-05-09T08:30:22.830633","exception":false,"start_time":"2023-05-09T08:30:22.792026","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-28T20:32:16.622577Z","iopub.execute_input":"2023-06-28T20:32:16.622962Z","iopub.status.idle":"2023-06-28T20:32:16.6417Z","shell.execute_reply.started":"2023-06-28T20:32:16.622929Z","shell.execute_reply":"2023-06-28T20:32:16.639789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If we look at the first entry of `train_terms.tsv`, we can see that it contains protein id(`A0A009IHW8`), the GO term(`GO:0008152`) and its aspect(`BPO`). ","metadata":{"papermill":{"duration":0.008764,"end_time":"2023-05-09T08:30:22.848867","exception":false,"start_time":"2023-05-09T08:30:22.840103","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# Look at the content of train_sequence.fasta","metadata":{}},{"cell_type":"code","source":"with open(\"/kaggle/input/cafa-5-protein-function-prediction/Train/train_sequences.fasta\", \"r\") as file:\n    fasta_100 = file.readlines()[:100]","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:32:29.722119Z","iopub.execute_input":"2023-06-28T20:32:29.722481Z","iopub.status.idle":"2023-06-28T20:32:31.419215Z","shell.execute_reply.started":"2023-06-28T20:32:29.722449Z","shell.execute_reply":"2023-06-28T20:32:31.418277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fasta_100[:10]","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:32:37.402253Z","iopub.execute_input":"2023-06-28T20:32:37.402658Z","iopub.status.idle":"2023-06-28T20:32:37.416648Z","shell.execute_reply.started":"2023-06-28T20:32:37.402598Z","shell.execute_reply":"2023-06-28T20:32:37.410997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Understand the fasta header: https://www.uniprot.org/help/fasta-headers","metadata":{}},{"cell_type":"markdown","source":"# Loading the protein embeddings\n\n\nWe will now load the pre calculated protein embeddings created by [Sergei Fironov](https://www.kaggle.com/sergeifironov) using the Rost Lab's T5 protein language model.\n\nIf the `tfembeds` is not yet on the input data of the notebook, you can add it to your enviromentby clicking on `Add Data` and search for `t5embeds` (make sure that it's the correct [one](https://www.kaggle.com/datasets/sergeifironov/t5embeds) ) and then click on the `+` beside it.\n\nThe protein embeddings to be used for training are recorded in `train_embeds.npy` and the corresponding protein ids are available in `train_ids.npy`.","metadata":{}},{"cell_type":"markdown","source":"First, we will load the protein ids of the protein embeddings in the train dataset contained in `train_ids.npy` into a numpy array.","metadata":{"papermill":{"duration":0.009256,"end_time":"2023-05-09T08:30:22.867158","exception":false,"start_time":"2023-05-09T08:30:22.857902","status":"completed"},"tags":[]}},{"cell_type":"code","source":"train_protein_ids = np.load('/kaggle/input/t5embeds/train_ids.npy')\nprint(train_protein_ids.shape)","metadata":{"papermill":{"duration":0.067806,"end_time":"2023-05-09T08:30:22.944355","exception":false,"start_time":"2023-05-09T08:30:22.876549","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-28T20:33:02.905318Z","iopub.execute_input":"2023-06-28T20:33:02.905704Z","iopub.status.idle":"2023-06-28T20:33:02.963576Z","shell.execute_reply.started":"2023-06-28T20:33:02.905669Z","shell.execute_reply":"2023-06-28T20:33:02.962476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `train_protein_ids` array consists of 142246 protein_ids. Let us print out the first 5 entries using the following code:","metadata":{"papermill":{"duration":0.009498,"end_time":"2023-05-09T08:30:22.963291","exception":false,"start_time":"2023-05-09T08:30:22.953793","status":"completed"},"tags":[]}},{"cell_type":"code","source":"train_protein_ids[:5]","metadata":{"papermill":{"duration":0.019907,"end_time":"2023-05-09T08:30:22.992625","exception":false,"start_time":"2023-05-09T08:30:22.972718","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-28T20:33:07.682832Z","iopub.execute_input":"2023-06-28T20:33:07.683207Z","iopub.status.idle":"2023-06-28T20:33:07.690276Z","shell.execute_reply.started":"2023-06-28T20:33:07.683175Z","shell.execute_reply":"2023-06-28T20:33:07.689255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<!-- Now, we will load`train_embeds.py` which contains the pre-calculated embeddings of the proteins in the train dataset. with protein_ids (`id`s we loaded previously from the **train_ids.npy**) into a numpy array. This array now contains the precalculated embeddings for the protein_ids( Ids we loaded above from **train_ids.npy**) needed for training. -->\n\nAfter loading the files as numpy arrays, we will convert them into Pandas dataframe.\n\nEach protein embedding is a vector of length 1024. We create the resulting dataframe such that there are 1024 columns to represent the values in each of the 1024 places in the vector.","metadata":{"papermill":{"duration":0.009375,"end_time":"2023-05-09T08:30:23.011402","exception":false,"start_time":"2023-05-09T08:30:23.002027","status":"completed"},"tags":[]}},{"cell_type":"code","source":"train_embeddings = np.load('/kaggle/input/t5embeds/train_embeds.npy')\n\n# Now lets convert embeddings numpy array(train_embeddings) into pandas dataframe.\ncolumn_num = train_embeddings.shape[1]\ntrain_df = pd.DataFrame(train_embeddings, columns = [\"Column_\" + str(i) for i in range(1, column_num+1)])\nprint(train_df.shape)","metadata":{"papermill":{"duration":9.719957,"end_time":"2023-05-09T08:30:32.741095","exception":false,"start_time":"2023-05-09T08:30:23.021138","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-28T20:33:18.512047Z","iopub.execute_input":"2023-06-28T20:33:18.512415Z","iopub.status.idle":"2023-06-28T20:33:29.251334Z","shell.execute_reply.started":"2023-06-28T20:33:18.512382Z","shell.execute_reply":"2023-06-28T20:33:29.250322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"papermill":{"duration":0.036828,"end_time":"2023-05-09T08:30:32.807222","exception":false,"start_time":"2023-05-09T08:30:32.770394","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-28T20:33:38.874763Z","iopub.execute_input":"2023-06-28T20:33:38.875157Z","iopub.status.idle":"2023-06-28T20:33:38.900064Z","shell.execute_reply.started":"2023-06-28T20:33:38.875122Z","shell.execute_reply":"2023-06-28T20:33:38.899164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Extract one label for a binary classification","metadata":{"papermill":{"duration":0.009208,"end_time":"2023-05-09T08:30:32.825978","exception":false,"start_time":"2023-05-09T08:30:32.81677","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"First we will extract all the needed labels(`GO term ID`) from `train_terms.tsv` file. There are more than 30,000 labels. In this exploration, we only extract one label for a binary classification.","metadata":{"papermill":{"duration":0.009674,"end_time":"2023-05-09T08:30:32.845065","exception":false,"start_time":"2023-05-09T08:30:32.835391","status":"completed"},"tags":[]}},{"cell_type":"code","source":"train_terms.shape","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:34:20.660856Z","iopub.execute_input":"2023-06-28T20:34:20.661237Z","iopub.status.idle":"2023-06-28T20:34:20.667388Z","shell.execute_reply.started":"2023-06-28T20:34:20.661202Z","shell.execute_reply":"2023-06-28T20:34:20.665985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_terms['term'].nunique()","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:34:23.614524Z","iopub.execute_input":"2023-06-28T20:34:23.614927Z","iopub.status.idle":"2023-06-28T20:34:24.094065Z","shell.execute_reply.started":"2023-06-28T20:34:23.61489Z","shell.execute_reply":"2023-06-28T20:34:24.09312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's plot the most frequent 100 `GO Term ID`s in `train_terms.tsv`.","metadata":{"papermill":{"duration":0.009238,"end_time":"2023-05-09T08:30:32.863785","exception":false,"start_time":"2023-05-09T08:30:32.854547","status":"completed"},"tags":[]}},{"cell_type":"code","source":"train_terms.term.value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:34:30.663644Z","iopub.execute_input":"2023-06-28T20:34:30.664005Z","iopub.status.idle":"2023-06-28T20:34:31.4899Z","shell.execute_reply.started":"2023-06-28T20:34:30.663974Z","shell.execute_reply":"2023-06-28T20:34:31.488765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Select first 100 values for plotting\nplot_df = train_terms['term'].value_counts().iloc[:100]\n\nfigure, axis = plt.subplots(1, 1, figsize=(12, 6))\n\nbp = sns.barplot(ax=axis, x=np.array(plot_df.index), y=plot_df.values)\nbp.set_xticklabels(bp.get_xticklabels(), rotation=90, size = 6)\naxis.set_title('Top 100 frequent GO term IDs')\nbp.set_xlabel(\"GO term IDs\", fontsize = 12)\nbp.set_ylabel(\"Count\", fontsize = 12)\nplt.show()","metadata":{"papermill":{"duration":1.592489,"end_time":"2023-05-09T08:30:34.465912","exception":false,"start_time":"2023-05-09T08:30:32.873423","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-28T20:34:41.278866Z","iopub.execute_input":"2023-06-28T20:34:41.279456Z","iopub.status.idle":"2023-06-28T20:34:43.181011Z","shell.execute_reply.started":"2023-06-28T20:34:41.279417Z","shell.execute_reply":"2023-06-28T20:34:43.180111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will now extrac the top 1 GO term Id.","metadata":{"papermill":{"duration":0.010458,"end_time":"2023-05-09T08:30:34.487707","exception":false,"start_time":"2023-05-09T08:30:34.477249","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Set the limit for label\nnum_of_labels = 1\n\n# Take value counts in descending order and fetch first `GO term ID` as label\nlabel_top1 = train_terms['term'].value_counts().index[:num_of_labels].tolist()[0]","metadata":{"papermill":{"duration":0.523976,"end_time":"2023-05-09T08:30:35.021974","exception":false,"start_time":"2023-05-09T08:30:34.497998","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-28T20:35:55.870003Z","iopub.execute_input":"2023-06-28T20:35:55.870377Z","iopub.status.idle":"2023-06-28T20:35:56.710274Z","shell.execute_reply.started":"2023-06-28T20:35:55.870344Z","shell.execute_reply":"2023-06-28T20:35:56.70915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"label_top1","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:36:02.032997Z","iopub.execute_input":"2023-06-28T20:36:02.03335Z","iopub.status.idle":"2023-06-28T20:36:02.038881Z","shell.execute_reply.started":"2023-06-28T20:36:02.033319Z","shell.execute_reply":"2023-06-28T20:36:02.038011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, we will create a new dataframe by filtering the train terms with the selected `GO Term ID`.","metadata":{"papermill":{"duration":0.009833,"end_time":"2023-05-09T08:30:35.042088","exception":false,"start_time":"2023-05-09T08:30:35.032255","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Fetch the train_terms data for the relevant labels only\ntrain_terms_updated = train_terms.loc[train_terms['term'] == label_top1]","metadata":{"papermill":{"duration":0.668657,"end_time":"2023-05-09T08:30:35.720953","exception":false,"start_time":"2023-05-09T08:30:35.052296","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-28T20:36:32.178698Z","iopub.execute_input":"2023-06-28T20:36:32.179071Z","iopub.status.idle":"2023-06-28T20:36:33.018662Z","shell.execute_reply.started":"2023-06-28T20:36:32.179037Z","shell.execute_reply":"2023-06-28T20:36:33.017677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_terms_updated.shape, train_terms.shape","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:36:35.606772Z","iopub.execute_input":"2023-06-28T20:36:35.607134Z","iopub.status.idle":"2023-06-28T20:36:35.614459Z","shell.execute_reply.started":"2023-06-28T20:36:35.607103Z","shell.execute_reply":"2023-06-28T20:36:35.613144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_terms_updated.EntryID.nunique()","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:36:44.33021Z","iopub.execute_input":"2023-06-28T20:36:44.33058Z","iopub.status.idle":"2023-06-28T20:36:44.36274Z","shell.execute_reply.started":"2023-06-28T20:36:44.330547Z","shell.execute_reply":"2023-06-28T20:36:44.361685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Use LogisticRegression to classify all proteins into GO:0005575 or not","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LogisticRegression","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:36:58.095714Z","iopub.execute_input":"2023-06-28T20:36:58.096089Z","iopub.status.idle":"2023-06-28T20:36:58.32054Z","shell.execute_reply.started":"2023-06-28T20:36:58.096055Z","shell.execute_reply":"2023-06-28T20:36:58.319568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lr = LogisticRegression()","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:36:58.851959Z","iopub.execute_input":"2023-06-28T20:36:58.852676Z","iopub.status.idle":"2023-06-28T20:36:58.857386Z","shell.execute_reply.started":"2023-06-28T20:36:58.852631Z","shell.execute_reply":"2023-06-28T20:36:58.856528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_protein_ids.shape","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:37:02.356902Z","iopub.execute_input":"2023-06-28T20:37:02.357292Z","iopub.status.idle":"2023-06-28T20:37:02.36498Z","shell.execute_reply.started":"2023-06-28T20:37:02.357259Z","shell.execute_reply":"2023-06-28T20:37:02.363763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create an empty dataframe of required size for storing the labels,\n# i.e, train_size x num_of_labels (142246 x 1)\ntrain_size = train_protein_ids.shape[0] # len(X)\ntrain_labels = np.zeros((train_size ,num_of_labels))\n\n# Convert from numpy to pandas series for better handling\nseries_train_protein_ids = pd.Series(train_protein_ids)","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:37:21.234003Z","iopub.execute_input":"2023-06-28T20:37:21.234364Z","iopub.status.idle":"2023-06-28T20:37:21.26342Z","shell.execute_reply.started":"2023-06-28T20:37:21.234331Z","shell.execute_reply":"2023-06-28T20:37:21.262419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# In the series_train_protein_ids pandas series, if a protein is in\n# the set of IDs whose term is label_top1, i.e., the set of train_terms_updated.EntryIDs, \n# then mark the train_lables as 1, else 0.\ntrain_labels[:,0] =  series_train_protein_ids.isin(train_terms_updated.EntryID).astype(float)","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:48:06.707292Z","iopub.execute_input":"2023-06-28T20:48:06.7077Z","iopub.status.idle":"2023-06-28T20:48:06.750112Z","shell.execute_reply.started":"2023-06-28T20:48:06.707659Z","shell.execute_reply":"2023-06-28T20:48:06.748794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels_df = pd.DataFrame(train_labels, columns=[label_top1])","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:52:01.711091Z","iopub.execute_input":"2023-06-28T20:52:01.711464Z","iopub.status.idle":"2023-06-28T20:52:01.717503Z","shell.execute_reply.started":"2023-06-28T20:52:01.711432Z","shell.execute_reply":"2023-06-28T20:52:01.715689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The final labels dataframe (`label_df`) is composed of 1 column and 142246 entries:","metadata":{"papermill":{"duration":0.010097,"end_time":"2023-05-09T08:38:51.971947","exception":false,"start_time":"2023-05-09T08:38:51.96185","status":"completed"},"tags":[]}},{"cell_type":"code","source":"labels_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:52:06.608424Z","iopub.execute_input":"2023-06-28T20:52:06.608799Z","iopub.status.idle":"2023-06-28T20:52:06.61811Z","shell.execute_reply.started":"2023-06-28T20:52:06.608766Z","shell.execute_reply":"2023-06-28T20:52:06.617197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training\n\nNext, we will use the protein embeddings to train a LogisticRegression classifier","metadata":{"papermill":{"duration":0.010523,"end_time":"2023-05-09T08:38:52.052433","exception":false,"start_time":"2023-05-09T08:38:52.04191","status":"completed"},"tags":[]}},{"cell_type":"code","source":"train_df.shape, labels_df.shape","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:52:25.617229Z","iopub.execute_input":"2023-06-28T20:52:25.61773Z","iopub.status.idle":"2023-06-28T20:52:25.625661Z","shell.execute_reply.started":"2023-06-28T20:52:25.617684Z","shell.execute_reply":"2023-06-28T20:52:25.62471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lr.fit(train_df, labels_df.iloc[:, 0])","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:53:30.35684Z","iopub.execute_input":"2023-06-28T20:53:30.357201Z","iopub.status.idle":"2023-06-28T20:53:53.1664Z","shell.execute_reply.started":"2023-06-28T20:53:30.35717Z","shell.execute_reply":"2023-06-28T20:53:53.165005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission","metadata":{"papermill":{"duration":0.021665,"end_time":"2023-05-09T08:41:01.780867","exception":false,"start_time":"2023-05-09T08:41:01.759202","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"For submission we will use the protein embeddings of the test data created by [Sergei Fironov](https://www.kaggle.com/sergeifironov) using the Rost Lab's T5 protein language model.","metadata":{"papermill":{"duration":0.02075,"end_time":"2023-05-09T08:41:01.82296","exception":false,"start_time":"2023-05-09T08:41:01.80221","status":"completed"},"tags":[]}},{"cell_type":"code","source":"test_embeddings = np.load('/kaggle/input/t5embeds/test_embeds.npy')\n\n# Convert test_embeddings to dataframe\ncolumn_num = test_embeddings.shape[1]\ntest_df = pd.DataFrame(test_embeddings, columns = [\"Column_\" + str(i) for i in range(1, column_num+1)])\nprint(test_df.shape)","metadata":{"papermill":{"duration":10.290827,"end_time":"2023-05-09T08:41:12.134919","exception":false,"start_time":"2023-05-09T08:41:01.844092","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-28T20:54:07.767328Z","iopub.execute_input":"2023-06-28T20:54:07.76774Z","iopub.status.idle":"2023-06-28T20:54:18.942505Z","shell.execute_reply.started":"2023-06-28T20:54:07.767707Z","shell.execute_reply":"2023-06-28T20:54:18.941318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `test_df` is composed of 1024 columns and 141865 entries. We can see all 1024 dimensions(results will be truncated since column length is too long) of our dataset by printing out the first 5 entries using the following code:","metadata":{"papermill":{"duration":0.020857,"end_time":"2023-05-09T08:41:12.17776","exception":false,"start_time":"2023-05-09T08:41:12.156903","status":"completed"},"tags":[]}},{"cell_type":"code","source":"test_df.head()","metadata":{"papermill":{"duration":0.050123,"end_time":"2023-05-09T08:41:12.248732","exception":false,"start_time":"2023-05-09T08:41:12.198609","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-28T20:54:18.944183Z","iopub.execute_input":"2023-06-28T20:54:18.945019Z","iopub.status.idle":"2023-06-28T20:54:18.972217Z","shell.execute_reply.started":"2023-06-28T20:54:18.94497Z","shell.execute_reply":"2023-06-28T20:54:18.971154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will now use the model to make predictions on the test embeddings. ","metadata":{}},{"cell_type":"code","source":"predictions =  lr.predict_proba(test_df)","metadata":{"papermill":{"duration":663.907351,"end_time":"2023-05-09T08:52:16.178461","exception":false,"start_time":"2023-05-09T08:41:12.27111","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-28T20:54:32.548458Z","iopub.execute_input":"2023-06-28T20:54:32.548853Z","iopub.status.idle":"2023-06-28T20:54:32.761408Z","shell.execute_reply.started":"2023-06-28T20:54:32.548817Z","shell.execute_reply":"2023-06-28T20:54:32.760118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:54:34.968087Z","iopub.execute_input":"2023-06-28T20:54:34.96845Z","iopub.status.idle":"2023-06-28T20:54:34.974659Z","shell.execute_reply.started":"2023-06-28T20:54:34.968417Z","shell.execute_reply":"2023-06-28T20:54:34.973667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions.shape","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:54:50.923612Z","iopub.execute_input":"2023-06-28T20:54:50.924019Z","iopub.status.idle":"2023-06-28T20:54:50.933289Z","shell.execute_reply.started":"2023-06-28T20:54:50.923985Z","shell.execute_reply":"2023-06-28T20:54:50.932024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_submission = pd.DataFrame(columns = ['Protein Id', 'GO Term Id','Prediction'])\ntest_protein_ids = np.load('/kaggle/input/t5embeds/test_ids.npy')","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:55:17.959315Z","iopub.execute_input":"2023-06-28T20:55:17.959737Z","iopub.status.idle":"2023-06-28T20:55:18.023507Z","shell.execute_reply.started":"2023-06-28T20:55:17.9597Z","shell.execute_reply":"2023-06-28T20:55:18.022483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_submission['Protein Id'] = test_protein_ids\ndf_submission['GO Term Id'] = 'GO:0005575'\ndf_submission['Prediction'] = predictions[:,1]","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:55:23.575241Z","iopub.execute_input":"2023-06-28T20:55:23.575611Z","iopub.status.idle":"2023-06-28T20:55:23.645731Z","shell.execute_reply.started":"2023-06-28T20:55:23.575578Z","shell.execute_reply":"2023-06-28T20:55:23.644787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_submission","metadata":{"execution":{"iopub.status.busy":"2023-06-28T20:55:27.197777Z","iopub.execute_input":"2023-06-28T20:55:27.198131Z","iopub.status.idle":"2023-06-28T20:55:27.211792Z","shell.execute_reply.started":"2023-06-28T20:55:27.198098Z","shell.execute_reply":"2023-06-28T20:55:27.210573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_submission.to_csv(\"submission.tsv\",header=False, index=False, sep=\"\\t\")","metadata":{"papermill":{"duration":0.063739,"end_time":"2023-05-09T08:52:16.292974","exception":false,"start_time":"2023-05-09T08:52:16.229235","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-06-28T20:57:39.605827Z","iopub.execute_input":"2023-06-28T20:57:39.606214Z","iopub.status.idle":"2023-06-28T20:57:40.216069Z","shell.execute_reply.started":"2023-06-28T20:57:39.60618Z","shell.execute_reply":"2023-06-28T20:57:40.215163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}