{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ? \n\nSimple Baseline to start with : Covert MultiLabel to MultiTarge  + Embeddings + Ridge \n\n    Features - precalculated embeddings for protein sequences. Thanks to Grandmaster Sergei Fironov for sharing protein emebedding calculated by T5 protein language model from the Rost Lab. \n    \n    Targets - multi-label is converted to mult-target (binary classification) task - i.e. for each sample we are preciting the probability that this label is assigned to that sample. In total there can be 40 000 labels - that is too much, so we choose only N the most frequent ones. \n    \n    After that - use any ML-model you like to make predictions. Start with Ridge as the he most simple and fast one. \n    ","metadata":{}},{"cell_type":"markdown","source":"Thanks to all  authors of the public notebooks and datasets which are quite helpful (please upvote them) and especially those ones:\n\nLEONID KULYK: https://www.kaggle.com/code/leonidkulyk/eda-cafa5-pfp-interactive-dags-plotly\n\nMARÍLIA PRATA: https://www.kaggle.com/code/mpwolke/cafa-5-protein-prediction\n\nDAREK KŁECZEK:  https://www.kaggle.com/code/thedrcat/cafa-eda\n\nD_KHATRI:  https://www.kaggle.com/code/dhruvkhatri/naive-submission-afa\n\n* Pretrained T5 protein embeddings: \n    * https://www.kaggle.com/datasets/danofer/uniprotkbswiss-prot-protein-embeddings\n\nGrandmaster Sergei Fironov shared protein emebedding calculated by T5 protein language model from the Rost Lab:  https://www.kaggle.com/datasets/sergeifironov/t5embeds\n\n","metadata":{}},{"cell_type":"markdown","source":"# Key param(s)\n\n","metadata":{}},{"cell_type":"code","source":"n_labels_to_consider = 1499 # We will choose only top frequent labels (in train) and predict only them. \nn_max_preds = 1499","metadata":{"execution":{"iopub.status.busy":"2023-05-10T19:21:23.613301Z","iopub.execute_input":"2023-05-10T19:21:23.614614Z","iopub.status.idle":"2023-05-10T19:21:23.65694Z","shell.execute_reply.started":"2023-05-10T19:21:23.614557Z","shell.execute_reply":"2023-05-10T19:21:23.655444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time\nt0start = time.time() \n\nimport numpy as np\nimport pandas as pd \nimport os\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.linear_model import Ridge,RidgeCV\nfrom sklearn.neural_network import MLPClassifier\n\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-05-10T19:21:23.660246Z","iopub.execute_input":"2023-05-10T19:21:23.661626Z","iopub.status.idle":"2023-05-10T19:21:24.406073Z","shell.execute_reply.started":"2023-05-10T19:21:23.661573Z","shell.execute_reply":"2023-05-10T19:21:24.405004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prepare multi-target Y  ( transition from multi-label task to multi target task - binary classifiction ). ","metadata":{}},{"cell_type":"markdown","source":"## Load train labels and select the most frequent ones","metadata":{}},{"cell_type":"code","source":"%%time\ntrainTerms = pd.read_csv(\"/kaggle/input/cafa-5-protein-function-prediction/Train/train_terms.tsv\",sep=\"\\t\")\nprint(trainTerms.shape)\ndisplay(trainTerms.head(2))\nvec_freqCount = (trainTerms['term'].value_counts())\nprint(vec_freqCount )","metadata":{"execution":{"iopub.status.busy":"2023-05-10T19:21:24.407609Z","iopub.execute_input":"2023-05-10T19:21:24.408232Z","iopub.status.idle":"2023-05-10T19:21:29.417468Z","shell.execute_reply.started":"2023-05-10T19:21:24.408192Z","shell.execute_reply":"2023-05-10T19:21:29.41611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## drop very rares\nvec_freqCount = vec_freqCount[vec_freqCount>=30]\nprint(vec_freqCount.shape[0])\nvec_freqCount.describe().round()","metadata":{"execution":{"iopub.status.busy":"2023-05-10T19:21:29.41915Z","iopub.execute_input":"2023-05-10T19:21:29.419624Z","iopub.status.idle":"2023-05-10T19:21:29.440956Z","shell.execute_reply.started":"2023-05-10T19:21:29.419572Z","shell.execute_reply":"2023-05-10T19:21:29.439086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"vec_freqCount[vec_freqCount>200].shape[0]","metadata":{"execution":{"iopub.status.busy":"2023-05-10T19:21:29.445675Z","iopub.execute_input":"2023-05-10T19:21:29.446196Z","iopub.status.idle":"2023-05-10T19:21:29.45575Z","shell.execute_reply.started":"2023-05-10T19:21:29.446153Z","shell.execute_reply":"2023-05-10T19:21:29.454177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print()\nlabels_to_consider = list(vec_freqCount.index[:n_labels_to_consider] )\nprint('n_labels_to_consider:', len(labels_to_consider), 'First 10:', labels_to_consider[:10] ) ","metadata":{"execution":{"iopub.status.busy":"2023-05-10T19:21:29.457736Z","iopub.execute_input":"2023-05-10T19:21:29.458127Z","iopub.status.idle":"2023-05-10T19:21:29.467634Z","shell.execute_reply.started":"2023-05-10T19:21:29.458092Z","shell.execute_reply":"2023-05-10T19:21:29.465945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load protein Ids in train","metadata":{}},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/t5embeds/train_ids.npy'\nvec_train_protein_ids = np.load(fn)\nprint(vec_train_protein_ids.shape)\nvec_train_protein_ids","metadata":{"execution":{"iopub.status.busy":"2023-05-10T19:21:29.469084Z","iopub.execute_input":"2023-05-10T19:21:29.469447Z","iopub.status.idle":"2023-05-10T19:21:29.541552Z","shell.execute_reply.started":"2023-05-10T19:21:29.469409Z","shell.execute_reply":"2023-05-10T19:21:29.540267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Prepare Y ","metadata":{}},{"cell_type":"code","source":"%%time \ntrain_size = 142246 # len(X)\nY = np.zeros( (train_size ,n_labels_to_consider) )\nprint(Y.shape)\n\nseries_train_protein_ids = pd.Series(vec_train_protein_ids ) # \n\ntrainTerms_smaller = trainTerms[ trainTerms['term'].isin( labels_to_consider ) ] # to speed-up the next step \nprint( trainTerms_smaller.shape)\n\nfor i in range(Y.shape[1]):\n    m = trainTerms_smaller['term'] ==  labels_to_consider[i]\n#     m.sum()\n    Y[:,i] =  series_train_protein_ids.isin(  set(trainTerms_smaller[m]['EntryID'] ) ).astype(float )\n    if (i % 10) == 0: \n        print(i, m.sum())\nY ","metadata":{"execution":{"iopub.status.busy":"2023-05-10T19:57:47.585984Z","iopub.execute_input":"2023-05-10T19:57:47.586966Z","iopub.status.idle":"2023-05-10T20:18:59.736348Z","shell.execute_reply.started":"2023-05-10T19:57:47.586914Z","shell.execute_reply":"2023-05-10T20:18:59.734939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \n# save for possible future reuse \nfn4saveY = 'Y_'+str(Y.shape[1])\nprint(fn4saveY)\nnp.save( fn4saveY , Y) ","metadata":{"execution":{"iopub.status.busy":"2023-05-10T20:19:06.807977Z","iopub.execute_input":"2023-05-10T20:19:06.808407Z","iopub.status.idle":"2023-05-10T20:19:10.329462Z","shell.execute_reply.started":"2023-05-10T20:19:06.808364Z","shell.execute_reply":"2023-05-10T20:19:10.327839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfn4save_labels = 'Y_'+str(Y.shape[1]) + '_labels'\nnp.save(fn4save_labels, labels_to_consider )","metadata":{"execution":{"iopub.status.busy":"2023-05-10T19:41:41.665632Z","iopub.execute_input":"2023-05-10T19:41:41.666026Z","iopub.status.idle":"2023-05-10T19:41:41.674617Z","shell.execute_reply.started":"2023-05-10T19:41:41.665989Z","shell.execute_reply":"2023-05-10T19:41:41.673233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print( list(np.load(fn4save_labels +'.npy' ))[:10] )","metadata":{"execution":{"iopub.status.busy":"2023-05-10T19:41:41.676599Z","iopub.execute_input":"2023-05-10T19:41:41.67724Z","iopub.status.idle":"2023-05-10T19:41:41.683933Z","shell.execute_reply.started":"2023-05-10T19:41:41.677185Z","shell.execute_reply":"2023-05-10T19:41:41.682289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \n# Someone may prefer  Y as dataframe \nif 1:\n    df_Y = pd.DataFrame(data = Y, columns = labels_to_consider)\n    display(df_Y.head(2))\n#     print( df.info().sum() )\n    print('memory_usage:', df_Y.memory_usage(index=True).sum() )\n    display(df_Y.describe() )    \n    fn4save =  'df_Y_'+str(Y.shape[1]) + '.csv'\n    df_Y.to_csv(fn4save)","metadata":{"execution":{"iopub.status.busy":"2023-05-10T19:41:41.685704Z","iopub.execute_input":"2023-05-10T19:41:41.686224Z","iopub.status.idle":"2023-05-10T19:44:23.411139Z","shell.execute_reply.started":"2023-05-10T19:41:41.686169Z","shell.execute_reply":"2023-05-10T19:44:23.409643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load train features - precalculated embeddings for the proteins","metadata":{}},{"cell_type":"code","source":"%%time\n\n# fn = '/kaggle/input/protein-embeddings-1/reduced_embeddings_file.npy'\n# fn = '/kaggle/input/protein-embeddings-1/embed_protbert_train_clip_1200_first_70000_prot.csv'\nfn = '/kaggle/input/t5embeds/train_embeds.npy'\n# fn = '/kaggle/input/t5embeds/test_embeds.npy'\n\nprint(fn)\nif '.csv' in fn:\n    df = pd.read_csv(fn, index_col = 0)\n    X = df.values\nelif '.npy' in fn:\n    X = np.load(fn)\nprint(X.shape)\nX","metadata":{"execution":{"iopub.status.busy":"2023-05-10T20:34:36.454685Z","iopub.execute_input":"2023-05-10T20:34:36.455519Z","iopub.status.idle":"2023-05-10T20:34:37.1474Z","shell.execute_reply.started":"2023-05-10T20:34:36.455468Z","shell.execute_reply":"2023-05-10T20:34:37.146001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load protein Ids ","metadata":{}},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/t5embeds/train_ids.npy'\nvec_train_protein_ids = np.load(fn)\nprint(vec_train_protein_ids.shape)\nvec_train_protein_ids","metadata":{"execution":{"iopub.status.busy":"2023-05-10T20:34:52.382152Z","iopub.execute_input":"2023-05-10T20:34:52.382642Z","iopub.status.idle":"2023-05-10T20:34:52.395728Z","shell.execute_reply.started":"2023-05-10T20:34:52.382597Z","shell.execute_reply":"2023-05-10T20:34:52.394715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Sanity check \n\nIds from the train data are the same as from the train labels data","metadata":{}},{"cell_type":"code","source":"s = set(vec_train_protein_ids) &set (trainTerms['EntryID'] )\nprint( len(s), len( X ) )  # get same numbers ","metadata":{"execution":{"iopub.status.busy":"2023-05-10T20:35:11.708831Z","iopub.execute_input":"2023-05-10T20:35:11.709317Z","iopub.status.idle":"2023-05-10T20:35:12.764341Z","shell.execute_reply.started":"2023-05-10T20:35:11.709272Z","shell.execute_reply":"2023-05-10T20:35:12.761298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prepare Train-Test split ","metadata":{"execution":{"iopub.status.busy":"2023-04-24T10:45:13.246257Z","iopub.execute_input":"2023-04-24T10:45:13.247139Z","iopub.status.idle":"2023-04-24T10:45:23.612468Z","shell.execute_reply.started":"2023-04-24T10:45:13.247095Z","shell.execute_reply":"2023-04-24T10:45:23.611491Z"}}},{"cell_type":"code","source":"IX = np.arange(len(X))\nIX_train, IX_test, _,_ = train_test_split( IX, IX, train_size=0.1, random_state=42)\nprint(len(IX_train), len(IX_test),  IX_train[:10], IX_test[:10] )","metadata":{"execution":{"iopub.status.busy":"2023-05-10T20:35:16.611099Z","iopub.execute_input":"2023-05-10T20:35:16.611563Z","iopub.status.idle":"2023-05-10T20:35:16.631949Z","shell.execute_reply.started":"2023-05-10T20:35:16.611524Z","shell.execute_reply":"2023-05-10T20:35:16.630265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling","metadata":{}},{"cell_type":"code","source":"from sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.ensemble import RandomForestClassifier\n\n# model = Ridge(alpha=1.0) # 0.805\nmodel = RidgeCV() # 0.8127 auc, with 15% train\n# model = RandomForestClassifier(n_estimators=200,  max_depth=14, min_samples_split=3, min_samples_leaf=1,n_jobs=-1) ## much slower... \n# model =MLPClassifier(hidden_layer_sizes=(512,256), early_stopping=True,\n#                      validation_fraction=0.05,learning_rate=\"adaptive\",learning_rate_init=0.005) # 0.59 rocauc , and slower\nstr_model_id = 'Ridge1'\n\ndf_models_stat = pd.DataFrame()\nmodel","metadata":{"execution":{"iopub.status.busy":"2023-05-10T20:35:25.189441Z","iopub.execute_input":"2023-05-10T20:35:25.190584Z","iopub.status.idle":"2023-05-10T20:35:25.204853Z","shell.execute_reply.started":"2023-05-10T20:35:25.190531Z","shell.execute_reply":"2023-05-10T20:35:25.20328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \nimport time\nfrom sklearn.metrics import roc_auc_score\n\nt0 = time.time()\nmodel.fit(X[IX_train,:],Y[IX_train,:])\nY_pred_test = model.predict(X[IX_test,:])\ntt = time.time() - t0\nprint(str_model_id, tt)\nl = []\nfor i in range(Y.shape[1]):\n    if len(np.unique(Y[IX_test,i]) ) > 1:\n        s = roc_auc_score(Y[IX_test,i], Y_pred_test[:,i]);\n    else:\n        s = 0.5\n    l.append(s)        \n    if i %10 == 0:\n        print(i, s)\ndf_models_stat.loc[str_model_id,'RocAuc Mean Test'] = np.mean(l)\ndf_models_stat.loc[str_model_id,'Time'] = np.round(tt,1)\ndf_models_stat.loc[str_model_id,'Test Size'] = len(IX_test)\ndf_models_stat","metadata":{"execution":{"iopub.status.busy":"2023-05-10T20:35:26.548582Z","iopub.execute_input":"2023-05-10T20:35:26.549329Z","iopub.status.idle":"2023-05-10T20:37:24.968295Z","shell.execute_reply.started":"2023-05-10T20:35:26.549283Z","shell.execute_reply":"2023-05-10T20:37:24.967018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.get_params()","metadata":{"execution":{"iopub.status.busy":"2023-05-10T20:37:24.970626Z","iopub.execute_input":"2023-05-10T20:37:24.971045Z","iopub.status.idle":"2023-05-10T20:37:24.980936Z","shell.execute_reply.started":"2023-05-10T20:37:24.971005Z","shell.execute_reply":"2023-05-10T20:37:24.979544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Scores statistics over targets ","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nplt.hist(l)\nplt.show()\npd.Series(l).describe()","metadata":{"execution":{"iopub.status.busy":"2023-05-10T20:37:24.982413Z","iopub.execute_input":"2023-05-10T20:37:24.98283Z","iopub.status.idle":"2023-05-10T20:37:25.259634Z","shell.execute_reply.started":"2023-05-10T20:37:24.982777Z","shell.execute_reply":"2023-05-10T20:37:25.258289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Retrain model on the the full sample ","metadata":{}},{"cell_type":"code","source":"%%time\nmodel.fit(X,Y)","metadata":{"execution":{"iopub.status.busy":"2023-05-10T20:37:25.262515Z","iopub.execute_input":"2023-05-10T20:37:25.262919Z","iopub.status.idle":"2023-05-10T20:38:52.122188Z","shell.execute_reply.started":"2023-05-10T20:37:25.262867Z","shell.execute_reply":"2023-05-10T20:38:52.12075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission preparations Step 1 - load features and calculate predictions ","metadata":{}},{"cell_type":"markdown","source":"## Load features for submission","metadata":{}},{"cell_type":"code","source":"%%time\n# fn = '/kaggle/input/protein-embeddings-1/reduced_embeddings_file.npy'\n# fn = '/kaggle/input/protein-embeddings-1/embed_protbert_train_clip_1200_first_70000_prot.csv'\n# fn = '/kaggle/input/t5embeds/train_embeds.npy'\nfn = '/kaggle/input/t5embeds/test_embeds.npy'\nprint(fn)\nX_submit = np.load(fn)\nprint(X_submit.shape)\n# X_submit","metadata":{"execution":{"iopub.status.busy":"2023-05-10T20:38:52.124174Z","iopub.execute_input":"2023-05-10T20:38:52.124685Z","iopub.status.idle":"2023-05-10T20:39:04.546004Z","shell.execute_reply.started":"2023-05-10T20:38:52.124633Z","shell.execute_reply":"2023-05-10T20:39:04.544604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Calculate prediction for submission","metadata":{}},{"cell_type":"code","source":"%%time\nY_submit =  model.predict(X_submit)\nprint(Y_submit.shape)","metadata":{"execution":{"iopub.status.busy":"2023-05-10T20:39:04.547856Z","iopub.execute_input":"2023-05-10T20:39:04.548607Z","iopub.status.idle":"2023-05-10T20:39:12.19982Z","shell.execute_reply.started":"2023-05-10T20:39:04.548565Z","shell.execute_reply":"2023-05-10T20:39:12.198245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission preparations Step 2 - prepare submision in desired format  ","metadata":{}},{"cell_type":"code","source":"%%time \ndf_finalSubmission = pd.DataFrame(columns = ['Protein Id', 'GO Term Id','Prediction'])","metadata":{"execution":{"iopub.status.busy":"2023-05-10T20:39:12.201249Z","iopub.execute_input":"2023-05-10T20:39:12.201644Z","iopub.status.idle":"2023-05-10T20:39:12.212382Z","shell.execute_reply.started":"2023-05-10T20:39:12.201601Z","shell.execute_reply":"2023-05-10T20:39:12.210854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load protein ids for the submission","metadata":{}},{"cell_type":"code","source":"%%time\nfn = '/kaggle/input/t5embeds/test_ids.npy'\nvec_test_protein_ids = np.load(fn)\nprint(vec_test_protein_ids.shape)\nvec_test_protein_ids","metadata":{"execution":{"iopub.status.busy":"2023-05-10T20:39:12.214031Z","iopub.execute_input":"2023-05-10T20:39:12.214387Z","iopub.status.idle":"2023-05-10T20:39:12.289168Z","shell.execute_reply.started":"2023-05-10T20:39:12.214352Z","shell.execute_reply":"2023-05-10T20:39:12.287808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## \"Melt\" protein ids ","metadata":{}},{"cell_type":"code","source":"%%time \nl = []\nfor k in list(vec_test_protein_ids):\n    l += [ k] * Y_submit.shape[1]\nprint(len(l), l[:20])    \n\ndf_finalSubmission['Protein Id'] = l","metadata":{"execution":{"iopub.status.busy":"2023-05-10T20:39:12.29138Z","iopub.execute_input":"2023-05-10T20:39:12.291976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# %%time \n# df_finalSubmission.head(3)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## \"Melt\" Labels (Gene ontology terms )","metadata":{"execution":{"iopub.status.busy":"2023-04-24T12:31:15.267721Z","iopub.execute_input":"2023-04-24T12:31:15.268181Z","iopub.status.idle":"2023-04-24T12:31:15.613742Z","shell.execute_reply.started":"2023-04-24T12:31:15.268143Z","shell.execute_reply":"2023-04-24T12:31:15.612514Z"}}},{"cell_type":"code","source":"df_finalSubmission['GO Term Id'] = labels_to_consider * Y_submit.shape[0]\n# df_finalSubmission.head(3)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Assign predictions ","metadata":{}},{"cell_type":"code","source":"df_finalSubmission['Prediction'] = Y_submit.ravel()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(df_finalSubmission)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### drop 0 preds and negatives\n* opt: sort by score, keep top K per Protein\n* warning : will be slooow with this many rows!","metadata":{}},{"cell_type":"code","source":"%%time\ndf_finalSubmission['Prediction'] = df_finalSubmission['Prediction'].round(3)\ndf_finalSubmission = df_finalSubmission[df_finalSubmission['Prediction']>0]\ndf_finalSubmission.shape[0]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_finalSubmission","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Save ","metadata":{}},{"cell_type":"code","source":"%%time \ndf_finalSubmission.to_csv(\"submission.tsv\",header=False, index=False, sep=\"\\t\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Show some info ","metadata":{}},{"cell_type":"code","source":"# %%time \n# df_finalSubmission.info()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \ndf_finalSubmission.describe()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nplt.figure(figsize = (15,4))\nplt.hist(df_finalSubmission['Prediction'].values, bins = 300 )\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_finalSubmission","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_finalSubmission.iloc[:,0:2].nunique()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_finalSubmission.shape[0]/141864 ## num proteins","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_finalSubmission.shape[0]/n_labels_to_consider","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}