{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Protein Neural Network\n\n$Protein$ $Neural$ $Network$ is a type of `deep learning neural network` that is used to `predict` the `properties of proteins`. They are `trained` on a `large dataset` of `protein sequences`/`structures`. The network learns to `identify patterns` in the data that are associated with `particular properties`. $Protein$ $Neural$ $Networks$ have been shown to be `very effective` at `predicting protein properties`. In many cases, they `outperform traditional` methods that rely on `hand-crafted features`\n\n<img src = \"https://scontent.fjai2-1.fna.fbcdn.net/v/t1.6435-9/37298434_654990401529724_543197132539035648_n.jpg?_nc_cat=102&ccb=1-7&_nc_sid=730e14&_nc_ohc=7hJcasqo1PkAX8fs55i&_nc_ht=scontent.fjai2-1.fna&oh=00_AfBLiYMGoJgSMX0w-kd690Y2GvRX4ga07ER2OXTSeH7o5w&oe=64CE544A\" width = 400>\n\n\n# 1 | Basic Terminologies 💻\n\n* $Structure$ $of$ $a$ $Protein$\n* $Gene$ $Ontology$ $(GO)$\n* $Principal$ $Component$ $Analaysis$\n* $Transformers$\n* $Categorical$ $Cross$ $Entropy$\n* $Adaptive$ $Moment$ $Estimation$ $ADAM$\n\n## 1.1 | Structure of a Protein\n\nSo what is a actually a **Protein...?**\n\nFirst of all lets understand the structure of an **Atom**\n\n<img src = \"https://www.sciencefacts.net/wp-content/uploads/2020/11/Parts-of-an-Atom-Diagram.jpg\"  width = 300>\n\nThere is a really good image I found of the `structure of atom`. Though there are many debates on the structure like this, but this `model is accepted universaly at this moment`.\n\nIn the centre we have the `Neucleus`. The `Neucleus` is made up of $2$ more structures named as `Neutron` and `Proton`. A `Proton` is `positively charged element` and a `Neutron` is a `neutral charged element`. A `Electron`, `negatively charged element`, `orbits` this `Neucleus` at some `distance apart`.\n\nThe `more we increase the number` of `Electrons` and `Protons`. The `bigger the atoms becomes`.\n\nThere are `different shells` where the `Electrons reside`. The `more closer the shell` is, the `less Electrons` it contrains. There are mainly $4$ shells. \n\n|||\n|---|---\n|$K$|$2$\n|$L$|$8$\n|$M$|$18$\n|$N$|$32$\n\nOnce an atom `fills its outer most shell` with `Electrons`. It becomes `stable atom` and try to `refuse any donation` or `recieve of extra atom`.\n\n<img src = \"https://cdn1.byjus.com/wp-content/uploads/2022/01/word-image128.png\" width = 400>\n\n`Different atoms combine` to `share Electrons` and become `stable Molecules` \n\n<img src = \"https://www.astrochem.org/sci_img/Amino_Acid_Structure.jpg\" width = 300>\n\nA `Amino Acid` is made up of mainly $4$ different atoms\n`[H , C , O , N]`. \n\n<img src = \"https://upload.wikimedia.org/wikipedia/commons/thumb/5/51/L-amino_acid_structure.svg/1200px-L-amino_acid_structure.svg.png\" alt = \"Bro use Light Theme\" width = 300 >\n\nWe also have a free `Electron Pair` of `Carbon` in this molecule, we call this as a `Side Chain` which can be of different types. Basically this `Side Chain` provide the flexibility to make `different types` of `Amino Acids`. This flexibilty allows for $20$ `different` `Amino Acids` \n\nWhen we join `Amino Acids` with `peptide bonds`, we get `Proteins`. Conncecting different types of `Amino Acids` ends up in different types of `Proteins`.\n\n## 1.2 | Gene Ontology (GO)\n\n$Gene$ $Ontology$ $(GO)$ is a `controlled vocabulary` that `describes the functions of genes` and gene products. It is constantly being updated as new information becomes available.\n\nThere are mainly $3$ `Ontologies`\n\n||||\n|---|---|---\n|$Biological$ $Process$|Describes the `biological processes`|A gene product might be involved in the process of `cell cycle`/`signal transduction.`\n|$Cellular$ $Component$|Describes the `cellular components`|A gene product might be located in the `nucleus`/`cytoplasm`.\n|$Molecular$ $Function$|Describes the `molecular functions`|A gene product might be involved in the `catalysis of a reaction`/`binding of a molecule`.\n\n**[Gene Ontology Documentation](http://geneontology.org/docs/ontology-documentation/)**\n\n## 1.3 | Principal Component Analysis\n\n<img src = \"https://media.makeameme.org/created/pca-the-cause.jpg\" width = 400>\n\n$Principal$ $Component$ $Analysis$ $PCA$ is a `statistical procedure` that uses an `orthogonal transformation` to convert a set of `correlated variables` into a set of `uncorrelated variables` called principal components.\n\nFirst we compute the Covariance Matrix with the formula \n\n$$Cov_{x , y} = \\frac{\\sum(x_i - x_{mean})(y_i - y_{mean})}{N-1}$$\n\nThen we find the Eign Values and Eign Vectors \n\n## 1.4 | Transformers \n\nA $Transformer$ is a `deep learning architecture` that relies on the `attention mechanism`. It is notable for requiring `less training time`. The model takes in `tokenized` `byte pair encoding` `input` tokens, and at each layer, `contextualizes each token` with other (unmasked) input tokens in `parallel` via attention mechanism.\n\n* $Attention$ $Mechanism$- The attention mechanism allows the transformer to `learn long-range dependencies` between `tokens` in a sequence. \n* $Self-Attention$ - Self-attention is a special type of attention that allows the transformer to `attend to itself`. This means that the transformer can learn relationships between different parts of the same sequence.\n* $Encoder-Decoder$ $Architecture$ - The transformer is an `Encoder-Decoder` architecture. This means that the transformer has two parts\n* * Encoder - takes the input sequence and produces a sequence of hidden states\n* * Decoder - takes these hidden states and produces the output sequence.\n\n$$a = Softmax(\\frac{KQ^T}{\\sqrt{d_k}})$$\n\n<img src = \"https://machinelearningmastery.com/wp-content/uploads/2021/08/attention_research_1.png\" width = 400>\n\n## 1.5 | Categorical Cross Entropy \n\n$Loss$ is a `measure` of `how well a model is performing` on a given task. It is `calculated` by `comparing` the `model's predictions` to the `ground truth labels`. The `lower the loss`, the `better the model` is performing.\n\nHere we will be using `Cross Entropy Loss` , a loss function that `measures` the `difference` between the `model's predicted probability distribution` and the `ground truth distribution`.\n\n## 1.6 | Adaptive Moment Estimation \n\nAn $Optimizer$ is an `algorithm` or function that `updates the weights and biases` of a neural network in order to `minimize a loss function`. \n\nHere we will be using the `Adam Optimizer`\n\n$Adam$ is an $Adaptive$ $Learning$ $Rate$ method, which works by `maintaining two moving averages of the gradients`\n* $Mean$ - Calculate the `momentum term`, which helps to `prevent` the `optimizer` from `getting stuck in local minima`\n* $Variance$ - Calculate the `learning rate`, which is `adjusted based on the magnitude of the gradients`.\n\n$$m_t = \\beta_1 * m_{t - 1} + (1 - \\beta_1) * w_t$$\n\n$$v_t = \\beta_2 * m_{t - 1} + (1 - \\beta_2) * w_t$$\n\n$$m_t = \\frac{m_t}{1 - \\beta_1^t}$$\n\n$$v_t = \\frac{v_t}{1 - \\beta_2^t}$$\n\n$$w_{t+1} = w_t - \\frac{n}{\\sqrt{v_t + e}} * m_t$$\n\n# 2 | Data 📊\n\nThe `goal` of this competition is to `predict the function of a set of proteins`. We will `develop a model trained` on the `amino-acid sequences` of the `proteins and on other data`. Our work `will help researchers` better `understand the function of proteins`, which is `important for discovering` `how cells, tissues, and organs work`.\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd \nimport re\nfrom Bio import SeqIO","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-07-07T03:28:04.05302Z","iopub.execute_input":"2023-07-07T03:28:04.053335Z","iopub.status.idle":"2023-07-07T03:28:04.181157Z","shell.execute_reply.started":"2023-07-07T03:28:04.053312Z","shell.execute_reply":"2023-07-07T03:28:04.180307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The $Training$ $Set$ contains all `proteins with annotated terms` that have been validated by \n* $Experimental$\n* $High-Throughput$ $Evidence$\n* [$Traceable$ $Author$ $Statement$](https://wiki.geneontology.org/index.php/Traceable_Author_Statement_(TAS)#:~:text=The%20TAS%20evidence%20code%20covers,annotations%20come%20from%20review%20articles.)\n* [$Inferred$ $by$ $Curator$ $(IC)$](https://wiki.geneontology.org/Inferred_by_Curator_(IC)) \n\n**Any other sources of Data are allowed**\n\n### 2.1.1.1 | Go-Basic.obo\n\nThe $Ontology$ data is in the `file go-basic.obo`. This file is in $OBO$ `Biology-Oriented Language`. The nodes in `the graph are indexed` by the `term name`\n```\nsubontology_roots = {'BPO':'GO:0008150',\n                     'CCO':'GO:0005575',\n                     'MFO':'GO:0003674'}\n```","metadata":{}},{"cell_type":"code","source":"with open('/kaggle/input/cafa-5-protein-function-prediction/Train/go-basic.obo') as file :\n    \n    content = file.read()\n    stanzas =  re.findall(r'\\[Term\\][\\s\\S]*?(?=\\n\\[|$)' , content)\n    \nprint(stanzas[0])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-07T03:28:11.161166Z","iopub.execute_input":"2023-07-07T03:28:11.161731Z","iopub.status.idle":"2023-07-07T03:28:12.54097Z","shell.execute_reply.started":"2023-07-07T03:28:11.161704Z","shell.execute_reply":"2023-07-07T03:28:12.539951Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.1.1.2 | Training Sequences.fasta\n\nThis file contains only `sequences` for `proteins` with `annotations` in the dataset `labeled proteins`.\n\nThis files are in `FASTA` format. \n\nThis file contains . To obtain the full set of protein sequences for unlabeled proteins, the Swiss-Prot and TrEMBL databases can be found here.","metadata":{}},{"cell_type":"code","source":"fasta_seq = SeqIO.parse(open(\"/kaggle/input/cafa-5-protein-function-prediction/Train/train_sequences.fasta\") ,'fasta')\n\nfasta_seq","metadata":{"execution":{"iopub.status.busy":"2023-07-07T03:28:12.824097Z","iopub.execute_input":"2023-07-07T03:28:12.824412Z","iopub.status.idle":"2023-07-07T03:28:12.851324Z","shell.execute_reply.started":"2023-07-07T03:28:12.824389Z","shell.execute_reply":"2023-07-07T03:28:12.850459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.1.1.3 | Train Terms.tsv\n\nThis file contains the list of annotated terms `ground truth` for the proteins in `train_sequences.fasta`. ","metadata":{}},{"cell_type":"code","source":"pd.read_csv(\"/kaggle/input/cafa-5-protein-function-prediction/Train/train_terms.tsv\" , sep = \"\\t\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-07T03:28:14.918582Z","iopub.execute_input":"2023-07-07T03:28:14.918895Z","iopub.status.idle":"2023-07-07T03:28:17.306396Z","shell.execute_reply.started":"2023-07-07T03:28:14.918871Z","shell.execute_reply":"2023-07-07T03:28:17.30551Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The first column indicates the `protein's UniProt accession ID`, the second is the `GO term ID`, and the third indicates in which `ontology the term appears`.\n\n### 2.1.1.4 | Train Taxonomy.tsv\n\nThis file contains the list of `proteins and the species to which they belong`, represented by a `taxonomic identifier` `taxon ID` number.","metadata":{}},{"cell_type":"code","source":"pd.read_csv(\"/kaggle/input/cafa-5-protein-function-prediction/Train/train_taxonomy.tsv\" , sep = \"\\t\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-07T03:28:17.307953Z","iopub.execute_input":"2023-07-07T03:28:17.308953Z","iopub.status.idle":"2023-07-07T03:28:17.386273Z","shell.execute_reply.started":"2023-07-07T03:28:17.308925Z","shell.execute_reply":"2023-07-07T03:28:17.385277Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2.1.1.5 | IA.txt\n\nIA.txt contains the information accretion (weights) for each GO term. These weights are used to compute weighted precision and recall, as described in the Evaluation section. \n\n## 2.1.2 | Test Set\n\nThe $Test$ $Set$ is `unknown at the beginning` of the competition. It will contain `protein sequences` `their functions` from the `test superset` that `gained experimental annotations` between the `submission-deadline` and the `time of evaluation`.\n\n# 3 | PyTorch DataLoader ⚙️","metadata":{}},{"cell_type":"code","source":"import numpy as np\nfrom transformers import BertModel, BertTokenizer\nimport torch\nfrom torch.utils.data import Dataset","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-07-08T15:15:19.891588Z","iopub.execute_input":"2023-07-08T15:15:19.892026Z","iopub.status.idle":"2023-07-08T15:15:19.897684Z","shell.execute_reply.started":"2023-07-08T15:15:19.891992Z","shell.execute_reply":"2023-07-08T15:15:19.896254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"First we will make a simple class...","metadata":{}},{"cell_type":"code","source":"class Pytorch_Dataset(Dataset):pass","metadata":{"execution":{"iopub.status.busy":"2023-07-08T15:15:26.014425Z","iopub.execute_input":"2023-07-08T15:15:26.014839Z","iopub.status.idle":"2023-07-08T15:15:26.020086Z","shell.execute_reply.started":"2023-07-08T15:15:26.014807Z","shell.execute_reply":"2023-07-08T15:15:26.018926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we will load the `Embedments` for the data \n\nHere we will be using the combination of `T5`/`Prot-BERT`\n\nEmbeds are like representation of a non-numercial elemenet to a list of numerical elemenet. ","metadata":{}},{"cell_type":"code","source":"class Pytorch_Dataset(Dataset):\n    \n    def __init__(self):\n        \n        super(Pytorch_Dataset).__init__()\n        \n        self.embeds = np.load(\"/kaggle/input/t5embeds/test_embeds.npy\")","metadata":{"execution":{"iopub.status.busy":"2023-07-08T15:15:28.499869Z","iopub.execute_input":"2023-07-08T15:15:28.500303Z","iopub.status.idle":"2023-07-08T15:15:28.506584Z","shell.execute_reply.started":"2023-07-08T15:15:28.500269Z","shell.execute_reply":"2023-07-08T15:15:28.50542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In case you wanna see, what these `embeds`looks like, open the below hidden cells\n\nAt this point we will only consider the first $10,000$ values","metadata":{}},{"cell_type":"code","source":"embeds = np.load(\"/kaggle/input/cafa-sample-embeddings/Embeddings/T5-Prot-BERT BFD/Train/Embeds.npy\")\n\nembeds , embeds.shape","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-07-08T15:17:02.083681Z","iopub.execute_input":"2023-07-08T15:17:02.084136Z","iopub.status.idle":"2023-07-08T15:17:02.118637Z","shell.execute_reply.started":"2023-07-08T15:17:02.0841Z","shell.execute_reply":"2023-07-08T15:17:02.117268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For the `training` we need the `targets`, which we will take from the `train_targets_top500.npy`.","metadata":{}},{"cell_type":"code","source":"class Pytorch_Dataset(Dataset):\n    \n    def __init__(self , datatype):\n        \n        super(Pytorch_Dataset).__init__()\n        \n        self.datatype  = datatype\n        \n        self.embeds = np.load(\"/kaggle/input/cafa-sample-embeddings/Embeddings/T5-Prot-BERT BFD/Train/Embeds.npy\")\n        \n        if self.datatype == \"train\" : \n            \n            self.targets = np.load(\"/kaggle/input/train-targets-top500/train_targets_top500.npy\")[:self.embeds.shape[0]]","metadata":{"execution":{"iopub.status.busy":"2023-07-08T15:18:46.590687Z","iopub.execute_input":"2023-07-08T15:18:46.591137Z","iopub.status.idle":"2023-07-08T15:18:46.598955Z","shell.execute_reply.started":"2023-07-08T15:18:46.591105Z","shell.execute_reply":"2023-07-08T15:18:46.597633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These are our targets","metadata":{}},{"cell_type":"code","source":"targets = np.load(\"/kaggle/input/train-targets-top500/train_targets_top500.npy\")[:embeds.shape[0]]\n\ntargets","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-07-08T15:18:53.065437Z","iopub.execute_input":"2023-07-08T15:18:53.065851Z","iopub.status.idle":"2023-07-08T15:18:57.171229Z","shell.execute_reply.started":"2023-07-08T15:18:53.06582Z","shell.execute_reply":"2023-07-08T15:18:57.170063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we will just apply some `Getters` in the class ","metadata":{}},{"cell_type":"code","source":"class Pytorch_Dataset(Dataset):\n\n    def __init__(self , datatype ):\n        super(Pytorch_Dataset).__init__()\n\n        self.datatype = datatype\n        self.embeds = np.load(\"/kaggle/input/cafa-sample-embeddings/Embeddings/T5-Prot-BERT BFD/Train/Embeds.npy\")\n        \n        if datatype == \"train\":\n\n            self.targets = np.load(\"/kaggle/input/train-targets-top500/train_targets_top500.npy\")[:self.embeds.shape[0]]\n            \n    def __len__(self):return self.targets.shape[0]\n\n    def __getitem__(self , index):\n\n        r_embed = torch.tensor(self.embeds[index] , dtype = torch.float32)\n\n        if self.datatype == \"train\":\n\n            r_targets = torch.tensor(self.targets[index], dtype = torch.float32)\n\n            return r_embed, r_targets\n\n        return r_embed","metadata":{"execution":{"iopub.status.busy":"2023-07-08T15:19:58.023453Z","iopub.execute_input":"2023-07-08T15:19:58.023898Z","iopub.status.idle":"2023-07-08T15:19:58.033524Z","shell.execute_reply.started":"2023-07-08T15:19:58.023864Z","shell.execute_reply":"2023-07-08T15:19:58.032164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = Pytorch_Dataset(datatype = \"train\")","metadata":{"execution":{"iopub.status.busy":"2023-07-08T15:20:14.803144Z","iopub.execute_input":"2023-07-08T15:20:14.803596Z","iopub.status.idle":"2023-07-08T15:20:15.15126Z","shell.execute_reply.started":"2023-07-08T15:20:14.803563Z","shell.execute_reply":"2023-07-08T15:20:15.149726Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_dataloader = torch.utils.data.DataLoader(train_data , batch_size = 128 , shuffle = True)","metadata":{"execution":{"iopub.status.busy":"2023-07-08T15:20:22.28532Z","iopub.execute_input":"2023-07-08T15:20:22.285781Z","iopub.status.idle":"2023-07-08T15:20:22.291386Z","shell.execute_reply.started":"2023-07-08T15:20:22.285744Z","shell.execute_reply":"2023-07-08T15:20:22.290469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4 | Model Setup 🤖️ \n\nNow we will train Simple models, and see the results\n\n# 4.1 | Multi Layer Perceptron 🧠\n\nA `MultiLayerPerceptron` is a simple model, which consists of just few perceptrons/nodes/neurons connected with each other. \n\nAs we have choosen for the `top-500` targets, our last layer of the network will contain 500 perceptrons, ","metadata":{}},{"cell_type":"code","source":"class MultiLayerPerceptron(torch.nn.Module):\n\n    def __init__(self):\n      \n        super(MultiLayerPerceptron, self).__init__()\n\n        # l_1 = Linear Layer 1 \n        # a_1 = Activation Layer 1\n\n        self.l_1 = torch.nn.Linear(1024 , 1000)\n        self.a_1 = torch.nn.ReLU()\n\n        self.l_2 = torch.nn.Linear(1000 , 900)\n        self.a_2 = torch.nn.ReLU()\n        \n        self.l_3 = torch.nn.Linear(900 , 800)\n        self.a_3 = torch.nn.ReLU()\n        \n        self.l_4 = torch.nn.Linear(800 , 700)\n        self.a_4 = torch.nn.ReLU()\n        \n        self.l_5 = torch.nn.Linear(700 , 600)\n        self.a_5 = torch.nn.ReLU()\n        \n        self.l_6 = torch.nn.Linear(600 , 500)\n    \n    def forward(self, x):\n        \n        x = self.l_1(x)\n        x = self.a_1(x)\n\n        x = self.l_2(x)\n        x = self.a_2(x)\n        \n        x = self.l_3(x)\n        x = self.a_3(x)\n        \n        x = self.l_4(x)\n        x = self.a_4(x)\n        \n        x = self.l_5(x)\n        x = self.a_5(x)\n        \n        x = self.l_6(x)\n        \n        return x","metadata":{"execution":{"iopub.status.busy":"2023-07-07T03:30:05.732094Z","iopub.execute_input":"2023-07-07T03:30:05.732425Z","iopub.status.idle":"2023-07-07T03:30:05.741455Z","shell.execute_reply.started":"2023-07-07T03:30:05.7324Z","shell.execute_reply":"2023-07-07T03:30:05.740393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I cannot use `Kaggle GPUs` for some reason (I dont know why), so I will be commenting the `cuda` parts ","metadata":{}},{"cell_type":"code","source":"MLP = MultiLayerPerceptron()\n\n# MLP = MultiLayerPerceptron().to(\"cuda\")","metadata":{"execution":{"iopub.status.busy":"2023-07-07T03:30:09.092159Z","iopub.execute_input":"2023-07-07T03:30:09.092477Z","iopub.status.idle":"2023-07-07T03:30:09.128502Z","shell.execute_reply.started":"2023-07-07T03:30:09.092452Z","shell.execute_reply":"2023-07-07T03:30:09.127575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4.2 | Single Layer Perceptron 🧠\n\nA $Single$ $Layer$ $Perceptron$ is a `Neural Network Architechture` that consists of $1$ `Hidden Layer`. According to the `Paper`, we should  be having $1$ `Hidden Layer` with $1,000$ `Perceptrons` in it. ","metadata":{}},{"cell_type":"code","source":"class SingleLayerPerceptron(torch.nn.Module):\n\n    def __init__(self):\n        \n        super(SingleLayerPerceptron, self).__init__()\n\n        self.linear1 = torch.nn.Linear(1024, 1012)\n        self.activation1 = torch.nn.ReLU()\n        \n        self.linear2 = torch.nn.Linear(1012, 500)\n\n    def forward(self, x):\n        \n        x = self.linear1(x)\n        x = self.activation1(x)\n        \n        x = self.linear2(x)\n        \n        return x","metadata":{"execution":{"iopub.status.busy":"2023-07-07T03:33:50.766326Z","iopub.execute_input":"2023-07-07T03:33:50.76666Z","iopub.status.idle":"2023-07-07T03:33:50.773566Z","shell.execute_reply.started":"2023-07-07T03:33:50.766637Z","shell.execute_reply":"2023-07-07T03:33:50.77235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"SLP = SingleLayerPerceptron()\n\n# SLP = SingleLayerPerceptron().to(\"cuda\")","metadata":{"execution":{"iopub.status.busy":"2023-07-07T03:33:51.615352Z","iopub.execute_input":"2023-07-07T03:33:51.615698Z","iopub.status.idle":"2023-07-07T03:33:51.632262Z","shell.execute_reply.started":"2023-07-07T03:33:51.615676Z","shell.execute_reply":"2023-07-07T03:33:51.631352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Loss Function 📉\n\nHere we will be using `Categorical Cross Entropy`","metadata":{}},{"cell_type":"code","source":"CrossEntropy = torch.nn.CrossEntropyLoss()","metadata":{"execution":{"iopub.status.busy":"2023-07-04T07:33:31.557407Z","iopub.execute_input":"2023-07-04T07:33:31.557861Z","iopub.status.idle":"2023-07-04T07:33:31.563777Z","shell.execute_reply.started":"2023-07-04T07:33:31.557829Z","shell.execute_reply":"2023-07-04T07:33:31.562381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Optimizer 💡\n\nHere we will be using `Adam`","metadata":{}},{"cell_type":"code","source":"optimizer = torch.optim.Adam(MLP.parameters(), lr = 0.0005)","metadata":{"execution":{"iopub.status.busy":"2023-07-04T07:33:32.041637Z","iopub.execute_input":"2023-07-04T07:33:32.042017Z","iopub.status.idle":"2023-07-04T07:33:32.04761Z","shell.execute_reply.started":"2023-07-04T07:33:32.04199Z","shell.execute_reply":"2023-07-04T07:33:32.0464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5 | Training Loop 🔁 ","metadata":{}},{"cell_type":"code","source":"from IPython.display import IFrame","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-07T03:45:17.409084Z","iopub.execute_input":"2023-07-07T03:45:17.409382Z","iopub.status.idle":"2023-07-07T03:45:17.41388Z","shell.execute_reply.started":"2023-07-07T03:45:17.409359Z","shell.execute_reply":"2023-07-07T03:45:17.412739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5.1 | Multi Layer Perceptron Training 🔁\n\nNow we will start the training loop \n\nI dont the exact reason, but everytime I try to access `GPU` for some training in `Kaggle`. `CUDA goes out of memory`. Thus I have trained the model on `Colab` and will imported the results to `Wandb`.\n\n```\nlos = []\n\nfor epochs in (range(5)):\n\n    losses = []\n\n    for x , y in tqdm.tqdm(train_dataloader):\n        \n        x = torch.tensor(x , dtype = torch.float32)\n        y = torch.tensor(x , dtype = troch.float32)\n        \n        \n#         x = torch.tensor(x , dtype = torch.float32).to(\"cuda\")\n#         y = torch.tensor(y , dtype = torch.float32).to(\"cuda\")\n\n        optimizer.zero_grad()\n        preds = MLP(x)\n\n        loss = CrossEntropy(preds, y)\n        losses.append(loss)\n\n    los.append(losses)\n```","metadata":{}},{"cell_type":"code","source":"IFrame(\"https://wandb.ai//ayushsinghal659/CAFA%7CMLP/reports/CAFA-MLP--Vmlldzo0Nzk1NjQx\" , 1300 , 400)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-04T07:34:33.838854Z","iopub.execute_input":"2023-07-04T07:34:33.839334Z","iopub.status.idle":"2023-07-04T07:34:33.848993Z","shell.execute_reply.started":"2023-07-04T07:34:33.839298Z","shell.execute_reply":"2023-07-04T07:34:33.847838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see we did not get good results, but we will try to omprove our results, by improving the model and by adding new one \n\n# 5.2 | Single Layer Perceptron Training 🔁\n\nNow we will train our `Single Layer Perceptron`\n\n```\n# torch.cuda.empty_cache()\n\nlos = []\n\nfor epochs in (range(4)):\n\n    losses = []\n    \n    # torch.cuda.empty_cache()\n    \n    for x , y in tqdm.tqdm(train_dataloader , total = 10000):\n\n        # torch.cuda.empty_cache()\n\n        x = torch.tensor(x , dtype = torch.float32)\n        y = torch.tensor(y , dtype = torch.float32)\n        \n        # x = torch.tensor(x , dtype = torch.float32).to(\"cuda\")\n        # y = torch.tensor(y , dtype = torch.float32).to(\"cuda\")\n\n        optimizer.zero_grad()\n        preds = SLP(x)\n\n        loss = CrossEntropy(preds, y)\n        losses.append(loss)\n\n    los.append(losses)\n```","metadata":{}},{"cell_type":"code","source":"IFrame(\"https://wandb.ai//ayushsinghal659/uncategorized/reports/CAFA-TEMPROT--Vmlldzo0ODIxNzI3\" , 1300 , 400)","metadata":{"execution":{"iopub.status.busy":"2023-07-07T03:45:21.710379Z","iopub.execute_input":"2023-07-07T03:45:21.711432Z","iopub.status.idle":"2023-07-07T03:45:21.717997Z","shell.execute_reply.started":"2023-07-07T03:45:21.711386Z","shell.execute_reply":"2023-07-07T03:45:21.71694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6 | TO DO LIST 📄\n\n```\nTO DO 1 : VISUALIZE THE DATA\n\nTO DO 2 : TRAIN A MODEL\n\nTO DO 3 : TRY DIFFERENT MODELS\n\nTO DO 4 : ADD WANDB SUPPORT\n\nTO DO 5 : ADD TENSORFLOW DATA LOADER\n\nTO DO 6 : TRAIN A TF MODEL\n\nTO DO 7 : IMPROVE RESULTS\n\nTO DO 8 : DECREASE TRAINING TIME\n\nTO DO 9 : DANCE \n```\n\n# 7 | Ending 🏁\n\n**THAT'S IT FOR TODAY GUYS**\n\n**WE WILL GO DEEPER INTO THE DATA IN THE UPCOMING VERSIONS**\n\n**PLEASE COMMENT YOUR THOUGHTS, HIHGLY APPRICIATED**\n\n**DONT FORGET TO MAKE AN UPVOTE, IF YOU LIKED MY WORK $:)$**\n\n<img src = \"https://i.imgflip.com/19aadg.jpg\">\n\n**PEACE OUT $:)$**","metadata":{}}]}