{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.11.13"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":41875,"databundleVersionId":5521661,"sourceType":"competition"},{"sourceId":116062,"databundleVersionId":14084779,"sourceType":"competition"}],"dockerImageVersionId":31153,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false},"papermill":{"default_parameters":{},"duration":69.341703,"end_time":"2025-11-09T06:03:58.48072","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2025-11-09T06:02:49.139017","version":"2.6.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Credits\nThis analysis is built upon the foundational work by @leonidkulyk from the CAFA 5 competition, who published the acclaimed notebook titled \"[EDA] 🧬CAFA5-PFP ~ 🕸️Interactive DAGs | 📊Plotly.\" I want to extend my gratitude for his insightful contribution.\n\nIn this notebook, I have adapted and utilized his original code to analyze this year's CAFA competition data.","metadata":{"papermill":{"duration":0.012774,"end_time":"2025-11-09T06:02:54.327468","exception":false,"start_time":"2025-11-09T06:02:54.314694","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<!-- Dark Mode Header -->\n<div style=\"background-color:#1e1e1e; padding: 30px; border-radius: 12px; text-align:center;\">\n    <h1 style=\"font-family: 'Consolas', monospace; font-size: 32px; font-weight: bold; color:#f5f5f5; margin-bottom:10px;\">\n        🧬 CAFA 5 Protein Function Prediction \n    </h1>\n    <p style=\"color:#949494; font-family: 'Consolas', monospace; font-size: 20px; text-align:center; margin-top:0;\">\n        Predict the biological function of a protein\n    </p>\n</div>\n<hr style=\"border:1px solid #333; margin:20px 0;\">\n\n","metadata":{"papermill":{"duration":0.009833,"end_time":"2025-11-09T06:02:54.348506","exception":false,"start_time":"2025-11-09T06:02:54.338673","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<!-- Dark Mode Overview Section -->\n<div style=\"background-color:#1e1e1e; padding: 30px; border-radius: 12px; color:#f5f5f5; font-family: 'Consolas', monospace;\">\n\n  <h2 style=\"font-size:32px; font-weight:bold; text-align:center; margin-bottom:20px;\">\n    (ಠಿ_ಠ) Overview\n  </h2>\n\n  <p style=\"font-size:16px; margin-bottom:12px;\">\n    🔴 <b>Goal</b>: Predict the function of proteins based on their amino acid sequences and other data.\n  </p>\n\n  <p style=\"font-size:16px; margin-bottom:12px;\">\n    ⚪ <b>Importance</b>: Accurate protein function assignment is <span style=\"color:#ffa500;\">crucial for understanding molecular biology</span>, discovering cellular mechanisms, and developing new therapies.\n  </p>\n\n  <p style=\"font-size:16px; margin-bottom:12px;\">\n    ⚪ <b>Context</b>: Proteins are composed of <span style=\"color:#00bfff;\">20 types of amino acids</span>. Humans have tens of thousands of proteins, each a chain of amino acids forming unique structures.\n  </p>\n\n  <p style=\"font-size:16px; margin-bottom:12px;\">\n    ⚪ <b>Challenges</b>: Proteins often have multiple functions and interact with many partners. Predicting functions is complex due to ambiguity, data integration, and structural diversity.\n  </p>\n\n  <p style=\"font-size:16px; margin-bottom:12px;\">\n    ⚪ <b>Host</b>: Function-COSI organizes the competition, bringing together computational biologists, experimental biologists, and biocurators.\n  </p>\n\n  <p style=\"font-size:16px; margin-bottom:12px;\">\n    ⚪ <b>Co-organizers</b>: Iowa State University, Northeastern University, University of Padova, and UniProt.\n  </p>\n\n  <p style=\"font-size:16px; margin-bottom:0;\">\n    ⚪ <b>Acknowledgments</b>: Supported by the organizers above and the International Society for Computational Biology.\n  </p>\n\n</div>\n<hr style=\"border:1px solid #333; margin:25px 0;\">\n","metadata":{"papermill":{"duration":0.009932,"end_time":"2025-11-09T06:02:54.368747","exception":false,"start_time":"2025-11-09T06:02:54.358815","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<!-- Dark Mode Table of Contents -->\n<a id=\"top\"></a>\n\n<div style=\"background-color:#1e1e1e; padding:25px 20px; border-radius:15px; text-align:center; font-family:'Consolas', monospace; color:#3c79f5; font-size:28px; font-weight:bold; box-shadow: 0 0 10px rgba(60,121,245,0.6); margin-bottom:20px;\">\n    Table of Contents\n</div>\n\n<div style=\"background-color:#2a2a2a; padding:25px 30px; border-radius:12px; font-family:'Consolas', monospace; font-size:16px; color:#f5f5f5; line-height:1.8;\">\n    <ul style=\"list-style-type:none; padding-left:0;\">\n        <li>🔹 <a href=\"#iid\" style=\"color:#61dafb; text-decoration:none;\">Install & Import & Define</a></li>\n        <li>🔹 <a href=\"#1\" style=\"color:#61dafb; text-decoration:none;\">1. Data overview</a></li>\n        <li>🔹 <a href=\"#2\" style=\"color:#61dafb; text-decoration:none;\">2. Training Set</a>\n            <ul style=\"list-style-type:none; padding-left:20px; margin-top:5px;\">\n                <li>⚪ <a href=\"#2.1\" style=\"color:#61dafb; text-decoration:none;\">2.1 Gene Ontology</a></li>\n                <li>⚪ <a href=\"#2.2\" style=\"color:#61dafb; text-decoration:none;\">2.2 Training sequences</a></li>\n                <li>⚪ <a href=\"#2.3\" style=\"color:#61dafb; text-decoration:none;\">2.3 Labels</a></li>\n                <li>⚪ <a href=\"#2.4\" style=\"color:#61dafb; text-decoration:none;\">2.4 Taxonomy</a></li>\n                <li>⚪ <a href=\"#2.5\" style=\"color:#61dafb; text-decoration:none;\">2.5 Information accretion</a></li>\n            </ul>\n        </li>\n    </ul>\n</div>\n","metadata":{"papermill":{"duration":0.010133,"end_time":"2025-11-09T06:02:54.389378","exception":false,"start_time":"2025-11-09T06:02:54.379245","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"\n<div style=\"\n    background-color:#1e1e1e; \n    color:#00ffff; \n    padding:15px 25px; \n    border-radius:12px; \n    font-family:'Consolas', monospace; \n    font-size:18px; \n    font-weight:bold; \n    text-align:center;\n    box-shadow: 0 0 10px rgba(0,255,255,0.3);\n\">\n🛠 Install & Import & Define\n</div>\n\n","metadata":{"papermill":{"duration":0.009886,"end_time":"2025-11-09T06:02:54.410057","exception":false,"start_time":"2025-11-09T06:02:54.400171","status":"completed"},"tags":[]}},{"cell_type":"code","source":"!pip install obonet -q\n!pip install pyvis -q","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:02:36.477934Z","iopub.execute_input":"2025-11-18T09:02:36.478205Z","iopub.status.idle":"2025-11-18T09:02:45.944121Z","shell.execute_reply.started":"2025-11-18T09:02:36.478175Z","shell.execute_reply":"2025-11-18T09:02:45.943079Z"},"papermill":{"duration":10.251619,"end_time":"2025-11-09T06:03:04.672258","exception":false,"start_time":"2025-11-09T06:02:54.420639","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!pip install Bio","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:02:45.945376Z","iopub.execute_input":"2025-11-18T09:02:45.94571Z","iopub.status.idle":"2025-11-18T09:02:52.245869Z","shell.execute_reply.started":"2025-11-18T09:02:45.945683Z","shell.execute_reply":"2025-11-18T09:02:52.244769Z"},"papermill":{"duration":6.644757,"end_time":"2025-11-09T06:03:11.327947","exception":false,"start_time":"2025-11-09T06:03:04.68319","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport json\nfrom PIL import Image\nfrom typing import Dict\nfrom collections import Counter\n\nimport random\nimport cv2\nimport obonet\nimport networkx\nimport pandas as pd\nimport numpy as np\nimport plotly.express as px\nimport plotly.graph_objects as go\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as mpatch\nfrom Bio import SeqIO\nfrom pyvis.network import Network\nimport plotly.io as pio\npio.renderers.default = 'iframe_connected'","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:02:52.247935Z","iopub.execute_input":"2025-11-18T09:02:52.248233Z","iopub.status.idle":"2025-11-18T09:02:55.920937Z","shell.execute_reply.started":"2025-11-18T09:02:52.248206Z","shell.execute_reply":"2025-11-18T09:02:55.920161Z"},"papermill":{"duration":6.791093,"end_time":"2025-11-09T06:03:18.13618","exception":false,"start_time":"2025-11-09T06:03:11.345087","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"\n    background-color:#1e1e1e; \n    color:#00ffff; \n    padding:15px 25px; \n    border-radius:12px; \n    font-family:'Consolas', monospace; \n    font-size:18px; \n    font-weight:bold; \n    text-align:center;\n    box-shadow: 0 0 10px rgba(0,255,255,0.3);\n\">\n🛠 Define config\n</div>\n","metadata":{"papermill":{"duration":0.011802,"end_time":"2025-11-09T06:03:18.159889","exception":false,"start_time":"2025-11-09T06:03:18.148087","status":"completed"},"tags":[]}},{"cell_type":"code","source":"class CFG:\n    train_go_obo_path: str = \"/kaggle/input/cafa-5-protein-function-prediction/Train/go-basic.obo\"\n    train_seq_fasta_path: str = \"/kaggle/input/cafa-5-protein-function-prediction/Train/train_sequences.fasta\"\n    train_terms_path: str = \"/kaggle/input/cafa-5-protein-function-prediction/Train/train_terms.tsv\"\n    train_taxonomy_path: str = \"/kaggle/input/cafa-5-protein-function-prediction/Train/train_taxonomy.tsv\"\n    train_ia_path: str = \"/kaggle/input/cafa-5-protein-function-prediction/IA.txt\"","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:02:55.921822Z","iopub.execute_input":"2025-11-18T09:02:55.922108Z","iopub.status.idle":"2025-11-18T09:02:55.927233Z","shell.execute_reply.started":"2025-11-18T09:02:55.92208Z","shell.execute_reply":"2025-11-18T09:02:55.926386Z"},"papermill":{"duration":0.019915,"end_time":"2025-11-09T06:03:18.191323","exception":false,"start_time":"2025-11-09T06:03:18.171408","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"\n    background-color:#1e1e1e; \n    color:#00ffff; \n    padding:15px 25px; \n    border-radius:12px; \n    font-family:'Consolas', monospace; \n    font-size:18px; \n    font-weight:bold; \n    text-align:center;\n    box-shadow: 0 0 10px rgba(0,255,255,0.3);\n\">\n🛠 Define utilization methods\n</div>\n","metadata":{"papermill":{"duration":0.011486,"end_time":"2025-11-09T06:03:18.214798","exception":false,"start_time":"2025-11-09T06:03:18.203312","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def plot_dag(graph, term, radius=1):\n    # create smaller subgraph\n    # radius - include all neighbors of distance<=radius from n (increse it to add further parent's branches).\n    ng_graph = networkx.ego_graph(graph, term, radius=radius)\n\n    for n in ng_graph.nodes(data=True):\n        # concatenate label of the node with its attribute\n        n[1][\"label\"] = n[0] + \" \" +n[1][\"name\"]\n\n    nt = Network(directed=True, notebook=True, cdn_resources=\"in_line\")\n    nt.from_nx(ng_graph)\n    return nt.show(\"network.html\")","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:02:55.92812Z","iopub.execute_input":"2025-11-18T09:02:55.92841Z","iopub.status.idle":"2025-11-18T09:02:55.950283Z","shell.execute_reply.started":"2025-11-18T09:02:55.928384Z","shell.execute_reply":"2025-11-18T09:02:55.949397Z"},"papermill":{"duration":0.022869,"end_time":"2025-11-09T06:03:18.249654","exception":false,"start_time":"2025-11-09T06:03:18.226785","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<a id=\"1\"></a>\n<div style=\"\n    background-color:#1e1e1e;\n    color:#61dafb;\n    padding:20px;\n    font-size:32px;\n    font-family:'Consolas', monospace;\n    text-align:center;\n    border-radius:15px;\n    box-shadow: 0 0 10px rgba(97,218,251,0.5) inset;\n\">\n<b>1. Data overview</b>\n</div>\n","metadata":{"papermill":{"duration":0.011352,"end_time":"2025-11-09T06:03:18.273185","exception":false,"start_time":"2025-11-09T06:03:18.261833","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<div style=\"background-color:#1e1e1e; color:#e0e0e0; padding:25px; border-radius:12px; font-family:'Consolas', monospace; font-size:16px; line-height:1.6;\">\n\n🔴 **The [Gene Ontology (GO)](http://geneontology.org/docs/ontology-documentation/)** is a concept hierarchy that describes the biological function of genes and gene products at different levels of abstraction (Ashburner et al., 2000). It is a good model to describe the multi-faceted nature of protein function.  \n\n⚪ GO is a `directed acyclic graph`. The nodes in this graph are functional descriptors (terms or classes) connected by relational ties (is_a, part_of, etc.). For example, terms *protein binding activity* and *binding activity* are related by an is_a relationship; however, the edge is often reversed to point from binding towards protein binding. This graph contains **three subgraphs (subontologies)**: **Molecular Function (MF), Biological Process (BP), and Cellular Component (CC)**. Each represents a different aspect of a protein's function: what it does on a molecular level (MF), which biological processes it participates in (BP), and where in the cell it is located (CC).  \n\n⚪ The protein's function is therefore represented by a subset of one or more subontologies. Annotations are supported by evidence codes, divided into experimental (from research papers) and non-experimental (inferred computationally). Read more about [GO evidence codes](http://geneontology.org/docs/guide-go-evidence-codes/).  \n\n🔴 In this competition, **experimentally determined term-protein assignments** are used as class labels. That is, if a protein is labeled with a term, it means this protein has this function validated by experimental evidence. By processing these annotated terms, we can generate a dataset of proteins and their ground truth labels. The absence of a term does not mean the protein lacks the function—just that it hasn’t been annotated yet. Proteins may have annotations from multiple subontologies.  \n\n</div>\n","metadata":{"papermill":{"duration":0.011254,"end_time":"2025-11-09T06:03:18.296383","exception":false,"start_time":"2025-11-09T06:03:18.285129","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<div style=\"background-color:#1e1e1e; color:#f0f0f0; padding:25px; border-radius:12px; \n            font-family:'Consolas', monospace; font-size:28px; text-align:center; \n            box-shadow: rgba(0, 0, 0, 0.2) 0px 2px 6px inset;\">\n    <b>2. Training Set</b>\n</div>\n","metadata":{"papermill":{"duration":0.011451,"end_time":"2025-11-09T06:03:18.319366","exception":false,"start_time":"2025-11-09T06:03:18.307915","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<p style=\"font-family: 'Consolas', monospace; font-size:16px; color:#e0e0e0; line-height:1.6;\">\n⚪ For the <i>training set</i>, we include all proteins with annotated terms that have been validated by experimental or high-throughput evidence, traceable author statement (<code>TAS</code>), or inferred by curator (<code>IC</code>). We use annotations from the UniProtKB release of 2022-11-17. Participants are not required to use these data and are also welcome to use any other data available to them.\n</p>\n\n<p style=\"font-family: 'Consolas', monospace; font-size:16px; color:#e0e0e0; line-height:1.6;\">\n⚪ For the <i>training set</i>, we include all proteins with annotated terms that have been validated by experimental or high-throughput evidence, traceable author statement (<code>TAS</code>), or inferred by curator (<code>IC</code>). We use annotations from the UniProtKB release of 2022-11-17. Participants are not required to use these data and are also welcome to use any other data available to them.\n</p>\n","metadata":{"papermill":{"duration":0.011304,"end_time":"2025-11-09T06:03:18.342712","exception":false,"start_time":"2025-11-09T06:03:18.331408","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<p style=\"font-family: 'Consolas', monospace; font-size:16px; color:#e0e0e0; line-height:1.6;\">\n❔ Let's consider each <i>training file</i> iteratively.\n</p>\n","metadata":{"papermill":{"duration":0.011196,"end_time":"2025-11-09T06:03:18.365677","exception":false,"start_time":"2025-11-09T06:03:18.354481","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<div style=\"background-color:#1e1e1e; color:#f0f0f0; padding:25px; border-radius:12px; \n            font-family:'Consolas', monospace; font-size:28px; text-align:center; \n            box-shadow: rgba(0, 0, 0, 0.2) 0px 2px 6px inset;\">\n    <b>2.1 Gene Ontology </b>\n</div>\n","metadata":{"papermill":{"duration":0.011377,"end_time":"2025-11-09T06:03:18.388505","exception":false,"start_time":"2025-11-09T06:03:18.377128","status":"completed"},"tags":[]}},{"cell_type":"code","source":"%%time\ngraph = obonet.read_obo(CFG.train_go_obo_path)","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:02:55.951131Z","iopub.execute_input":"2025-11-18T09:02:55.95144Z","iopub.status.idle":"2025-11-18T09:03:02.50191Z","shell.execute_reply.started":"2025-11-18T09:02:55.951409Z","shell.execute_reply":"2025-11-18T09:03:02.50119Z"},"papermill":{"duration":7.043599,"end_time":"2025-11-09T06:03:25.443546","exception":false,"start_time":"2025-11-09T06:03:18.399947","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">❔ Number of nodes.</p>","metadata":{"papermill":{"duration":0.012524,"end_time":"2025-11-09T06:03:25.467849","exception":false,"start_time":"2025-11-09T06:03:25.455325","status":"completed"},"tags":[]}},{"cell_type":"code","source":"print(f\"Number of nodes: {len(graph)}\")","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:02.502643Z","iopub.execute_input":"2025-11-18T09:03:02.50285Z","iopub.status.idle":"2025-11-18T09:03:02.507715Z","shell.execute_reply.started":"2025-11-18T09:03:02.502832Z","shell.execute_reply":"2025-11-18T09:03:02.506738Z"},"papermill":{"duration":0.019235,"end_time":"2025-11-09T06:03:25.49884","exception":false,"start_time":"2025-11-09T06:03:25.479605","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">❔ Number of edges.</p>","metadata":{"papermill":{"duration":0.013058,"end_time":"2025-11-09T06:03:25.524008","exception":false,"start_time":"2025-11-09T06:03:25.51095","status":"completed"},"tags":[]}},{"cell_type":"code","source":"print(f\"Number of edges: {graph.number_of_edges()}\")","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:02.508621Z","iopub.execute_input":"2025-11-18T09:03:02.50884Z","iopub.status.idle":"2025-11-18T09:03:02.60677Z","shell.execute_reply.started":"2025-11-18T09:03:02.508821Z","shell.execute_reply":"2025-11-18T09:03:02.605954Z"},"papermill":{"duration":0.100549,"end_time":"2025-11-09T06:03:25.636888","exception":false,"start_time":"2025-11-09T06:03:25.536339","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">🔴 To display a graph, you need to focus on a specific term, let's take term <code>GO:0034655</code> as an example.</p> ","metadata":{"papermill":{"duration":0.011778,"end_time":"2025-11-09T06:03:25.66101","exception":false,"start_time":"2025-11-09T06:03:25.649232","status":"completed"},"tags":[]}},{"cell_type":"code","source":"term = \"GO:0034655\"","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:02.609856Z","iopub.execute_input":"2025-11-18T09:03:02.610122Z","iopub.status.idle":"2025-11-18T09:03:02.614289Z","shell.execute_reply.started":"2025-11-18T09:03:02.610098Z","shell.execute_reply":"2025-11-18T09:03:02.613372Z"},"papermill":{"duration":0.019801,"end_time":"2025-11-09T06:03:25.692982","exception":false,"start_time":"2025-11-09T06:03:25.673181","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">❔ Lookup <code>nucleobase-containing compound catabolic process</code> node properties (term GO:0034655).</p>","metadata":{"papermill":{"duration":0.011584,"end_time":"2025-11-09T06:03:25.716668","exception":false,"start_time":"2025-11-09T06:03:25.705084","status":"completed"},"tags":[]}},{"cell_type":"code","source":"graph.nodes[term]","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:02.615209Z","iopub.execute_input":"2025-11-18T09:03:02.615475Z","iopub.status.idle":"2025-11-18T09:03:02.634935Z","shell.execute_reply.started":"2025-11-18T09:03:02.61545Z","shell.execute_reply":"2025-11-18T09:03:02.633957Z"},"papermill":{"duration":0.021819,"end_time":"2025-11-09T06:03:25.750326","exception":false,"start_time":"2025-11-09T06:03:25.728507","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">❔ Let's plot DAG for the term GO:0034655.</p>","metadata":{"papermill":{"duration":0.011516,"end_time":"2025-11-09T06:03:25.77395","exception":false,"start_time":"2025-11-09T06:03:25.762434","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">❔ Let's plot DAG for the term GO:0034655.</p>","metadata":{"papermill":{"duration":0.011388,"end_time":"2025-11-09T06:03:25.797426","exception":false,"start_time":"2025-11-09T06:03:25.786038","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def plot_dag(graph, term, radius=1):\n    \"\"\"\n    Plots a DAG subgraph around the specified term using pyvis.\n    Saves a separate HTML file for each term and radius.\n    \"\"\"\n    # Create smaller subgraph\n    ng_graph = networkx.ego_graph(graph, term, radius=radius, undirected=False)\n\n    # Prepare node labels\n    for n in ng_graph.nodes(data=True):\n        n[1][\"label\"] = n[0] + \" \" + n[1].get(\"name\", \"\")\n\n    # Initialize pyvis network\n    nt = Network(\n        height=\"800px\",\n        width=\"100%\",\n        directed=True,\n        notebook=True,\n        bgcolor=\"#1e1e1e\",\n        font_color=\"white\",\n        cdn_resources=\"in_line\"\n    )\n    nt.from_nx(ng_graph)\n\n    # Optional styling\n    nt.set_options(\"\"\"\n    var options = {\n      \"nodes\": {\"color\":{\"background\":\"#1f78b4\",\"border\":\"#ffffff\"},\"font\":{\"color\":\"white\",\"size\":14,\"face\":\"Consolas\"}},\n      \"edges\": {\"color\":\"white\",\"arrows\":{\"to\":{\"enabled\":true}}},\n      \"physics\": {\"enabled\":true,\"stabilization\":{\"iterations\":200}}\n    }\n    \"\"\")\n\n    # Dynamic filename\n    filename = f\"network_{term}_radius{radius}.html\"\n    return nt.show(filename)\n","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:02.635812Z","iopub.execute_input":"2025-11-18T09:03:02.636122Z","iopub.status.idle":"2025-11-18T09:03:02.653405Z","shell.execute_reply.started":"2025-11-18T09:03:02.636094Z","shell.execute_reply":"2025-11-18T09:03:02.652651Z"},"papermill":{"duration":0.022132,"end_time":"2025-11-09T06:03:25.83138","exception":false,"start_time":"2025-11-09T06:03:25.809248","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_dag(graph, term, radius= 1)","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:02.654159Z","iopub.execute_input":"2025-11-18T09:03:02.654408Z","iopub.status.idle":"2025-11-18T09:03:02.800593Z","shell.execute_reply.started":"2025-11-18T09:03:02.654388Z","shell.execute_reply":"2025-11-18T09:03:02.799449Z"},"papermill":{"duration":0.14317,"end_time":"2025-11-09T06:03:25.986892","exception":false,"start_time":"2025-11-09T06:03:25.843722","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">❔ Now let's look how's full DAG looks like for the selected term. To do that just increase value of the radius parameter. <code>radius</code> - responsible include all neighbors of distance ≤ radius from n.</p>","metadata":{"papermill":{"duration":0.011882,"end_time":"2025-11-09T06:03:26.010972","exception":false,"start_time":"2025-11-09T06:03:25.99909","status":"completed"},"tags":[]}},{"cell_type":"code","source":"plot_dag(graph, term, radius=50000)","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:02.801513Z","iopub.execute_input":"2025-11-18T09:03:02.801761Z","iopub.status.idle":"2025-11-18T09:03:02.926433Z","shell.execute_reply.started":"2025-11-18T09:03:02.801741Z","shell.execute_reply":"2025-11-18T09:03:02.925513Z"},"papermill":{"duration":0.139039,"end_time":"2025-11-09T06:03:26.163006","exception":false,"start_time":"2025-11-09T06:03:26.023967","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">⚪ In the first graph, only the nodes are connected by relational ties between them, that is, which are located in <code>is_a</code>. But the second one shows a complete graph in which you can see how the connection will look with a non-peripheral relationship between all the nodes in the graph.</p> ","metadata":{"papermill":{"duration":0.011551,"end_time":"2025-11-09T06:03:26.186911","exception":false,"start_time":"2025-11-09T06:03:26.17536","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">🔴 Let's look at another term. The term <code>GO:00048420</code> taken as an example.</p> ","metadata":{"papermill":{"duration":0.011699,"end_time":"2025-11-09T06:03:26.210523","exception":false,"start_time":"2025-11-09T06:03:26.198824","status":"completed"},"tags":[]}},{"cell_type":"code","source":"term = \"GO:0004842\"","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:02.927376Z","iopub.execute_input":"2025-11-18T09:03:02.927812Z","iopub.status.idle":"2025-11-18T09:03:02.932036Z","shell.execute_reply.started":"2025-11-18T09:03:02.927788Z","shell.execute_reply":"2025-11-18T09:03:02.931073Z"},"papermill":{"duration":0.01928,"end_time":"2025-11-09T06:03:26.24171","exception":false,"start_time":"2025-11-09T06:03:26.22243","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">❔ Lookup <code>cellular nitrogen compound catabolic process</code> node properties (term GO:0044270).</p>","metadata":{"papermill":{"duration":0.014383,"end_time":"2025-11-09T06:03:26.268175","exception":false,"start_time":"2025-11-09T06:03:26.253792","status":"completed"},"tags":[]}},{"cell_type":"code","source":"graph.nodes[term]","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:02.933042Z","iopub.execute_input":"2025-11-18T09:03:02.933467Z","iopub.status.idle":"2025-11-18T09:03:02.949859Z","shell.execute_reply.started":"2025-11-18T09:03:02.933445Z","shell.execute_reply":"2025-11-18T09:03:02.949128Z"},"papermill":{"duration":0.022433,"end_time":"2025-11-09T06:03:26.302854","exception":false,"start_time":"2025-11-09T06:03:26.280421","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_dag(graph, term, radius=3)","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:02.950728Z","iopub.execute_input":"2025-11-18T09:03:02.951204Z","iopub.status.idle":"2025-11-18T09:03:03.081709Z","shell.execute_reply.started":"2025-11-18T09:03:02.95118Z","shell.execute_reply":"2025-11-18T09:03:03.080846Z"},"papermill":{"duration":0.137509,"end_time":"2025-11-09T06:03:26.454389","exception":false,"start_time":"2025-11-09T06:03:26.31688","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">❔ Now let's look how's full DAG looks like for the selected term. To do that just increase value of the radius parameter. <code>radius</code> - responsible include all neighbors of distance ≤ radius from n.</p>","metadata":{"papermill":{"duration":0.012264,"end_time":"2025-11-09T06:03:26.479896","exception":false,"start_time":"2025-11-09T06:03:26.467632","status":"completed"},"tags":[]}},{"cell_type":"code","source":"plot_dag(graph, term, radius=1000)","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:03.082619Z","iopub.execute_input":"2025-11-18T09:03:03.082909Z","iopub.status.idle":"2025-11-18T09:03:03.202719Z","shell.execute_reply.started":"2025-11-18T09:03:03.082886Z","shell.execute_reply":"2025-11-18T09:03:03.201825Z"},"papermill":{"duration":0.135645,"end_time":"2025-11-09T06:03:26.628261","exception":false,"start_time":"2025-11-09T06:03:26.492616","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"background-color:#1e1e1e; color:#f0f0f0; padding:25px; border-radius:12px; \n            font-family:'Consolas', monospace; font-size:28px; text-align:center; \n            box-shadow: rgba(0, 0, 0, 0.2) 0px 2px 6px inset;\">\n    <b>2.2 Training Sequences </b>\n</div>\n","metadata":{"papermill":{"duration":0.012452,"end_time":"2025-11-09T06:03:26.654011","exception":false,"start_time":"2025-11-09T06:03:26.641559","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<div style=\"background-color:#1e1e1e; \n            color:#c5c5c5; \n            font-family:'Consolas', monospace; \n            font-size:16px; \n            line-height:1.6; \n            padding:20px; \n            border-radius:12px;\n            box-shadow: rgba(0,0,0,0.5) 0px 2px 6px inset;\">\n\n<p>⚪ <code>Training sequences</code>: <b>train_sequences.fasta</b> contains the protein sequences for the training dataset.</p>\n\n<p>⚪ These files are in <a href=\"https://en.wikipedia.org/wiki/FASTA_format\" style=\"color:#61dafb;\"><strong>FASTA format</strong></a>, a standard format for describing protein sequences. The proteins were all retrieved from the <a href=\"https://www.uniprot.org/\" style=\"color:#61dafb;\"><strong>UniProt dataset</strong></a> curated at the European Bioinformatics Institute.</p>\n\n<p>⚪ The header contains the protein's UniProt accession ID and additional information about the protein. Most protein sequences were extracted from the Swiss-Prot database, but a subset of proteins not represented in Swiss-Prot were extracted from the TrEMBL database. In both cases, sequences come from the 2022_05 release (14-Dec-2022). More info <a href=\"https://www.uniprot.org/help/uniprotkb_sections\" style=\"color:#61dafb;\"><strong>here</strong></a>.</p>\n\n<p>⚪ The <code>train_sequences.fasta</code> file indicates the database source. For example, <code>sp|P9WHI7|RECN_MYCT</code> in the FASTA header indicates a Swiss-Prot protein (<code>sp</code>) with UniProt ID <code>P9WHI7</code> and gene name <code>RECN_MYCT</code>. TrEMBL sequences use <code>tr</code> instead. Both Swiss-Prot and TrEMBL are part of UniProtKB.</p>\n\n<p>⚪ This file contains only sequences for proteins with annotations (labeled proteins). To get the full set of protein sequences for unlabeled proteins, download Swiss-Prot and TrEMBL <a href=\"https://www.uniprot.org/help/downloads\" style=\"color:#61dafb;\"><strong>here</strong></a>.</p>\n\n</div>\n","metadata":{"papermill":{"duration":0.012796,"end_time":"2025-11-09T06:03:26.68038","exception":false,"start_time":"2025-11-09T06:03:26.667584","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">🔴 To read and analyze the protein sequences from the <b>train_sequences.fasta</b> file, we can use the <code>Biopython</code> package.</p> ","metadata":{"papermill":{"duration":0.012731,"end_time":"2025-11-09T06:03:26.705903","exception":false,"start_time":"2025-11-09T06:03:26.693172","status":"completed"},"tags":[]}},{"cell_type":"code","source":"print(\"Sequence example:\\n\\n\", next(iter(SeqIO.parse(CFG.train_seq_fasta_path, \"fasta\"))))","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:03.203623Z","iopub.execute_input":"2025-11-18T09:03:03.203892Z","iopub.status.idle":"2025-11-18T09:03:03.216733Z","shell.execute_reply.started":"2025-11-18T09:03:03.203867Z","shell.execute_reply":"2025-11-18T09:03:03.215743Z"},"papermill":{"duration":0.029375,"end_time":"2025-11-09T06:03:26.748337","exception":false,"start_time":"2025-11-09T06:03:26.718962","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">❔ Let's count the number of sequences.</p>","metadata":{"papermill":{"duration":0.012586,"end_time":"2025-11-09T06:03:26.774425","exception":false,"start_time":"2025-11-09T06:03:26.761839","status":"completed"},"tags":[]}},{"cell_type":"code","source":"sequences = SeqIO.parse(CFG.train_seq_fasta_path, \"fasta\")\nnum_sequences = sum(1 for seq in sequences)\n\nprint(\"Number of sequences:\", num_sequences)","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:03.218116Z","iopub.execute_input":"2025-11-18T09:03:03.218434Z","iopub.status.idle":"2025-11-18T09:03:05.0884Z","shell.execute_reply.started":"2025-11-18T09:03:03.218407Z","shell.execute_reply":"2025-11-18T09:03:05.08757Z"},"papermill":{"duration":0.882432,"end_time":"2025-11-09T06:03:27.669608","exception":false,"start_time":"2025-11-09T06:03:26.787176","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">❔ Let's plot the length distribution of the protein sequences.</p>","metadata":{"papermill":{"duration":0.012705,"end_time":"2025-11-09T06:03:27.695416","exception":false,"start_time":"2025-11-09T06:03:27.682711","status":"completed"},"tags":[]}},{"cell_type":"code","source":"from Bio import SeqIO\nimport plotly.express as px\n\n# Load sequences\nsequences = SeqIO.parse(CFG.train_seq_fasta_path, \"fasta\")\n\n# Get the length of each sequence\nlengths = [len(seq) for seq in sequences]\n\n# Create histogram\nfig = px.histogram(\n    x=lengths,\n    nbins=1000,\n    color_discrete_sequence=['goldenrod']\n)\n\n# Dark mode layout\nfig.update_layout(\n    template='plotly_dark',  # Enables dark mode\n    title={\n        'text': \"Distribution of Protein Sequence Lengths\",\n        'y':0.95,\n        'x':0.5,\n        'xanchor': 'center',\n        'yanchor': 'top',\n        'font': {'size': 24, 'color': 'goldenrod'}\n    },\n    xaxis_title=\"Sequence Length\",\n    yaxis_title=\"Count\",\n    xaxis=dict(showgrid=True, gridcolor='gray', zeroline=False),\n    yaxis=dict(showgrid=True, gridcolor='gray', zeroline=False),\n    plot_bgcolor='#1e1e1e',  # Dark background\n    paper_bgcolor='#1e1e1e', # Dark surrounding\n)\n\nfig.show()\n","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:05.089169Z","iopub.execute_input":"2025-11-18T09:03:05.08946Z","iopub.status.idle":"2025-11-18T09:03:07.757158Z","shell.execute_reply.started":"2025-11-18T09:03:05.089438Z","shell.execute_reply":"2025-11-18T09:03:07.756359Z"},"papermill":{"duration":2.372884,"end_time":"2025-11-09T06:03:30.08132","exception":false,"start_time":"2025-11-09T06:03:27.708436","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"np.percentile(lengths, 99)","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:07.758127Z","iopub.execute_input":"2025-11-18T09:03:07.758461Z","iopub.status.idle":"2025-11-18T09:03:07.778786Z","shell.execute_reply.started":"2025-11-18T09:03:07.758433Z","shell.execute_reply":"2025-11-18T09:03:07.777824Z"},"papermill":{"duration":0.035641,"end_time":"2025-11-09T06:03:30.13103","exception":false,"start_time":"2025-11-09T06:03:30.095389","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">⚪ This also means that only <b>1%</b> of the data is present after the value <b>2375</b>.</p>","metadata":{"papermill":{"duration":0.013018,"end_time":"2025-11-09T06:03:30.157417","exception":false,"start_time":"2025-11-09T06:03:30.144399","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">❔ Let's calculate <b>the amino acid composition</b> of each protein sequence. Amino acid composition is the frequency distribution of amino acids in a protein sequence. It can provide valuable information about the protein's structure and function.</p>","metadata":{"papermill":{"duration":0.013078,"end_time":"2025-11-09T06:03:30.184317","exception":false,"start_time":"2025-11-09T06:03:30.171239","status":"completed"},"tags":[]}},{"cell_type":"code","source":"from Bio import SeqIO\nfrom collections import Counter\nimport plotly.express as px\n\n# Load sequences\nrecords = SeqIO.parse(CFG.train_seq_fasta_path, \"fasta\")\n\n# Create a list of all amino acids\naa_list = [aa for record in records for aa in record.seq]\n\n# Count frequency of each amino acid\naa_count = Counter(aa_list)\n\n# Plot horizontal bar chart\nfig = px.bar(\n    x=list(aa_count.values()),\n    y=list(aa_count.keys()),\n    orientation='h',\n    color_discrete_sequence=['darkslateblue'],\n    height=700\n)\n\n# Dark mode layout\nfig.update_layout(\n    template='plotly_dark',\n    title={\n        'text': \"Amino Acid Composition\",\n        'y':0.95,\n        'x':0.5,\n        'xanchor': 'center',\n        'yanchor': 'top',\n        'font': {'size': 24, 'color': 'darkorange'}\n    },\n    xaxis_title=\"Frequency\",\n    yaxis_title=\"Amino Acid\",\n    xaxis=dict(showgrid=True, gridcolor='gray', zeroline=False),\n    yaxis=dict(showgrid=False),\n    plot_bgcolor='#1e1e1e',\n    paper_bgcolor='#1e1e1e'\n)\n\nfig.show()\n","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:07.779732Z","iopub.execute_input":"2025-11-18T09:03:07.78Z","iopub.status.idle":"2025-11-18T09:03:15.757745Z","shell.execute_reply.started":"2025-11-18T09:03:07.779971Z","shell.execute_reply":"2025-11-18T09:03:15.756861Z"},"papermill":{"duration":4.440107,"end_time":"2025-11-09T06:03:34.637455","exception":false,"start_time":"2025-11-09T06:03:30.197348","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">⚪ Here are a few observations that can be made from the obtained frequency values:</p>\n\n* <p style=\"font-family: consolas; font-size: 16px;\">The most common amino acids in this dataset are leucine (L), serine (S), alanine (A), and glutamic (E), tyrosine (Y). These amino acids are known to be abundant in proteins and play important roles in protein structure and function.</p>\n* <p style=\"font-family: consolas; font-size: 16px;\">The least common amino acids in this dataset are cysteine (C), methionine (M), tryptophan (W), and histidine (H). These amino acids are typically less abundant in proteins, but they can be important for specific functions, such as catalysis, metal binding, or protein-protein interactions.</p>\n* <p style=\"font-family: consolas; font-size: 16px;\">The absence of ambiguous amino acids (X, B, Z) and rare amino acids (O, U) in the dataset suggests that sequences are not incomplete or contain errors unlike CAFA 5.</p>","metadata":{"papermill":{"duration":0.013037,"end_time":"2025-11-09T06:03:34.663956","exception":false,"start_time":"2025-11-09T06:03:34.650919","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<div style=\"background-color:#1e1e1e; color:#f0f0f0; padding:25px; border-radius:12px; \n            font-family:'Consolas', monospace; font-size:28px; text-align:center; \n            box-shadow: rgba(0, 0, 0, 0.2) 0px 2px 6px inset;\">\n    <b>2.3 Labels </b>\n</div>\n","metadata":{"papermill":{"duration":0.012937,"end_time":"2025-11-09T06:03:34.690375","exception":false,"start_time":"2025-11-09T06:03:34.677438","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">⚪ <code>Labels</code>: <b>train_terms.tsv</b> contains the list of annotated terms (ground truth) for the proteins in train_sequences.fasta.</p> \n\n* <p style=\"font-family: consolas; font-size: 16px;\">The first column indicates the protein's UniProt accession ID.</p>\n* <p style=\"font-family: consolas; font-size: 16px;\">The second is the GO term ID.</p>\n* <p style=\"font-family: consolas; font-size: 16px;\">The third indicates in which ontology the term appears (BPO, CCO or MFO). BPO, CCO, and MFO are abbreviations for different categories of gene ontology terms. <code>BPO</code>: Biological Process Ontology, which describes biological processes, functions, and pathways. <code>CCO</code>: Cellular Component Ontology, which describes the components of a cell or its extracellular environment. <code>MFO</code>: Molecular Function Ontology, which describes the biochemical activities or capabilities of proteins and other molecules. These categories are used in gene ontology to classify genes and gene products based on their biological roles and functions. By using these categories, researchers can better understand the functions and interactions of different genes and gene products within a biological system.</p>","metadata":{"papermill":{"duration":0.012967,"end_time":"2025-11-09T06:03:34.716682","exception":false,"start_time":"2025-11-09T06:03:34.703715","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">⚪ Load the train terms dataframe.</p> ","metadata":{"papermill":{"duration":0.013192,"end_time":"2025-11-09T06:03:34.743195","exception":false,"start_time":"2025-11-09T06:03:34.730003","status":"completed"},"tags":[]}},{"cell_type":"code","source":"train_terms_df = pd.read_csv(CFG.train_terms_path, sep=\"\\t\")","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:15.758527Z","iopub.execute_input":"2025-11-18T09:03:15.758811Z","iopub.status.idle":"2025-11-18T09:03:19.405554Z","shell.execute_reply.started":"2025-11-18T09:03:15.758781Z","shell.execute_reply":"2025-11-18T09:03:19.404766Z"},"papermill":{"duration":0.430562,"end_time":"2025-11-09T06:03:35.186707","exception":false,"start_time":"2025-11-09T06:03:34.756145","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">❔ Let's look how's data looks like.</p>","metadata":{"papermill":{"duration":0.013407,"end_time":"2025-11-09T06:03:35.213941","exception":false,"start_time":"2025-11-09T06:03:35.200534","status":"completed"},"tags":[]}},{"cell_type":"code","source":"train_terms_df.head()","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:19.406666Z","iopub.execute_input":"2025-11-18T09:03:19.406994Z","iopub.status.idle":"2025-11-18T09:03:19.428566Z","shell.execute_reply.started":"2025-11-18T09:03:19.406967Z","shell.execute_reply":"2025-11-18T09:03:19.427612Z"},"papermill":{"duration":0.038925,"end_time":"2025-11-09T06:03:35.266225","exception":false,"start_time":"2025-11-09T06:03:35.2273","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">❔ Display main information about dataframe columns.</p>","metadata":{"papermill":{"duration":0.014519,"end_time":"2025-11-09T06:03:35.294669","exception":false,"start_time":"2025-11-09T06:03:35.28015","status":"completed"},"tags":[]}},{"cell_type":"code","source":"train_terms_df.describe()","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:19.429501Z","iopub.execute_input":"2025-11-18T09:03:19.429819Z","iopub.status.idle":"2025-11-18T09:03:21.611961Z","shell.execute_reply.started":"2025-11-18T09:03:19.429788Z","shell.execute_reply":"2025-11-18T09:03:21.61108Z"},"papermill":{"duration":0.344462,"end_time":"2025-11-09T06:03:35.652474","exception":false,"start_time":"2025-11-09T06:03:35.308012","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">❔ Now let's plot pie distribution of aspect values.</p>","metadata":{"papermill":{"duration":0.013536,"end_time":"2025-11-09T06:03:35.680677","exception":false,"start_time":"2025-11-09T06:03:35.667141","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import plotly.express as px\n\n# Aspect counts\naspect_counts = train_terms_df.aspect.value_counts()\n\n# Pie chart\nfig = px.pie(\n    values=aspect_counts.values,\n    names=aspect_counts.index,\n    color_discrete_sequence=px.colors.sequential.Viridis\n)\n\n# Dark mode layout\nfig.update_traces(\n    textposition='inside',\n    textfont_size=14,\n    textinfo='percent+label',\n    marker=dict(line=dict(color='#1e1e1e', width=2))\n)\nfig.update_layout(\n    template='plotly_dark',\n    title={\n        'text': \"Pie Distribution of Aspect Values\",\n        'y':0.95,\n        'x':0.5,\n        'xanchor': 'center',\n        'yanchor': 'top',\n        'font': {'size': 22, 'color': 'lightblue'}\n    },\n    legend_title_text='Aspect:',\n    paper_bgcolor='#1e1e1e',\n    plot_bgcolor='#1e1e1e'\n)\n\nfig.show()\n","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:21.612882Z","iopub.execute_input":"2025-11-18T09:03:21.613614Z","iopub.status.idle":"2025-11-18T09:03:22.124208Z","shell.execute_reply.started":"2025-11-18T09:03:21.613586Z","shell.execute_reply":"2025-11-18T09:03:22.123466Z"},"papermill":{"duration":0.35558,"end_time":"2025-11-09T06:03:36.050206","exception":false,"start_time":"2025-11-09T06:03:35.694626","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"background-color:#1e1e1e; color:#f0f0f0; padding:25px; border-radius:12px; \n            font-family:'Consolas', monospace; font-size:28px; text-align:center; \n            box-shadow: rgba(0, 0, 0, 0.2) 0px 2px 6px inset;\">\n    <b>2.4 Taxonomy </b>\n</div>","metadata":{"papermill":{"duration":0.013361,"end_time":"2025-11-09T06:03:36.077227","exception":false,"start_time":"2025-11-09T06:03:36.063866","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">⚪ <code>Taxonomy</code>: <b>train_taxonomy.tsv</b> contains the list of proteins and the species to which they belong, represented by a \"taxonomic identifier\" (taxon ID) number. The first column is the protein UniProt accession ID and the second is the taxon ID. More information about taxonomies can he found <a href=\"https://www.uniprot.org/help/taxonomic_identifier\"><strong>here</strong></a>.</p>","metadata":{"papermill":{"duration":0.013222,"end_time":"2025-11-09T06:03:36.104084","exception":false,"start_time":"2025-11-09T06:03:36.090862","status":"completed"},"tags":[]}},{"cell_type":"code","source":"train_taxonomy_df = pd.read_csv(CFG.train_taxonomy_path, sep=\"\\t\")","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:22.128193Z","iopub.execute_input":"2025-11-18T09:03:22.128524Z","iopub.status.idle":"2025-11-18T09:03:22.245035Z","shell.execute_reply.started":"2025-11-18T09:03:22.128503Z","shell.execute_reply":"2025-11-18T09:03:22.244348Z"},"papermill":{"duration":0.08858,"end_time":"2025-11-09T06:03:36.206089","exception":false,"start_time":"2025-11-09T06:03:36.117509","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">❔ Let's look how's data looks like.</p>","metadata":{"papermill":{"duration":0.013844,"end_time":"2025-11-09T06:03:36.234283","exception":false,"start_time":"2025-11-09T06:03:36.220439","status":"completed"},"tags":[]}},{"cell_type":"code","source":"train_taxonomy_df.head()","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:22.245957Z","iopub.execute_input":"2025-11-18T09:03:22.246268Z","iopub.status.idle":"2025-11-18T09:03:22.254884Z","shell.execute_reply.started":"2025-11-18T09:03:22.246239Z","shell.execute_reply":"2025-11-18T09:03:22.253845Z"},"papermill":{"duration":0.025804,"end_time":"2025-11-09T06:03:36.273945","exception":false,"start_time":"2025-11-09T06:03:36.248141","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">❔ Display dataframe length.</p>","metadata":{"papermill":{"duration":0.015581,"end_time":"2025-11-09T06:03:36.303821","exception":false,"start_time":"2025-11-09T06:03:36.28824","status":"completed"},"tags":[]}},{"cell_type":"code","source":"len(train_taxonomy_df)","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:22.256096Z","iopub.execute_input":"2025-11-18T09:03:22.256523Z","iopub.status.idle":"2025-11-18T09:03:22.275927Z","shell.execute_reply.started":"2025-11-18T09:03:22.256492Z","shell.execute_reply":"2025-11-18T09:03:22.275151Z"},"papermill":{"duration":0.02306,"end_time":"2025-11-09T06:03:36.340932","exception":false,"start_time":"2025-11-09T06:03:36.317872","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"background-color:#1e1e1e; color:#f0f0f0; padding:25px; border-radius:12px; \n            font-family:'Consolas', monospace; font-size:28px; text-align:center; \n            box-shadow: rgba(0, 0, 0, 0.2) 0px 2px 6px inset;\">\n    <b>2.5 Information Accreation </b>\n</div>\n","metadata":{"papermill":{"duration":0.014234,"end_time":"2025-11-09T06:03:36.369972","exception":false,"start_time":"2025-11-09T06:03:36.355738","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"<p style=\"font-family: consolas; font-size: 16px;\">⚪ <code>Information accretion</code>: <b>IA.txt</b> contains the information accretion (weights) for each GO term. These weights are used to compute weighted precision and recall, as described in the Evaluation section of the competition.</p>","metadata":{"papermill":{"duration":0.013651,"end_time":"2025-11-09T06:03:36.397459","exception":false,"start_time":"2025-11-09T06:03:36.383808","status":"completed"},"tags":[]}},{"cell_type":"code","source":"limit = 10\n\nwith open(\"/kaggle/input/cafa-5-protein-function-prediction/IA.txt\") as f:\n    ia_weights = [x.replace(\"\\n\", \"\").split(\"\\t\") for x in f.readlines()]\n\nia_weights[:limit]\n\n","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:22.276811Z","iopub.execute_input":"2025-11-18T09:03:22.277164Z","iopub.status.idle":"2025-11-18T09:03:22.348717Z","shell.execute_reply.started":"2025-11-18T09:03:22.277135Z","shell.execute_reply":"2025-11-18T09:03:22.347779Z"},"papermill":{"duration":0.548318,"end_time":"2025-11-09T06:03:36.959408","exception":false,"start_time":"2025-11-09T06:03:36.41109","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 3. Test Set Analysis\n# ⚪ The test set consists of two main files:\n# - testsuperset.fasta: Contains protein sequences for prediction, with headers including UniProt ID and taxon ID\n# - testsuperset-taxon-list.tsv: Provides taxon ID information for proteins in the test superset\n\nfrom Bio import SeqIO\nfrom collections import Counter\nimport pandas as pd\nimport numpy as np\nimport plotly.express as px\nimport plotly.graph_objects as go\n\n# Update config with test set paths\nclass CFG:\n    # ... (保持原有训练集路径)\n    test_seq_fasta_path: str = \"/kaggle/input/cafa-5-protein-function-prediction/Test (Targets)/testsuperset.fasta\"\n    test_taxon_path: str = \"/kaggle/input/cafa-5-protein-function-prediction/Test (Targets)/testsuperset-taxon-list.tsv\"\n    train_taxonomy_path : str = \"/kaggle/input/cafa-5-protein-function-prediction/Train/train_taxonomy.tsv\"\n    train_seq_fasta_path : str = \"/kaggle/input/cafa-5-protein-function-prediction/Train/train_sequences.fasta\"\n    \n# 3.1 Test Sequences Analysis\n# ❔ Let's first analyze the protein sequences in testsuperset.fasta\n# 修正训练集分类数据的列名\n# train_taxonomy_df = pd.read_csv(\n#     CFG.train_taxonomy_path, \n#     sep=\"\\t\", \n#     header=None,  # 声明文件没有表头\n#     names=['EntryID', 'Species']  # 手动指定列名\n# )\n# Count number of test sequences\ntest_sequences = SeqIO.parse(CFG.test_seq_fasta_path, \"fasta\")\nnum_test_sequences = sum(1 for seq in test_sequences)\nprint(f\"Number of test sequences: {num_test_sequences}\")\n\n# Sequence length distribution\ntest_sequences = SeqIO.parse(CFG.test_seq_fasta_path, \"fasta\")\ntest_lengths = [len(seq) for seq in test_sequences]\n\n# Plot length distribution\nfig = px.histogram(\n    x=test_lengths,\n    nbins=1000,\n    color_discrete_sequence=['crimson']\n)\n\nfig.update_layout(\n    template='plotly_dark',\n    title={\n        'text': \"Distribution of Test Protein Sequence Lengths\",\n        'y':0.95,\n        'x':0.5,\n        'xanchor': 'center',\n        'yanchor': 'top',\n        'font': {'size': 24, 'color': 'crimson'}\n    },\n    xaxis_title=\"Sequence Length\",\n    yaxis_title=\"Count\",\n    xaxis=dict(showgrid=True, gridcolor='gray', zeroline=False),\n    yaxis=dict(showgrid=True, gridcolor='gray', zeroline=False),\n    plot_bgcolor='#1e1e1e',\n    paper_bgcolor='#1e1e1e',\n)\n\nfig.show()\n\n# 99th percentile of test sequence lengths\nprint(f\"99th percentile of test sequence lengths: {np.percentile(test_lengths, 99)}\")\n\n# Amino acid composition in test set\ntest_sequences = SeqIO.parse(CFG.test_seq_fasta_path, \"fasta\")\ntest_aa_list = [aa for record in test_sequences for aa in record.seq]\ntest_aa_count = Counter(test_aa_list)\n\n# Plot amino acid composition\nfig = px.bar(\n    x=list(test_aa_count.values()),\n    y=list(test_aa_count.keys()),\n    orientation='h',\n    color_discrete_sequence=['indigo'],\n    height=700\n)\n\nfig.update_layout(\n    template='plotly_dark',\n    title={\n        'text': \"Test Set Amino Acid Composition\",\n        'y':0.95,\n        'x':0.5,\n        'xanchor': 'center',\n        'yanchor': 'top',\n        'font': {'size': 24, 'color': 'indigo'}\n    },\n    xaxis_title=\"Frequency\",\n    yaxis_title=\"Amino Acid\",\n    xaxis=dict(showgrid=True, gridcolor='gray', zeroline=False),\n    yaxis=dict(showgrid=False),\n    plot_bgcolor='#1e1e1e',\n    paper_bgcolor='#1e1e1e'\n)\n\nfig.show()\n\n# 3.2 Test Taxonomy Analysis\n# ❔ 从测试集Fasta头中提取物种编号（Taxon ID）\n\n# 解析测试集Fasta，提取每个序列的物种编号\ntest_sequences = SeqIO.parse(CFG.test_seq_fasta_path, \"fasta\")\ntest_taxon_data = []\nfor seq in test_sequences:\n    entry_id = seq.id\n    taxon_id = seq.description.split()[1]  # 序列头格式：\"EntryID TaxonID\"，分割后取第二个元素\n    test_taxon_data.append({'EntryID': entry_id, 'Species': taxon_id})\n\n# 构建测试集分类数据DataFrame\ntest_taxonomy_df = pd.DataFrame(test_taxon_data)\ntest_taxonomy_df['Species'] = test_taxonomy_df['Species'].astype(str)  # 统一为字符串类型\nprint(\"测试集物种数据（从Fasta提取）预览：\")\ndisplay(test_taxonomy_df.head())\n\n# Basic statistics\nprint(f\"Number of test proteins in taxon list: {len(test_taxonomy_df)}\")\nprint(f\"Number of unique taxon IDs in test set: {test_taxonomy_df['Species'].nunique()}\")\n\n# Top 10 most common species in test set\ntop_test_species = test_taxonomy_df['Species'].value_counts().head(10)\n\n# Plot top species distribution\nfig = px.bar(\n    x=top_test_species.values,\n    y=top_test_species.index,\n    orientation='h',\n    color_discrete_sequence=['limegreen']\n)\n\nfig.update_layout(\n    template='plotly_dark',\n    title={\n        'text': \"Top 10 Most Common Species in Test Set\",\n        'y':0.95,\n        'x':0.5,\n        'xanchor': 'center',\n        'yanchor': 'top',\n        'font': {'size': 22, 'color': 'limegreen'}\n    },\n    xaxis_title=\"Number of Proteins\",\n    yaxis_title=\"Species (Taxon ID)\",\n    xaxis=dict(showgrid=True, gridcolor='gray', zeroline=False),\n    yaxis=dict(showgrid=False),\n    plot_bgcolor='#1e1e1e',\n    paper_bgcolor='#1e1e1e',\n    height=600\n)\n\nfig.show()\n\n# ==============================================================================\n# 2. 加载训练集分类数据\n# ==============================================================================\ntrain_taxonomy_df = pd.read_csv(\n    CFG.train_taxonomy_path, \n    sep=\"\\t\", \n    header=None, \n    names=['EntryID', 'Species']\n)\n# ==============================================================================\n# 关键：直接使用原始的物种编号（Taxon ID）进行对比\n# 确保训练集和测试集的物种编号都是字符串类型，以避免类型错误\n# ==============================================================================\n\n# ==============================================================================\n# 关键：在使用 Matplotlib 绘图前，配置支持中文的字体\n# ==============================================================================\nimport matplotlib.pyplot as plt\n\n# 设置字体以支持中文显示\n# 'WenQuanYi Micro Hei' 和 'Heiti TC' 是 Linux 和 macOS 上常见的中文字体\n# 如果找不到，可以尝试其他字体，如 'SimHei' (Windows)\nplt.rcParams['font.sans-serif'] = ['WenQuanYi Micro Hei', 'Heiti TC', 'DejaVu Sans']\n# 解决负号 '-' 在图表中显示为方框的问题\nplt.rcParams['axes.unicode_minus'] = False\n\n# ==============================================================================\n# 关键：在使用 Matplotlib 绘图前，配置支持中文的字体\n# ==============================================================================\nimport matplotlib.pyplot as plt\n\n# 设置字体以支持中文显示\n# 'WenQuanYi Micro Hei' 和 'Heiti TC' 是 Linux 和 macOS 上常见的中文字体\n# 如果找不到，可以尝试其他字体，如 'SimHei' (Windows)\nplt.rcParams['font.sans-serif'] = ['WenQuanYi Micro Hei', 'Heiti TC', 'DejaVu Sans']\n# 解决负号 '-' 在图表中显示为方框的问题\nplt.rcParams['axes.unicode_minus'] = False\n\n# ==============================================================================\n# 1. 重新加载并预处理训练集分类数据\n# ==============================================================================\n# ==============================================================================\n# 1. 重新加载并预处理训练集分类数据\n# ==============================================================================\nprint(\"🔄 Reloading and preprocessing training taxonomy data...\")\ntry:\n    train_taxonomy_df = pd.read_csv(\n        CFG.train_taxonomy_path, \n        sep=\"\\t\", \n        header=None, \n        names=['EntryID', 'Species']\n    )\n    train_taxonomy_df['Species'] = train_taxonomy_df['Species'].astype(str)\n    print(\"✅ Training taxonomy data reloaded successfully.\")\nexcept Exception as e:\n    print(f\"❌ Failed to reload training taxonomy data. Error: {e}\")\n\n# ==============================================================================\n# 2. 从测试集Fasta文件中提取原始物种编号\n# ==============================================================================\nprint(\"\\n🔄 Extracting taxon IDs from test FASTA file...\")\ntry:\n    test_sequences = SeqIO.parse(CFG.test_seq_fasta_path, \"fasta\")\n    test_taxon_data = []\n    for seq in test_sequences:\n        entry_id = seq.id\n        taxon_id = seq.description.split()[1]\n        test_taxon_data.append({'EntryID': entry_id, 'Species': taxon_id})\n    \n    test_taxonomy_df = pd.DataFrame(test_taxon_data)\n    test_taxonomy_df['Species'] = test_taxonomy_df['Species'].astype(str)\n    print(\"✅ Test taxon IDs extracted successfully.\")\nexcept Exception as e:\n    print(f\"❌ Failed to extract test taxon IDs. Error: {e}\")\n\n# ==============================================================================\n# 3. 统计物种数量并可视化（完全英文版本）\n# ==============================================================================\ntry:\n    # 计算物种数量\n    train_species_count = train_taxonomy_df['Species'].nunique()\n    test_species_count = test_taxonomy_df['Species'].nunique()\n    train_species_set = set(train_taxonomy_df['Species'].unique())\n    test_species_set = set(test_taxonomy_df['Species'].unique())\n    common_species_count = len(train_species_set.intersection(test_species_set))\n    \n    print(\"\\n📊 Species Count Statistics:\")\n    print(f\"Total species in Training Set: {train_species_count}\")\n    print(f\"Total species in Test Set: {test_species_count}\")\n    print(f\"Common species between both sets: {common_species_count}\")\n    print(f\"Species unique to Training Set: {train_species_count - common_species_count}\")\n    print(f\"Species unique to Test Set: {test_species_count - common_species_count}\")\n    \n    # --- 绘图代码 (完全英文) ---\n    import matplotlib.pyplot as plt\n\n    fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(18, 7))\n    \n    # 柱状图\n    categories = ['Training Set', 'Test Set']\n    counts = [train_species_count, test_species_count]\n    bars = ax1.bar(categories, counts, color=['royalblue', 'crimson'], width=0.5)\n    ax1.set_title('Total Number of Species: Training vs Test Set', fontsize=16, pad=20)\n    ax1.set_ylabel('Number of Species', fontsize=14)\n    ax1.grid(axis='y', linestyle='--', alpha=0.7)\n    for bar, count in zip(bars, counts):\n        height = bar.get_height()\n        ax1.text(bar.get_x() + bar.get_width()/2., height + 100,\n                f'{count}', ha='center', va='bottom', fontsize=14)\n    \n    # 饼图\n    labels = ['Common Species', 'Test Set Unique']\n    sizes = [common_species_count, test_species_count - common_species_count]\n    colors = ['#ff9999', '#99ff99']\n    wedges, texts, autotexts = ax2.pie(sizes, labels=labels, colors=colors, \n                                       autopct='%1.1f%%', startangle=90,\n                                       textprops={'fontsize': 12})\n    ax2.set_title('Species Distribution', fontsize=16, pad=20)\n    \n    plt.tight_layout()\n    plt.show()\n    \nexcept Exception as e:\n    print(f\"\\n⚠️ Species count statistics or plotting failed. Detailed error: {type(e).__name__} - {e}\")\n\n# ==============================================================================\n# 4. 进行物种分布对比（Plotly部分）\n# ==============================================================================\ntry:\n    top_train_species = train_taxonomy_df['Species'].value_counts().head(10)\n    top_test_species = test_taxonomy_df['Species'].value_counts().head(10)\n    \n    print(\"\\n===== Test Set Top 10 Species Distribution (Taxon ID) =====\")\n    print(top_test_species)\n    print(\"\\n===== Training Set Top 10 Species Distribution (Taxon ID) =====\")\n    print(top_train_species)\n    \n    compare_df = pd.DataFrame({\n        'Test Set': top_test_species,\n        'Training Set': top_train_species\n    }).fillna(0)\n    \n    compare_df['Total'] = compare_df['Test Set'] + compare_df['Training Set']\n    compare_df = compare_df.sort_values(by='Total', ascending=True).tail(15)\n    \n    fig = go.Figure(data=[\n        go.Bar(name='Test Set', y=compare_df.index, x=compare_df['Test Set'], orientation='h', marker_color='royalblue'),\n        go.Bar(name='Training Set', y=compare_df.index, x=compare_df['Training Set'], orientation='h', marker_color='crimson')\n    ])\n    \n    fig.update_layout(\n        template='plotly_dark',\n        barmode='group',\n        title={\n            'text': \"Comparison of Species Distribution by Taxon ID (Test vs Training)\",\n            'y':0.95,\n            'x':0.5,\n            'xanchor': 'center',\n            'yanchor': 'top'\n        },\n        xaxis_title=\"Number of Proteins\",\n        yaxis_title=\"Species (Taxon ID)\",\n        plot_bgcolor='#1e1e1e',\n        paper_bgcolor='#1e1e1e',\n        height=700,\n        legend=dict(orientation='h', yanchor='bottom', y=1.02, xanchor='right', x=1)\n    )\n    \n    fig.show()\n    print(\"\\n✅ Comparison plot generated successfully!\")\n\nexcept Exception as e:\n    print(f\"\\n⚠️ Comparison failed. Detailed error: {type(e).__name__} - {e}\")","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:22.35003Z","iopub.execute_input":"2025-11-18T09:03:22.350363Z","iopub.status.idle":"2025-11-18T09:03:34.113932Z","shell.execute_reply.started":"2025-11-18T09:03:22.350331Z","shell.execute_reply":"2025-11-18T09:03:34.113205Z"},"papermill":{"duration":16.513939,"end_time":"2025-11-09T06:03:53.487614","exception":false,"start_time":"2025-11-09T06:03:36.973675","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ==============================================================================\n# 分析训练集和测试集蛋白序列的重合比例\n# ==============================================================================\ntry:\n    # 1. 提取训练集所有蛋白序列\n    print(\"🔄 Extracting training set sequences...\")\n    train_sequences = SeqIO.parse(CFG.train_seq_fasta_path, \"fasta\")\n    train_seq_set = set()\n    train_seq_count = 0\n    \n    for seq in train_sequences:\n        train_seq_count += 1\n        # 将序列转换为字符串并添加到集合中（自动去重）\n        train_seq_set.add(str(seq.seq))\n    \n    print(f\"✅ Training set: {train_seq_count} total sequences, {len(train_seq_set)} unique sequences\")\n    \n    # 2. 提取测试集所有蛋白序列\n    print(\"\\n🔄 Extracting test set sequences...\")\n    test_sequences = SeqIO.parse(CFG.test_seq_fasta_path, \"fasta\")\n    test_seq_set = set()\n    test_seq_count = 0\n    \n    for seq in test_sequences:\n        test_seq_count += 1\n        test_seq_set.add(str(seq.seq))\n    \n    print(f\"✅ Test set: {test_seq_count} total sequences, {len(test_seq_set)} unique sequences\")\n    \n    # 3. 计算重合序列\n    overlapping_sequences = train_seq_set.intersection(test_seq_set)\n    overlapping_count = len(overlapping_sequences)\n    \n    # 4. 计算比例\n    train_overlap_ratio = overlapping_count / len(train_seq_set) * 100 if len(train_seq_set) > 0 else 0\n    test_overlap_ratio = overlapping_count / len(test_seq_set) * 100 if len(test_seq_set) > 0 else 0\n    \n    # 输出统计结果\n    print(\"\\n📊 Sequence Overlap Analysis:\")\n    print(f\"Number of overlapping unique sequences: {overlapping_count}\")\n    print(f\"Overlap ratio in training set: {train_overlap_ratio:.2f}%\")\n    print(f\"Overlap ratio in test set: {test_overlap_ratio:.2f}%\")\n    \n    # 5. 可视化展示\n    fig, ((ax1, ax2), (ax3, ax4)) = plt.subplots(2, 2, figsize=(18, 12))\n    \n    # 子图1：序列总数对比\n    categories = ['Training Set', 'Test Set']\n    total_counts = [train_seq_count, test_seq_count]\n    bars1 = ax1.bar(categories, total_counts, color=['royalblue', 'crimson'], alpha=0.8)\n    ax1.set_title('Total Number of Sequences', fontsize=14, pad=15)\n    ax1.set_ylabel('Sequence Count')\n    ax1.grid(axis='y', linestyle='--', alpha=0.7)\n    for bar, count in zip(bars1, total_counts):\n        height = bar.get_height()\n        ax1.text(bar.get_x() + bar.get_width()/2., height + 500,\n                f'{count}', ha='center', va='bottom', fontsize=12)\n    \n    # 子图2：唯一序列数对比\n    unique_counts = [len(train_seq_set), len(test_seq_set)]\n    bars2 = ax2.bar(categories, unique_counts, color=['royalblue', 'crimson'], alpha=0.8)\n    ax2.set_title('Number of Unique Sequences', fontsize=14, pad=15)\n    ax2.set_ylabel('Unique Sequence Count')\n    ax2.grid(axis='y', linestyle='--', alpha=0.7)\n    for bar, count in zip(bars2, unique_counts):\n        height = bar.get_height()\n        ax2.text(bar.get_x() + bar.get_width()/2., height + 500,\n                f'{count}', ha='center', va='bottom', fontsize=12)\n    \n    # 子图3：重合序列占比饼图\n    labels3 = ['Overlapping Sequences', 'Unique to Training']\n    sizes3 = [overlapping_count, len(train_seq_set) - overlapping_count]\n    colors3 = ['#ff9999', '#66b3ff']\n    ax3.pie(sizes3, labels=labels3, colors=colors3, autopct='%1.1f%%', startangle=90)\n    ax3.set_title('Sequence Distribution in Training Set', fontsize=14, pad=15)\n    \n    # 子图4：重合序列占比饼图（测试集视角）\n    labels4 = ['Overlapping Sequences', 'Unique to Test']\n    sizes4 = [overlapping_count, len(test_seq_set) - overlapping_count]\n    colors4 = ['#ff9999', '#99ff99']\n    ax4.pie(sizes4, labels=labels4, colors=colors4, autopct='%1.1f%%', startangle=90)\n    ax4.set_title('Sequence Distribution in Test Set', fontsize=14, pad=15)\n    \n    plt.tight_layout()\n    plt.show()\n    \n    # 6. 额外统计：训练集内部重复序列情况\n    train_duplicate_count = train_seq_count - len(train_seq_set)\n    test_duplicate_count = test_seq_count - len(test_seq_set)\n    \n    print(\"\\n📈 Duplicate Sequence Analysis:\")\n    print(f\"Duplicate sequences in training set: {train_duplicate_count} ({train_duplicate_count/train_seq_count*100:.2f}%)\")\n    print(f\"Duplicate sequences in test set: {test_duplicate_count} ({test_duplicate_count/test_seq_count*100:.2f}%)\")\n    \nexcept Exception as e:\n    print(f\"\\n⚠️ Sequence overlap analysis failed. Detailed error: {type(e).__name__} - {e}\")","metadata":{"execution":{"iopub.status.busy":"2025-11-18T09:03:34.114673Z","iopub.execute_input":"2025-11-18T09:03:34.114884Z","iopub.status.idle":"2025-11-18T09:03:36.595499Z","shell.execute_reply.started":"2025-11-18T09:03:34.114866Z","shell.execute_reply":"2025-11-18T09:03:36.594591Z"},"papermill":{"duration":2.437999,"end_time":"2025-11-09T06:03:55.942459","exception":false,"start_time":"2025-11-09T06:03:53.50446","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null}]}