{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":118765,"databundleVersionId":15231210,"sourceType":"competition"},{"sourceId":11775065,"sourceType":"datasetVersion","datasetId":7392749},{"sourceId":11788894,"sourceType":"datasetVersion","datasetId":7402173}],"dockerImageVersionId":31012,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### [Stanford RNA 3D Folding Part 2](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/code?competitionId=118765&sortBy=scoreDescending&excludeNonAccessedDatasources=true)\n\nSolve RNA structure prediction, one of biology's remaining grand challenges\n\n\n| | | &nbsp; | | &nbsp; | &nbsp; | &nbsp; |\n|:-|:-:| :-: | :-: | :-: | :-: | :-: |\n| [Model_1](#Model_1) | [0.177](https://www.kaggle.com/code/rhijudas/mmseqs2-3d-rna-template-identification-part-2?scriptVersionId=290653201) |&nbsp;v.1&nbsp;| [MMseqs2 3D RNA Template Identification Part 2](https://www.kaggle.com/code/rhijudas/mmseqs2-3d-rna-template-identification-part-2) | expert | [Rhiju Das](https://www.kaggle.com/rhijudas) | United States |\n| [Model_2](#Model_2) | [0.114](https://www.kaggle.com/code/khaaliswooden/stanford-rna-3d-a-form-helix-baseline?scriptVersionId=290830613) |&nbsp;v.1&nbsp;| [Stanford RNA 3D - A-Form Helix Baseline](https://www.kaggle.com/code/khaaliswooden/stanford-rna-3d-a-form-helix-baseline) | contributor | [Khaalis Wooden](https://www.kaggle.com/khaaliswooden) | United States |\n| [Model_3](#Model_3) | [0.111](https://www.kaggle.com/code/mdferozahmedafm/stanford-rna-3d-folding-part2-baseline?scriptVersionId=290801544) |&nbsp;v.2&nbsp;| [Stanford RNA 3D Folding Part2 Baseline](https://www.kaggle.com/code/mdferozahmedafm/stanford-rna-3d-folding-part2-baseline) | expert | [Md Feroz Ahmed](https://www.kaggle.com/mdferozahmedafm) | Bangladesh |\n| [Model_4](#Model_4) | []() |&nbsp;&nbsp;| []() |  | []() | |\n|||||||\n| [Model_5](#Model_5) | [0.110]() |&nbsp;v.9&nbsp;| [RNA 3D baseline CPU](https://www.kaggle.com/code/antonoof/rna-3d-baseline-cpu) | master      | [Antonoof](https://www.kaggle.com/antonoof) | Russia |\n| [Model_6](#Model_6) | [0.109]() |&nbsp;v.3&nbsp;| [baseline on CPU](https://www.kaggle.com/code/krylovalexander/baseline-on-cpu) | contributor | [Krylov Alexander](https://www.kaggle.com/krylovalexander) | Russia |\n| [Model_7](#Model_7) | [0.102]() |&nbsp;v.2&nbsp;| [Stanford RNA 3D Folding Part 2](https://www.kaggle.com/code/hmnshudhmn24/stanford-rna-3d-folding-part-2) | contributor | [Himanshu Dhiman](https://www.kaggle.com/hmnshudhmn24) | India |\n| [Model_8](#Model_8) | [0.099]() |&nbsp;v.1&nbsp;| [Helical model](https://www.kaggle.com/code/smartmanoj/helical-model) | contributor | [Smart Manoj](https://www.kaggle.com/smartmanoj) | India |\n||||||||\n|||| **ensemble models**| | **main weight** |\n|  | [0.102](https://www.kaggle.com/code/antonoof/rna-3d-baseline-cpu) |  | [ 5., 6., 7., 8. ] |  | [ +25 +25 +25 +25 ] |\n|  | [0.110](https://www.kaggle.com/code/nina2025/stanford-rna-direct-addition?scriptVersionId=290842420) |  | [ 5., 6., 7., 8. ]  |  | [ +40 +30 +20 +10 ] |\n|||||||\n|  | [0.094](https://www.kaggle.com/code/nina2025/stanford-rna-3d-folding-p2-ensemble-of-solutions?scriptVersionId=290883176) | v.1  | [ 1., 2., 3., 5. ] | | [ +25 +25 +25 +25 ] |\n|  | [0.095](https://www.kaggle.com/code/nina2025/stanford-rna-3d-folding-p2-ensemble-of-solutions?scriptVersionId=290892542) | v.2 | [ 1., 2., 3., 5. ] | | [ +40 +30 +20 +10 ] |\n|  | [?]() | v.5 | [ 2., 3., 5., 8. ] | | [ +90 +02 +07 +01 ] |","metadata":{}},{"cell_type":"code","source":"ensemble_of_Solutions = {\n    'Model_2':0.90,\n    'Model_3':0.02,\n    'Model_5':0.07,\n    'Model_8':0.01\n}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:43:27.864283Z","iopub.execute_input":"2026-01-09T05:43:27.864631Z","iopub.status.idle":"2026-01-09T05:43:27.873714Z","shell.execute_reply.started":"2026-01-09T05:43:27.864605Z","shell.execute_reply":"2026-01-09T05:43:27.872456Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **Model_1**","metadata":{}},{"cell_type":"markdown","source":"### Identify 3D templates for RNA targets\nThis notebook prepares a mock submission as an illustration of how to derive 3D templates for targets in the Stanford 3D RNA folding competition. See also https://www.kaggle.com/datasets/rhijudas/rna-3d-folding-templates/data for precomputed outputs useful for training models for RNA 3D folding.\n\n### Important options","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"if 'Model_1' in ensemble_of_Solutions: \n\n    # if False, don't check on temporal_cutoff -- during training on known structures, will get lots of leakage. \n    # If going for Early Sharing Prize, set to True.\n    CHECK_TEMPORAL_CUTOFF = True\n    \n    # Number of templates to use. \n    # Here set to 5 to prepare a mock submission\n    # Should make larger (e.g., 40) if using templates for modeling. \n    MAX_TEMPLATES = 5\n    \n    # Better to use nan when preparing files for templates, to allow easy recognition of which coordinates are missing.\n    # But for this example, using 0.0 to avoid errors in scoring the final submission.csv\n    NULL_VALUE = 0.0  \n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:44:50.679658Z","iopub.execute_input":"2026-01-09T05:44:50.680007Z","iopub.status.idle":"2026-01-09T05:44:50.685949Z","shell.execute_reply.started":"2026-01-09T05:44:50.679984Z","shell.execute_reply":"2026-01-09T05:44:50.684389Z"},"jupyter":{"source_hidden":true},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Download MMseqs2","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"if 'Model_1' in ensemble_of_Solutions: \n\n    #!wget https://mmseqs.com/latest/mmseqs-linux-avx2.tar.gz\n    #!tar xvfz /kaggle/working/mmseqs-linux-avx2.tar.gz\n    !rsync -avL /kaggle/input/mmseqs2/mmseqs /kaggle/working/\n    !chmod 755 /kaggle/working/mmseqs/bin/mmseqs","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:44:54.713271Z","iopub.execute_input":"2026-01-09T05:44:54.713613Z","iopub.status.idle":"2026-01-09T05:44:55.721859Z","shell.execute_reply.started":"2026-01-09T05:44:54.713587Z","shell.execute_reply":"2026-01-09T05:44:55.720334Z"},"_kg_hide-input":true,"_kg_hide-output":true,"jupyter":{"source_hidden":true,"outputs_hidden":true},"collapsed":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Create DB based on FASTA of all PDB nucleic acid sequences, which is part of PDB_RNA dataset","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"if 'Model_1' in ensemble_of_Solutions: \n\n    !/kaggle/working/mmseqs/bin/mmseqs createdb /kaggle/input/stanford-rna-3d-folding-2/PDB_RNA/pdb_seqres_NA.fasta pdb_seqres_NA --dbtype 2 # > MMseqs_createDB.log","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:45:03.647277Z","iopub.execute_input":"2026-01-09T05:45:03.647618Z","iopub.status.idle":"2026-01-09T05:45:04.038884Z","shell.execute_reply.started":"2026-01-09T05:45:03.647592Z","shell.execute_reply":"2026-01-09T05:45:04.037234Z"},"collapsed":true,"jupyter":{"source_hidden":true,"outputs_hidden":true},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Find templates for targets in test_sequences.csv by aligning to PDB with MMseqs2","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"markdown","source":"### Need to  convert test_sequences.csv file to FASTA","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"if 'Model_1' in ensemble_of_Solutions: \n\n    import csv\n    input_file='/kaggle/input/stanford-rna-3d-folding-2/test_sequences.csv'\n    output_file='test_sequences.fasta'\n    with open(input_file, 'r', newline='') as csv_file, open(output_file, 'w') as fasta_file:\n        csv_reader = csv.reader(csv_file, quotechar='\"', delimiter=',', quoting=csv.QUOTE_ALL, skipinitialspace=True)\n        next(csv_reader)  # Skip the header row\n        for row in csv_reader:\n            if len(row) >= 2:\n                fasta_file.write(f\">{row[0]}\\n{row[1]}\\n\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:45:10.995662Z","iopub.execute_input":"2026-01-09T05:45:10.996039Z","iopub.status.idle":"2026-01-09T05:45:11.011039Z","shell.execute_reply.started":"2026-01-09T05:45:10.996009Z","shell.execute_reply":"2026-01-09T05:45:11.010097Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if 'Model_1' in ensemble_of_Solutions: \n\n    !/kaggle/working/mmseqs/bin/mmseqs easy-search /kaggle/working/test_sequences.fasta /kaggle/working/pdb_seqres_NA testResult.txt tmp --search-type 3 --format-output \"query,target,evalue,qstart,qend,tstart,tend,qaln,taln\" ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:45:14.91431Z","iopub.execute_input":"2026-01-09T05:45:14.914646Z","iopub.status.idle":"2026-01-09T05:45:33.42353Z","shell.execute_reply.started":"2026-01-09T05:45:14.914623Z","shell.execute_reply":"2026-01-09T05:45:33.421928Z"},"_kg_hide-input":true,"_kg_hide-output":true,"jupyter":{"outputs_hidden":true},"collapsed":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Assemble file of template coordinates by going through .cif files found for each target by MMseqs2","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"if 'Model_1' in ensemble_of_Solutions: \n\n    !pip3 install /kaggle/input/biopython/biopython-1.85-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl --no-deps","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:45:45.349378Z","iopub.execute_input":"2026-01-09T05:45:45.349732Z","iopub.status.idle":"2026-01-09T05:45:50.232617Z","shell.execute_reply.started":"2026-01-09T05:45:45.3497Z","shell.execute_reply":"2026-01-09T05:45:50.231318Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if 'Model_1' in ensemble_of_Solutions: \n\n    from Bio import SeqIO,PDB,BiopythonWarning\n    from Bio.PDB.MMCIF2Dict import MMCIF2Dict\n    from Bio.Seq import Seq\n    from Bio.PDB import MMCIFParser\n    import numpy as np\n    import pandas as pd\n    import os\n    import gzip\n    import sys\n    import warnings\n    import argparse\n    from datetime import datetime\n    \n    # Suppress warnings\n    warnings.simplefilter('ignore', BiopythonWarning)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:45:53.613402Z","iopub.execute_input":"2026-01-09T05:45:53.613746Z","iopub.status.idle":"2026-01-09T05:45:54.074397Z","shell.execute_reply.started":"2026-01-09T05:45:53.613718Z","shell.execute_reply":"2026-01-09T05:45:54.073359Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Basic options and file locations","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"if 'Model_1' in ensemble_of_Solutions: \n\n    sequences_file = '/kaggle/input/stanford-rna-3d-folding-2/test_sequences.csv'\n    mmseqs_results_file = '/kaggle/working/testResult.txt' # created by MMseqs command lines\n    outfile = 'subm_1.csv'\n    cif_dir = '/kaggle/input/stanford-rna-3d-folding-2/PDB_RNA'\n    \n    # variables not in use here:\n    id_map_file = ''\n    start_idx = 0\n    end_idx = 0","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:45:58.395259Z","iopub.execute_input":"2026-01-09T05:45:58.395829Z","iopub.status.idle":"2026-01-09T05:45:58.401507Z","shell.execute_reply.started":"2026-01-09T05:45:58.3958Z","shell.execute_reply":"2026-01-09T05:45:58.400154Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Helper functions","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"if 'Model_1' in ensemble_of_Solutions: \n\n    def clean_res_name( res_name ):\n        if res_name in ['A', 'C', 'G', 'U']:\n            return res_name\n        else: # can be modified residue with 3-letter name.\n            return 'X'\n    \n    def extract_title_release_date( cif_path ):\n    \n        if cif_path.endswith('.gz'):\n            with gzip.open(cif_path, 'rt') as cif_file:\n                mmcif_dict = MMCIF2Dict(cif_file)\n        else:\n            mmcif_dict = MMCIF2Dict(cif_path)\n    \n        possible_title_fields = [\n            '_struct.title',\n            '_entry.title',\n            '_struct_keywords.pdbx_keywords'\n        ]\n    \n        pdb_title = None\n        for field in possible_title_fields:\n            if field in mmcif_dict:\n                pdb_title = mmcif_dict[field]\n                if isinstance(pdb_title, list):\n                    pdb_title = ' '.join(pdb_title)\n                break\n    \n        possible_date_fields = [\n            '_pdbx_database_status.initial_release_date',\n            '_pdbx_database_status.recvd_initial_deposition_date',\n            '_database_PDB_rev.date'\n        ]\n    \n        release_date = None\n        for field in possible_date_fields:\n            if field in mmcif_dict:\n                release_date = mmcif_dict[field]\n                if isinstance(release_date, list):\n                    release_date = release_date[0]  # Take the first date if it's a list\n                break\n    \n        return pdb_title, release_date\n    \n    \n    def extract_rna_sequence(cif_path,chain_id):\n    \n        if cif_path.endswith('.gz'):\n            with gzip.open(cif_path, 'rt') as cif_file:\n                mmcif_dict = MMCIF2Dict(cif_file)\n        else:\n            mmcif_dict = MMCIF2Dict(cif_path)\n    \n        pdb_sequence = None\n        pdb_chain_id = None\n        chain_seq_nums = None\n    \n        # Extract _pdbx_poly_seq_scheme information\n        strand_id  = mmcif_dict.get('_pdbx_poly_seq_scheme.pdb_strand_id',[])\n        mon_id     = mmcif_dict.get('_pdbx_poly_seq_scheme.mon_id',[])\n        pdb_mon_id = mmcif_dict.get('_pdbx_poly_seq_scheme.pdb_mon_id',[])\n        pdb_seq_num = mmcif_dict.get('_pdbx_poly_seq_scheme.pdb_seq_num',[])\n        chain_ids = list(set(strand_id))\n        seq_chains = []\n    \n        full_sequence = ''\n        pdb_chain_sequence = ''\n        pdb_chain_seq_nums = []\n        for (strand,mon,pdb_mon,pdb_num) in zip(strand_id,mon_id,pdb_mon_id,pdb_seq_num):\n            if strand==chain_id:\n                full_sequence += clean_res_name( mon )\n                pdb_chain_sequence += clean_res_name( pdb_mon )\n                pdb_chain_seq_nums.append( pdb_num)\n    \n        #print(full_sequence)\n        #print(pdb_chain_sequence)\n        #print(pdb_chain_seq_nums)\n    \n        return full_sequence,pdb_chain_sequence,pdb_chain_seq_nums\n    \n    def get_c1prime_labels(cif_path, chain_id, alignment, chain_seq_nums):\n        \"\"\"\n        Extract C1' coordinates for an RNA chain based on a reference sequence alignment.\n    \n        This function uses Biopython to parse a CIF file, finds the specified chain,\n        and extracts C1' coordinates for RNA residues. It aligns these coordinates\n        with a reference sequence, handling gaps and missing residues.\n    \n        Parameters:\n        cif_path (str): Path to the CIF file.\n        chain_id (str): Chain identifier in the CIF file.\n        alignment (list): A list containing two elements:\n                          alignment[0]: List of residues for the reference sequence (A,C,G,U,-)\n                          alignment[1]: List of residues for the chain sequence (A,C,G,U,X,-)\n        chain_seq_nums (list): numbers of residues in PDB\n    \n        Returns:\n        list of tuples: Each tuple contains (resname, resid, x, y, z), where:\n                        resname: Residue name (A, C, G, or U) from the reference sequence\n                        resid: Residue ID (1, 2, 3, ...) based on position in reference sequence\n                        x, y, z: C1' coordinates (nan or NULL_VALUE for missing residues/atoms)\n    \n        The length of the returned list is equal to the number of non-gap residues\n        in the reference sequence.\n        \"\"\"\n        # Parse the CIF file\n        parser = MMCIFParser()\n        if cif_path.endswith('.gz'):\n            with gzip.open(cif_path, 'rt') as gz_file:\n                structure = parser.get_structure('RNA', gz_file )\n        else:\n            structure = parser.get_structure('RNA', cif_path)\n    \n        # Get the specified chain\n        chain = structure[0][chain_id]\n    \n        # getting residues out of chain is complex -- easier to get a list ahead of time.\n        residues = {}\n        for residue in chain: residues[ residue.id[1] ] = residue\n    \n        chain_seq = ''.join( [clean_res_name(residue.get_resname()) for residue in chain ] )\n        #print(chain_seq)\n    \n        # Initialize the result list\n        result = []\n    \n        # Counter for residue ID in reference sequence\n        ref_resid = 0\n        chain_idx = 0\n        for ref_res, chain_res in zip(alignment[0], alignment[1]):\n            if chain_res != '-': chain_idx += 1\n            if ref_res != '-':\n                ref_resid += 1\n                if chain_res == '-': # or chain_res == 'X':\n                    # Missing residue in chain or unknown residue\n                    result.append((ref_res, ref_resid, NULL_VALUE, NULL_VALUE, NULL_VALUE, -1e18))\n                else:\n                    # Find the corresponding residue in the chain\n                    try:\n                        #chain_seq_num = int(chain_seq_nums[ref_resid-1])\n                        chain_seq_num = int(chain_seq_nums[chain_idx-1])\n                        residue = residues[chain_seq_num]\n                        c1_prime = residue['C1\\'']\n                        coords = c1_prime.coord\n                        if residue.get_resname() != chain_res:\n                            print( 'WARNING!',ref_resid,chain_idx,chain_seq_num,residue.get_resname(),chain_res)\n                        result.append((ref_res, ref_resid, coords[0], coords[1], coords[2], residue.id[1]))\n                    except KeyError:\n                        # C1' atom not found\n                        result.append((ref_res, ref_resid, NULL_VALUE, NULL_VALUE, NULL_VALUE, -1e18))\n                    except Exception as e:\n                        # Any other error (e.g., residue not found)\n                        result.append((ref_res, ref_resid, NULL_VALUE, NULL_VALUE, NULL_VALUE, -1e18))\n    \n        return result\n    \n    def is_before_or_on(d1, d2):\n        date1 = pd.to_datetime(d1)\n        date2 = pd.to_datetime(d2)\n        return date1 <= date2\n    \n    def read_id_map(id_map_file):\n        if len(id_map_file)==0: return None\n        id_map = {}\n        try:\n            with open(id_map_file, newline='') as f:\n                reader = csv.DictReader(f)\n                if 'orig' not in reader.fieldnames or 'new' not in reader.fieldnames:\n                    print(\"Warning: ID map file does not contain the fields 'orig' and 'new'. Using original IDs instead.\")\n                    return id_map\n                for row in reader:\n                    id_map[row['orig']] = row['new']\n        except FileNotFoundError:\n            print(f\"Warning: ID map file {id_map_file} not found. Using original IDs instead.\", file=sys.stderr)\n        except Exception as exc:\n            print(f\"Error reading {id_map_file}: {exc}\", file=sys.stderr)\n        return id_map\n    \n    def read_release_dates( release_data_file ):\n        release_dates = {}\n        # must have format Entry ID, Release Date\n        with open(release_data_file, newline='') as f:\n            reader = csv.DictReader(f)\n            for row in reader:\n                release_dates[row['Entry ID']] = row['Release Date']\n    \n        return release_dates","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:46:01.915376Z","iopub.execute_input":"2026-01-09T05:46:01.915804Z","iopub.status.idle":"2026-01-09T05:46:01.93856Z","shell.execute_reply.started":"2026-01-09T05:46:01.915769Z","shell.execute_reply":"2026-01-09T05:46:01.937257Z"},"_kg_hide-input":true,"_kg_hide-output":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Setup","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"code","source":"if 'Model_1' in ensemble_of_Solutions: \n\n    # Prepare to collect output data\n    output_labels = []\n    \n    # Read the FASTA file\n    df = pd.read_csv( sequences_file )\n    targets = df['target_id'].to_list()\n    sequences = df['sequence'].to_list()\n    temporal_cutoffs = df['temporal_cutoff'].to_list()\n    \n    aln_lines = []\n    for line in open( mmseqs_results_file ).readlines():\n        # query,template,eval,qstart,qend,tstart,tend,qaln,taln\n        aln_lines.append( line.strip().split() )\n    \n    id_map = read_id_map( id_map_file )\n    \n    release_dates = read_release_dates( cif_dir + '/pdb_release_dates_NA.csv' )\n    \n    if start_idx == 0 and end_idx == 0: # do all targets by default\n        start_idx = 1\n        end_idx = len(targets)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:46:09.777079Z","iopub.execute_input":"2026-01-09T05:46:09.777396Z","iopub.status.idle":"2026-01-09T05:46:09.832947Z","shell.execute_reply.started":"2026-01-09T05:46:09.777377Z","shell.execute_reply":"2026-01-09T05:46:09.831862Z"},"_kg_hide-input":true,"_kg_hide-output":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if 'Model_1' in ensemble_of_Solutions: \n    \n    num_targets = 0\n    count = 0\n    for target,sequence,temporal_cutoff in zip(targets,sequences,temporal_cutoffs):\n        count += 1\n        if (count < start_idx) or (count > end_idx): continue\n    \n        # look for alignments and fill out C1' templates\n        templates = []\n        for aln_line in aln_lines:\n            if len(aln_line)!=9: continue # some kind of overflow in some alignments?\n    \n            query,template,eval,qstart,qend,tstart,tend,qaln,taln = aln_line\n    \n            if query != target: continue\n    \n            if int(qend)<int(qstart): continue # aligned to reverse complement!\n    \n            pdb_id,chain_id = template.split('_')\n    \n            # need to do alignment\n            cif_path = os.path.join(cif_dir, f'{pdb_id.lower()}.cif')\n            if not os.path.isfile( cif_path ): continue # occasional alignment to DNA, ignore!\n    \n            release_date = release_dates[pdb_id.upper()] # pulled from PDB server\n    \n            if CHECK_TEMPORAL_CUTOFF and is_before_or_on(temporal_cutoff,release_date): continue\n    \n            # these release dates in the CIF files can be buggy!\n            title,release_date_unreliable = extract_title_release_date( cif_path )\n    \n            print('\\n',target,temporal_cutoff,\"   \",template)\n            if title: print(f\"PDB Title: {title}\")\n            if release_date: print(f\"PDB Release Date: {release_date}\")\n    \n            # sometimes there is a mismatch between PDB's fasta files and what's actually stored in coordinates,\n            # so best to get the actual residue numbers for the chain\n            chain_full_sequence,chain_sequence,chain_seq_nums = extract_rna_sequence(cif_path,chain_id)\n    \n            # get 3d data\n            alignment = []\n            qstart=int(qstart)\n            qend=int(qend)\n            tstart=int(tstart)\n            tend=int(tend)\n            alignment.append( sequence[:(qstart-1)] + '-'*(tstart-1) + qaln + sequence[qend:]  )\n            alignment.append( '-'*(qstart-1)        + 'X'*(tstart-1) + taln + '-'*(len(sequence)-qend) )\n            print( alignment[0] )\n            print( alignment[1] )\n            c1prime_data = get_c1prime_labels( cif_path, chain_id, alignment, chain_seq_nums )\n    \n            # mismatch in FASTA sequence and the polyx info in the CIF file\n            if len(c1prime_data) != len(sequence):\n                print( 'WARNING! len(c1prime_data) != len(sequence)', 'len c1prime_data', len(c1prime_data), 'len sequence', len(sequence), 'qstart',qstart,'len qaln',len(qaln),'qend',qend)\n                continue\n    \n            templates.append( c1prime_data )\n    \n            if len(templates) >= MAX_TEMPLATES: break\n    \n        print( \"Found\", len(templates), \"templates for\", target,'\\n' )\n    \n        mapped_target = target\n        if not id_map is None: mapped_target = id_map[target]\n    \n        for i in range(len(sequence)):\n            output_label = {\n                \"ID\": f'{mapped_target}_{i+1}',\n                \"resname\": sequence[i],\n                \"resid\": i+1,\n            }\n    \n            # output templates\n            for n in range(len(templates)):\n                res,resid,x,y,z,pdb_seqnum = templates[n][i]\n                assert( resid == i+1 )\n                output_label[ f\"x_{n+1}\" ] = x\n                output_label[ f\"y_{n+1}\" ] = y\n                output_label[ f\"z_{n+1}\" ] = z\n    \n            # pad with blank models\n            for n in range(len(templates),MAX_TEMPLATES):\n                output_label[ f\"x_{n+1}\" ] = NULL_VALUE\n                output_label[ f\"y_{n+1}\" ] = NULL_VALUE\n                output_label[ f\"z_{n+1}\" ] = NULL_VALUE\n            output_labels.append( output_label )\n    \n        num_targets += 1\n        # if num_targets > 1: break # for debug!\n    \n    \n    print(f'Completed {num_targets} targets\\n')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:46:13.697613Z","iopub.execute_input":"2026-01-09T05:46:13.697977Z","iopub.status.idle":"2026-01-09T05:50:04.822189Z","shell.execute_reply.started":"2026-01-09T05:46:13.69795Z","shell.execute_reply":"2026-01-09T05:50:04.821117Z"},"_kg_hide-input":true,"_kg_hide-output":true,"jupyter":{"outputs_hidden":true,"source_hidden":true},"collapsed":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if 'Model_1' in ensemble_of_Solutions: \n    \n    # Create a DataFrame and write to CSV\n    \n    def output_csv( output_data, outfile ):\n        df = pd.DataFrame(output_data)\n        df.to_csv(outfile, index=False)\n        print(f\"Output written to {outfile}\")\n        return df\n    \n    df1 = output_csv( output_labels, outfile )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:50:53.258756Z","iopub.execute_input":"2026-01-09T05:50:53.259098Z","iopub.status.idle":"2026-01-09T05:50:53.436276Z","shell.execute_reply.started":"2026-01-09T05:50:53.259078Z","shell.execute_reply":"2026-01-09T05:50:53.43512Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **Model_2**","metadata":{}},{"cell_type":"code","source":"if 'Model_2' in ensemble_of_Solutions:\n    \n    # Simple baseline that generates A-form RNA helix coordinates with variations.\n    # Ready for immediate submission - no additional datasets required.\n    \n    \n    print('--------------------- cell.2.1')\n    \n    \n    import pandas as pd\n    import numpy as np\n    from scipy.spatial.transform import Rotation\n    \n    # A-form RNA helix parameters (Angstroms)\n    A_FORM_PARAMS = {\n        'rise'  : 2.8,    # Rise per residue\n        'twist' : 32.7,   # Twist angle (degrees)\n        'radius': 9.0     # Helix radius\n    }\n    \n    def generate_helix_coords(sequence, variation_idx=0):\n        \"\"\"Generate A-form helix C1' coordinates for an RNA sequence.\"\"\"\n        length = len(sequence)\n        np.random.seed(variation_idx * 42)\n        variation = 0.9 + 0.2 * np.random.random()\n        \n        params = {\n            'rise'  : A_FORM_PARAMS['rise'] * (0.9 + 0.2 * np.random.random()),\n            'twist' : np.deg2rad(A_FORM_PARAMS['twist'] * (0.9 + 0.2 * np.random.random())),\n            'radius': A_FORM_PARAMS['radius'] * variation\n        }\n        \n        coords = []\n        for i in range(length):\n            angle = i * params['twist']\n            x = params['radius'] * np.cos(angle)\n            y = params['radius'] * np.sin(angle)\n            z = i * params['rise']\n            coords.append([x, y, z])\n        \n        coords = np.array(coords)\n        \n        # Apply random rotation for variation\n        if variation_idx > 0:\n            rotation = Rotation.from_euler('xyz', \n                [np.random.uniform(-30, 30) for _ in range(3)], \n                degrees=True).as_matrix()\n            coords = coords @ rotation\n        \n        # Center the structure\n        coords = coords - np.mean(coords, axis=0)\n        return coords\n    \n    print(\"Functions defined.\")\n    \n    \n    print('--------------------- cell.2.2')\n    \n    \n    # Load data\n    sample = pd.read_csv('/kaggle/input/stanford-rna-3d-folding-2/sample_submission.csv')\n    test_df = pd.read_csv('/kaggle/input/stanford-rna-3d-folding-2/test_sequences.csv')\n    \n    # Build sequence lookup\n    seq_lookup = dict(zip(test_df['target_id'], test_df['sequence']))\n    \n    print(f\"Loaded {len(test_df)} sequences\")\n    print(f\"Total residues: {len(sample)}\")\n    \n    \n    print('--------------------- cell.2.3')\n    \n    \n    # Generate predictions\n    all_data = []\n    current_target = None\n    current_models = None\n    \n    for idx, row in sample.iterrows():\n        id_parts = row['ID'].rsplit('_', 1)\n        target_id = id_parts[0]\n        resid = int(id_parts[1])\n        \n        # Generate models for new target\n        if target_id != current_target:\n            current_target = target_id\n            sequence = seq_lookup[target_id]\n            current_models = [generate_helix_coords(sequence, i) for i in range(5)]\n            print(f\"Processing {target_id} (len={len(sequence)})\")\n        \n        # Get coordinates for this residue\n        i = resid - 1\n        row_data = {\n            'ID': row['ID'],\n            'resname': row['resname'],\n            'resid': row['resid']\n        }\n        \n        for model_idx, model_coords in enumerate(current_models, 1):\n            row_data[f'x_{model_idx}'] = float(model_coords[i, 0])\n            row_data[f'y_{model_idx}'] = float(model_coords[i, 1])\n            row_data[f'z_{model_idx}'] = float(model_coords[i, 2])\n        \n        all_data.append(row_data)\n    \n    print(f\"\\nProcessed {len(all_data)} residues\")\n    \n    \n    print('--------------------- cell.2.4')\n    \n    \n    # Create submission DataFrame\n    submission_df = pd.DataFrame(all_data)\n    \n    # Ensure correct column order\n    column_order = ['ID', 'resname', 'resid']\n    for i in range(1, 6):\n        column_order.extend([f'x_{i}', f'y_{i}', f'z_{i}'])\n    \n    submission_df = submission_df[column_order]\n    \n    # Save submission\n    submission_df.to_csv('subm_2.csv', index=False, float_format='%.3f')\n    \n    # Validate\n    print(f\"Submission shape: {submission_df.shape}\")\n    print(f\"ID match: {(submission_df['ID'] == sample['ID']).all()}\")\n    print(f\"resname match: {(submission_df['resname'] == sample['resname']).all()}\")\n    print(f\"resid match: {(submission_df['resid'] == sample['resid']).all()}\")\n    \n    submission_df.head(10)\n    \n    df2 = submission_df","metadata":{"trusted":true,"jupyter":{"outputs_hidden":true,"source_hidden":true},"execution":{"iopub.status.busy":"2026-01-09T05:52:30.828466Z","iopub.execute_input":"2026-01-09T05:52:30.828824Z","iopub.status.idle":"2026-01-09T05:52:32.125416Z","shell.execute_reply.started":"2026-01-09T05:52:30.828795Z","shell.execute_reply":"2026-01-09T05:52:32.124272Z"},"_kg_hide-input":true,"_kg_hide-output":true,"collapsed":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **Model_3**","metadata":{}},{"cell_type":"code","source":"if 'Model_3' in ensemble_of_Solutions: \n        \n    print('--------------------- cell.3.1')\n    \n    import os\n    import random\n    import tempfile\n    import subprocess\n    import numpy as np\n    import pandas as pd\n    \n    from scipy.spatial.transform import Rotation\n    \n    # an attempt to work from an article\n    \n    class TemplateBasedModeler:\n        \n        def __init__(self):\n            self.a_form_params = {\n                'rise': 2.8,\n                'twist': 32.7,\n                'radius': 9.0\n            }\n        \n        def search_templates(self, sequence, max_templates=5):\n            templates = []\n            length = len(sequence)\n            \n            for i in range(max_templates):\n                variation = 0.8 + 0.4 * random.random()\n                params = {\n                    'rise': self.a_form_params['rise'] * (0.9 + 0.2 * random.random()),\n                    'twist': np.deg2rad(self.a_form_params['twist'] * (0.9 + 0.2 * random.random())),\n                    'radius': self.a_form_params['radius'] * variation\n                }\n                \n                coords = []\n                for j in range(length):\n                    angle = j * params['twist']\n                    x = params['radius'] * np.cos(angle)\n                    y = params['radius'] * np.sin(angle)\n                    z = j * params['rise']\n                    coords.append([x, y, z])\n                \n                template = {\n                    'id': f'template_{i}',\n                    'coords': np.array(coords),\n                    'confidence': 0.9 - i*0.15\n                }\n                templates.append(template)\n            \n            return templates\n        \n        def align_templates(self, templates, sequence):\n            aligned = []\n            \n            for template in templates:\n                coords = template['coords']\n                centered = coords - np.mean(coords, axis=0)\n                \n                if random.random() > 0.3:\n                    rotation = Rotation.random().as_matrix()\n                    centered = centered @ rotation\n                \n                template['aligned_coords'] = centered\n                aligned.append(template)\n            \n            aligned.sort(key=lambda x: x['confidence'], reverse=True)\n            return aligned[:5]\n    \n    class RNAProPredictor:\n        \n        def __init__(self):\n            self.template_modeler = TemplateBasedModeler()\n            self.use_tbm = True\n            \n        def _generate_helix(self, length, params):\n            coords = []\n            for i in range(length):\n                angle = i * params['twist']\n                x = params['radius'] * np.cos(angle)\n                y = params['radius'] * np.sin(angle)\n                z = i * params['rise']\n                coords.append([x, y, z])\n            return np.array(coords)\n        \n        def _add_secondary_structure(self, sequence, coords):\n            length = len(sequence)\n            for i in range(length - 4):\n                pairs = [('G', 'C'), ('C', 'G'), ('A', 'U'), ('U', 'A')]\n                if (sequence[i], sequence[i+4]) in pairs:\n                    noise = np.random.normal(0, 0.2, 3)\n                    coords[i] += noise\n                    coords[i+4] -= noise * 0.8\n            return coords\n        \n        def _generate_de_novo_model(self, sequence):\n            length = len(sequence)\n            variation = 0.8 + 0.4 * random.random()\n            params = {\n                'rise': 2.8 * (0.9 + 0.2 * random.random()),\n                'twist': np.deg2rad(32.7 * (0.9 + 0.2 * random.random())),\n                'radius': 9.0 * variation\n            }\n            \n            base_coords = self._generate_helix(length, params)\n            base_coords = self._add_secondary_structure(sequence, base_coords)\n            \n            if random.random() > 0.7:\n                rotation = Rotation.from_euler('xyz', \n                    [random.uniform(-30, 30) for _ in range(3)], degrees=True).as_matrix()\n                base_coords = base_coords @ rotation\n            \n            return base_coords\n        \n        def _generate_variations(self, base_model, sequence, num_variations):\n            variations = []\n            length = len(sequence)\n            \n            for i in range(num_variations):\n                if i == 0:\n                    model = base_model.copy()\n                else:\n                    model = base_model.copy()\n                    scale = 0.1 + i * 0.05\n                    \n                    for j in range(length):\n                        noise = np.random.normal(0, scale, 3)\n                        model[j] += noise\n                    \n                    if i >= 2:\n                        rotation = Rotation.random().as_matrix()\n                        model = model @ rotation\n                \n                variations.append(model)\n            \n            return variations\n        \n        def predict_ensemble(self, sequence, num_models=5):\n            models = []\n            \n            if self.use_tbm and len(sequence) > 10:\n                templates = self.template_modeler.search_templates(sequence)\n                aligned_templates = self.template_modeler.align_templates(templates, sequence)\n                \n                for i, template in enumerate(aligned_templates[:min(3, len(aligned_templates))]):\n                    base = template['aligned_coords']\n                    vars_from_template = self._generate_variations(base, sequence, 2)\n                    models.extend(vars_from_template)\n            \n            if len(models) < num_models:\n                de_novo_base = self._generate_de_novo_model(sequence)\n                de_novo_vars = self._generate_variations(de_novo_base, sequence, \n                                                        num_models - len(models))\n                models.extend(de_novo_vars)\n            \n            return models[:num_models]\n    \n    def create_kaggle_submission(test_file, output_file='submission.csv'):\n        predictor = RNAProPredictor()\n        test_df = pd.read_csv(test_file)\n        \n        all_data = []\n        \n        for _, row in test_df.iterrows():\n            target_id = row['target_id']\n            sequence = row['sequence']\n            \n            models = predictor.predict_ensemble(sequence, 5)\n            \n            for i, nucleotide in enumerate(sequence):\n                residue_id = f\"{target_id}_{i+1}\"\n                row_data = {\n                    'ID': residue_id,\n                    'resname': nucleotide,\n                    'resid': i + 1\n                }\n                \n                for model_idx, model_coords in enumerate(models, 1):\n                    if model_idx <= 5:\n                        row_data[f'x_{model_idx}'] = float(model_coords[i, 0])\n                        row_data[f'y_{model_idx}'] = float(model_coords[i, 1])\n                        row_data[f'z_{model_idx}'] = float(model_coords[i, 2])\n                \n                all_data.append(row_data)\n        \n        submission_df = pd.DataFrame(all_data)\n        \n        column_order = ['ID', 'resname', 'resid']\n        for i in range(1, 6):\n            column_order.extend([f'x_{i}', f'y_{i}', f'z_{i}'])\n        \n        submission_df = submission_df[column_order]\n        submission_df.to_csv(output_file, index=False, float_format='%.3f')\n        \n        return submission_df\n    \n    submission = create_kaggle_submission(\n        '/kaggle/input/stanford-rna-3d-folding-2/test_sequences.csv',\n        'subm_3.csv'\n    )\n    submission.head()\n    \n    df3 = submission","metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2026-01-09T05:52:48.062539Z","iopub.execute_input":"2026-01-09T05:52:48.063539Z","iopub.status.idle":"2026-01-09T05:52:48.724359Z","shell.execute_reply.started":"2026-01-09T05:52:48.063509Z","shell.execute_reply":"2026-01-09T05:52:48.723137Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Model_4","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **Model_5**","metadata":{}},{"cell_type":"code","source":"if 'Model_5' in ensemble_of_Solutions: \n    \n    import os\n    import numpy as np\n    import pandas as pd\n    from tqdm import tqdm\n    \n    DATA_DIR = \"/kaggle/input/stanford-rna-3d-folding-2\" # Update path based on environment\n    SUBMISSION_FILE = \"subm_5.csv\"\n    TEST_FILE = os.path.join(DATA_DIR, \"test_sequences.csv\")\n    MSA_DIR = os.path.join(DATA_DIR, \"MSA\")\n    \n    # --- 2. MODEL DEFINITION ---\n    class RNAStructurePredictor:\n        \"\"\"\n        Placeholder for your RNA Folding Model.\n        In a real scenario, this would load weights from a pre-trained model \n        and use the MSA and Sequence as input features.\n        \"\"\"\n        def predict(self, sequence, msa_path=None):\n            # sequence: string of ACGU\n            # msa_path: path to the .fasta MSA file for this target\n            L = len(sequence)\n            \n            # We need to return 5 sets of coordinates (5 predictions)\n            # Shape: (5, L, 3) -> 5 models, L residues, (x, y, z)\n            # Here we initialize with zeros or random numbers as a baseline\n            preds = np.random.normal(loc=0, scale=10, size=(5, L, 3))\n            \n            # NOTE: In a real model, you'd process the MSA here:\n            # if msa_path and os.path.exists(msa_path):\n            #     msa_data = load_msa(msa_path)\n            #     preds = my_model.inference(sequence, msa_data)\n            \n            return preds\n    \n    # --- 3. PIPELINE EXECUTION ---\n    def run_inference():\n        # Check if test file exists (Kaggle hidden test set)\n        if not os.path.exists(TEST_FILE):\n            # Fallback for local testing or validation\n            TEST_FILE_ALT = os.path.join(DATA_DIR, \"validation_sequences.csv\")\n            test_df = pd.read_csv(TEST_FILE_ALT)\n        else:\n            test_df = pd.read_csv(TEST_FILE)\n    \n        predictor = RNAStructurePredictor()\n        all_submission_rows = []\n    \n        print(f\"Starting inference on {len(test_df)} sequences...\")\n    \n        for _, row in tqdm(test_df.iterrows(), total=len(test_df)):\n            target_id = row['target_id']\n            sequence = row['sequence']\n            msa_path = os.path.join(MSA_DIR, f\"{target_id}.MSA.fasta\")\n            \n            # Generate 5 structures for the entire sequence\n            # shape: (5, L, 3)\n            coords_5_models = predictor.predict(sequence, msa_path)\n            \n            # Unpack predictions into the CSV format\n            # Format: ID, resname, resid, x_1, y_1, z_1, ..., x_5, y_5, z_5\n            for i, resname in enumerate(sequence):\n                resid = i + 1\n                row_id = f\"{target_id}_{resid}\"\n                \n                # Extract (x,y,z) for this specific residue across all 5 models\n                res_coords = []\n                for m_idx in range(5):\n                    x, y, z = coords_5_models[m_idx, i, :]\n                    res_coords.extend([x, y, z])\n                \n                all_submission_rows.append([row_id, resname, resid] + res_coords)\n    \n        # --- 4. FORMATTING & CLIPPING ---\n        columns = ['ID', 'resname', 'resid']\n        for i in range(1, 6):\n            columns += [f'x_{i}', f'y_{i}', f'z_{i}']\n    \n        sub_df = pd.DataFrame(all_submission_rows, columns=columns)\n        # CRITICAL: Clip coordinates to prevent PDB format errors (-999.999 to 9999.999)\n        coord_cols = [c for c in sub_df.columns if c.startswith(('x_', 'y_', 'z_'))]\n        sub_df[coord_cols] = sub_df[coord_cols].clip(lower=-999.999, upper=9999.999)\n        # Save\n        sub_df.to_csv(SUBMISSION_FILE, index=False)\n        print(f\"Submission saved to {SUBMISSION_FILE}\")\n        return sub_df\n    \n    df5 = run_inference()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:53:31.780112Z","iopub.execute_input":"2026-01-09T05:53:31.780419Z","iopub.status.idle":"2026-01-09T05:53:32.227227Z","shell.execute_reply.started":"2026-01-09T05:53:31.780398Z","shell.execute_reply":"2026-01-09T05:53:32.226141Z"},"_kg_hide-input":true,"_kg_hide-output":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **Model_6**","metadata":{}},{"cell_type":"code","source":"if 'Model_6' in ensemble_of_Solutions: \n    \n    import os\n    import math\n    import numpy as np\n    import pandas as pd\n    \n    K_CONTEXT = 2\n    SMOOTH_WIN = 9\n    NOISE_STD = 0.10\n    BEND_ANGLE_STD = 0.035\n    BEND_CORR = 0.992\n    SEEDS = (1, 2, 3, 4, 5)\n    \n    CLIP_MIN = -999.999\n    CLIP_MAX = 9999.999\n    \n    def pad_seq(seq, k=2, pad=\"N\"):\n        return pad * k + seq + pad * k\n    \n    def rot_from_axis_angle(axis, angle):\n        ax = np.asarray(axis, dtype=np.float64)\n        n = np.linalg.norm(ax) + 1e-12\n        ax = ax / n\n        x, y, z = ax\n        c = math.cos(angle)\n        s = math.sin(angle)\n        C = 1.0 - c\n        return np.array([\n            [c + x*x*C, x*y*C - z*s, x*z*C + y*s],\n            [y*x*C + z*s, c + y*y*C, y*z*C - x*s],\n            [z*x*C - y*s, z*y*C + x*s, c + z*z*C]\n        ], dtype=np.float64)\n    \n    def smooth_steps(steps, win):\n        if win <= 1:\n            return steps\n        pad = win // 2\n        s = np.pad(steps, ((pad, pad), (0, 0)), mode=\"edge\")\n        out = np.empty_like(steps)\n        for i in range(len(steps)):\n            out[i] = s[i:i + win].mean(axis=0)\n        return out\n    \n    def bend_steps(steps, rng, angle_std, corr):\n        steps = steps.astype(np.float64, copy=True)\n        R = np.eye(3, dtype=np.float64)\n        w = np.zeros(3, dtype=np.float64)\n        for i in range(len(steps)):\n            w = corr * w + rng.normal(0.0, angle_std, size=3)\n            ang = float(np.linalg.norm(w))\n            if ang > 1e-12:\n                R = R @ rot_from_axis_angle(w, ang)\n            steps[i] = (R @ steps[i]).ravel()\n        return steps.astype(np.float32)\n    \n    def integrate(steps):\n        return np.cumsum(steps, axis=0)\n    \n    def fast_target_id_from_ID(ids: pd.Series) -> pd.Series:\n        s = ids.astype(\"string\")\n        return s.str.rsplit(\"_\", n=1, expand=True)[0]\n    \n    def detect_k_cols(path):\n        head = pd.read_csv(path, nrows=0)\n        ks = []\n        for c in head.columns:\n            if c.startswith(\"x_\"):\n                try:\n                    ks.append(int(c.split(\"_\")[1]))\n                except:\n                    pass\n        return max(ks) if ks else 1\n    \n    def load_labels(path):\n        K = detect_k_cols(path)\n        usecols = [\"ID\", \"resid\"]\n        for k in range(1, K + 1):\n            usecols += [f\"x_{k}\", f\"y_{k}\", f\"z_{k}\"]\n        try:\n            df = pd.read_csv(\n                path,\n                usecols=usecols,\n                dtype={\"ID\": \"string\", \"resid\": \"int32\"},\n                engine=\"pyarrow\",\n            )\n        except Exception:\n            df = pd.read_csv(\n                path,\n                usecols=usecols,\n                dtype={\"ID\": \"string\", \"resid\": \"int32\"},\n            )\n        df[\"target_id\"] = fast_target_id_from_ID(df[\"ID\"])\n        return df, K\n    \n    def build_step_stats(train_sequences_csv, train_labels_csv):\n        seq_df = pd.read_csv(train_sequences_csv, usecols=[\"target_id\", \"sequence\"])\n        seq_map = dict(zip(seq_df[\"target_id\"].values, seq_df[\"sequence\"].values))\n    \n        lab, K = load_labels(train_labels_csv)\n        lab = lab[lab[\"target_id\"].isin(seq_map.keys())].copy()\n        lab = lab.sort_values([\"target_id\", \"resid\"])\n    \n        sums5, cnt5 = {}, {}\n        sums3, cnt3 = {}, {}\n        sums1, cnt1 = {}, {}\n        sumg = np.zeros(3, dtype=np.float64)\n        cntg = 0\n    \n        for tid, g in lab.groupby(\"target_id\", sort=False):\n            seq = seq_map.get(tid)\n            if not seq:\n                continue\n            L = len(seq)\n            pseq = pad_seq(seq, k=K_CONTEXT)\n    \n            resid_all = g[\"resid\"].to_numpy(np.int32)\n            ok = (resid_all >= 1) & (resid_all <= L)\n            if not np.any(ok):\n                continue\n    \n            g = g.iloc[np.flatnonzero(ok)]\n            resid_all = resid_all[ok]\n            order = np.argsort(resid_all)\n            g = g.iloc[order]\n            resid_all = resid_all[order]\n    \n            for k in range(1, K + 1):\n                x = g.get(f\"x_{k}\")\n                y = g.get(f\"y_{k}\")\n                z = g.get(f\"z_{k}\")\n                if x is None or y is None or z is None:\n                    continue\n    \n                xv = x.to_numpy(np.float32, copy=False)\n                yv = y.to_numpy(np.float32, copy=False)\n                zv = z.to_numpy(np.float32, copy=False)\n    \n                obs = ~np.isnan(xv) & ~np.isnan(yv) & ~np.isnan(zv)\n                if not np.any(obs):\n                    continue\n    \n                pos = resid_all[obs] - 1\n                coords = np.stack([xv[obs], yv[obs], zv[obs]], axis=1).astype(np.float32)\n    \n                steps = np.zeros_like(coords, dtype=np.float32)\n                if len(pos) > 1:\n                    cont = (pos[1:] == pos[:-1] + 1)\n                    steps[1:][cont] = coords[1:][cont] - coords[:-1][cont]\n    \n                for j in range(len(pos)):\n                    i = int(pos[j])\n                    v = steps[j].astype(np.float64)\n    \n                    k5 = pseq[i:i + 5]\n                    k3 = pseq[i + 1:i + 4]\n                    k1 = pseq[i + 2:i + 3]\n    \n                    if k5 in sums5:\n                        sums5[k5] += v; cnt5[k5] += 1\n                    else:\n                        sums5[k5] = v.copy(); cnt5[k5] = 1\n    \n                    if k3 in sums3:\n                        sums3[k3] += v; cnt3[k3] += 1\n                    else:\n                        sums3[k3] = v.copy(); cnt3[k3] = 1\n    \n                    if k1 in sums1:\n                        sums1[k1] += v; cnt1[k1] += 1\n                    else:\n                        sums1[k1] = v.copy(); cnt1[k1] = 1\n    \n                    sumg += v\n                    cntg += 1\n    \n        mean5 = {k: (sums5[k] / cnt5[k]).astype(np.float32) for k in sums5}\n        mean3 = {k: (sums3[k] / cnt3[k]).astype(np.float32) for k in sums3}\n        mean1 = {k: (sums1[k] / cnt1[k]).astype(np.float32) for k in sums1}\n        meang = (sumg / max(cntg, 1)).astype(np.float32)\n        return mean5, mean3, mean1, meang\n    \n    def predict_steps(seq, mean5, mean3, mean1, meang):\n        L = len(seq)\n        pseq = pad_seq(seq, k=K_CONTEXT)\n        out = np.zeros((L, 3), dtype=np.float32)\n        for i in range(L):\n            k5 = pseq[i:i + 5]\n            v = mean5.get(k5)\n            if v is not None:\n                out[i] = v; continue\n            k3 = pseq[i + 1:i + 4]\n            v = mean3.get(k3)\n            if v is not None:\n                out[i] = v; continue\n            k1 = pseq[i + 2:i + 3]\n            v = mean1.get(k1)\n            if v is not None:\n                out[i] = v; continue\n            out[i] = meang\n        out[0] = 0.0\n        return out\n    \n    def make_submission(test_sequences_csv, mean5, mean3, mean1, meang, out_path):\n        test_df = pd.read_csv(test_sequences_csv, usecols=[\"target_id\", \"sequence\"])\n        rows = []\n    \n        for _, r in test_df.iterrows():\n            tid = r[\"target_id\"]\n            seq = r[\"sequence\"]\n            L = len(seq)\n    \n            base = predict_steps(seq, mean5, mean3, mean1, meang)\n            variants = [base]\n    \n            for s in SEEDS[1:]:\n                rng = np.random.default_rng(10_000 + s)\n                st = bend_steps(base, rng, angle_std=BEND_ANGLE_STD, corr=BEND_CORR)\n                st = st + rng.normal(0.0, NOISE_STD, size=st.shape).astype(np.float32)\n                st = smooth_steps(st, SMOOTH_WIN)\n                st[0] = 0.0\n                variants.append(st)\n    \n            coords = [integrate(st) for st in variants]\n            coords = [np.clip(c, CLIP_MIN, CLIP_MAX) for c in coords]\n    \n            for i in range(L):\n                row = {\"ID\": f\"{tid}_{i+1}\", \"resname\": seq[i], \"resid\": i + 1}\n                for k in range(5):\n                    row[f\"x_{k+1}\"] = float(coords[k][i, 0])\n                    row[f\"y_{k+1}\"] = float(coords[k][i, 1])\n                    row[f\"z_{k+1}\"] = float(coords[k][i, 2])\n                rows.append(row)\n    \n        sub = pd.DataFrame(rows)\n        sub.to_csv(out_path, index=False)\n        print(\"saved:\", out_path, sub.shape)\n        return sub\n    \n    train_sequences_csv = \"/kaggle/input/stanford-rna-3d-folding-2/train_sequences.csv\"\n    train_labels_csv    = \"/kaggle/input/stanford-rna-3d-folding-2/train_labels.csv\"\n    test_sequences_csv  = \"/kaggle/input/stanford-rna-3d-folding-2/test_sequences.csv\"\n    out_path            = \"/kaggle/working/subm_6.csv\"\n\n    mean5, mean3, mean1, meang = build_step_stats(train_sequences_csv, train_labels_csv)\n    \n    df6 = make_submission(test_sequences_csv, mean5, mean3, mean1, meang, out_path)","metadata":{"trusted":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2026-01-09T05:53:52.243952Z","iopub.execute_input":"2026-01-09T05:53:52.244292Z","iopub.status.idle":"2026-01-09T05:53:52.27994Z","shell.execute_reply.started":"2026-01-09T05:53:52.244269Z","shell.execute_reply":"2026-01-09T05:53:52.278554Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **Model_7**","metadata":{}},{"cell_type":"code","source":"if 'Model_7' in ensemble_of_Solutions: \n    \n    import os\n    import numpy as np\n    import pandas as pd\n    from tqdm import tqdm\n    \n    # --- 1. CONFIGURATION ---\n    DATA_DIR = \"/kaggle/input/stanford-rna-3d-folding-2\" # Update path based on environment\n    SUBMISSION_FILE = \"subm_7.csv\"\n    TEST_FILE = os.path.join(DATA_DIR, \"test_sequences.csv\")\n    MSA_DIR = os.path.join(DATA_DIR, \"MSA\")\n    \n    # --- 2. MODEL DEFINITION ---\n    class RNAStructurePredictor:\n        \"\"\"\n        Placeholder for your RNA Folding Model.\n        In a real scenario, this would load weights from a pre-trained model \n        and use the MSA and Sequence as input features.\n        \"\"\"\n        def predict(self, sequence, msa_path=None):\n            # sequence: string of ACGU\n            # msa_path: path to the .fasta MSA file for this target\n            L = len(sequence)\n            \n            # We need to return 5 sets of coordinates (5 predictions)\n            # Shape: (5, L, 3) -> 5 models, L residues, (x, y, z)\n            # Here we initialize with zeros or random numbers as a baseline\n            preds = np.random.normal(loc=0, scale=10, size=(5, L, 3))\n            \n            # NOTE: In a real model, you'd process the MSA here:\n            # if msa_path and os.path.exists(msa_path):\n            #     msa_data = load_msa(msa_path)\n            #     preds = my_model.inference(sequence, msa_data)\n            \n            return preds\n    \n    # --- 3. PIPELINE EXECUTION ---\n    def run_inference():\n        # Check if test file exists (Kaggle hidden test set)\n        if not os.path.exists(TEST_FILE):\n            # Fallback for local testing or validation\n            TEST_FILE_ALT = os.path.join(DATA_DIR, \"validation_sequences.csv\")\n            test_df = pd.read_csv(TEST_FILE_ALT)\n        else:\n            test_df = pd.read_csv(TEST_FILE)\n    \n        predictor = RNAStructurePredictor()\n        all_submission_rows = []\n    \n        print(f\"Starting inference on {len(test_df)} sequences...\")\n    \n        for _, row in tqdm(test_df.iterrows(), total=len(test_df)):\n            target_id = row['target_id']\n            sequence = row['sequence']\n            msa_path = os.path.join(MSA_DIR, f\"{target_id}.MSA.fasta\")\n            \n            # Generate 5 structures for the entire sequence\n            # shape: (5, L, 3)\n            coords_5_models = predictor.predict(sequence, msa_path)\n            \n            # Unpack predictions into the CSV format\n            # Format: ID, resname, resid, x_1, y_1, z_1, ..., x_5, y_5, z_5\n            for i, resname in enumerate(sequence):\n                resid = i + 1\n                row_id = f\"{target_id}_{resid}\"\n                \n                # Extract (x,y,z) for this specific residue across all 5 models\n                res_coords = []\n                for m_idx in range(5):\n                    x, y, z = coords_5_models[m_idx, i, :]\n                    res_coords.extend([x, y, z])\n                \n                all_submission_rows.append([row_id, resname, resid] + res_coords)\n    \n        # --- 4. FORMATTING & CLIPPING ---\n        columns = ['ID', 'resname', 'resid']\n        for i in range(1, 6):\n            columns += [f'x_{i}', f'y_{i}', f'z_{i}']\n    \n        df_subm3 = pd.DataFrame(all_submission_rows, columns=columns)\n        # CRITICAL: Clip coordinates to prevent PDB format errors (-999.999 to 9999.999)\n        coord_cols = [c for c in df_subm3.columns if c.startswith(('x_', 'y_', 'z_'))]\n        df_subm3[coord_cols] = df_subm3[coord_cols].clip(lower=-999.999, upper=9999.999)\n        # Save\n        df_subm3.to_csv(SUBMISSION_FILE, index=False)\n        print(f\"Submission saved to {SUBMISSION_FILE}\")\n        return df_subm3\n        \n    \n    df7 = run_inference()","metadata":{"trusted":true,"_kg_hide-output":true,"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2026-01-09T05:53:59.325557Z","iopub.execute_input":"2026-01-09T05:53:59.325954Z","iopub.status.idle":"2026-01-09T05:53:59.338135Z","shell.execute_reply.started":"2026-01-09T05:53:59.325928Z","shell.execute_reply":"2026-01-09T05:53:59.336898Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **Model_8**","metadata":{}},{"cell_type":"code","source":"if 'Model_8' in ensemble_of_Solutions: \n     \n    import numpy as np\n    import pandas as pd\n    from pathlib import Path\n    \n    # Constants\n    HELIX_RISE      =  2.81\n    HELIX_TWIST_DEG = 32.77\n    HELIX_RADIUS    = 10.21\n    \n    print(\"=\" * 21)\n    print(\"Stanford RNA 3D Folding Part 2 - Baseline\")\n    print(\"=\" * 21)\n    \n    \n    def random_rotation(coords: np.ndarray, seed: int) -> np.ndarray:\n        \"\"\"Apply random rotation.\"\"\"\n        rng = np.random.RandomState(seed)\n        alpha, beta, gamma = rng.uniform(0, 2*np.pi, 3)\n    \n        ca, sa = np.cos(alpha), np.sin(alpha)\n        cb, sb = np.cos(beta), np.sin(beta)\n        cg, sg = np.cos(gamma), np.sin(gamma)\n    \n        R = np.array([\n            [cg*cb*ca - sg*sa, -cg*cb*sa - sg*ca, cg*sb],\n            [sg*cb*ca + cg*sa, -sg*cb*sa + cg*ca, sg*sb],\n            [-sb*ca, sb*sa, cb]\n        ], dtype=np.float32)\n    \n        return (coords - coords.mean(axis=0)) @ R.T\n    \n    \n    def predict_ensemble(sequence: str, num_models: int = 5) -> list:\n        \"\"\"Generate ensemble of diverse predictions.\"\"\"\n        sequence = sequence.upper().replace('T', 'U')\n    \n        models = []\n        for i in range(num_models):\n            coords = simple_fold(sequence, seed=i * 7919)\n            if i > 0:\n                coords = random_rotation(coords, seed=i * 17)\n            models.append(coords)\n    \n        return models\n    \n    \n    def simple_fold(sequence: str, seed: int = 42) -> np.ndarray:\n        \"\"\"Simple helical folding.\"\"\"\n        n = len(sequence)\n        rng = np.random.RandomState(seed)\n    \n        coords = np.zeros((n, 3), dtype=np.float32)\n        twist_rad = np.radians(HELIX_TWIST_DEG)\n    \n        for i in range(n):\n            angle = i * twist_rad\n            coords[i] = [\n                HELIX_RADIUS * np.cos(angle),\n                HELIX_RADIUS * np.sin(angle),\n                i * HELIX_RISE\n            ]\n    \n        coords += rng.randn(n, 3).astype(np.float32) * 0.5\n        coords -= coords.mean(axis=0)\n    \n        return coords\n    \n    \n    # Load test data\n    INPUT_DIR = Path(\"/kaggle/input/stanford-rna-3d-folding-2\")\n    test_path = INPUT_DIR / \"test_sequences.csv\"\n    \n    print(f\"Looking for test data at: {test_path}\")\n    \n    if test_path.exists():\n        test_df = pd.read_csv(test_path)\n        print(f\"Loaded {len(test_df)} test sequences\")\n    else:\n        print(f\"Error: Test file not found at {test_path}\")\n        # List available files\n        print(f\"Available in input dir: {list(INPUT_DIR.glob('*'))}\")\n        raise FileNotFoundError(f\"Test file not found at {test_path}\")\n    \n    # Generate predictions\n    rows = []\n    \n    for idx, row in test_df.iterrows():\n        target_id = row['target_id']\n        sequence = str(row['sequence'])\n    \n        # Clean sequence\n        if '\\n' in sequence:\n            sequence = sequence.split('\\n')[0]\n        sequence = ''.join(c for c in sequence.upper() if c in 'AUGC')\n    \n        if len(sequence) == 0:\n            print(f\"Warning: Empty sequence for {target_id}\")\n            continue\n    \n        if idx % 5 == 0:\n            print(f\"[{idx+1}/{len(test_df)}] {target_id}: {len(sequence)} nt\")\n    \n        models = predict_ensemble(sequence, num_models=5)\n    \n        for resid, nucleotide in enumerate(sequence, start=1):\n            row_data = {\n                'ID': f\"{target_id}_{resid}\",\n                'resname': nucleotide,\n                'resid': resid\n            }\n    \n            for m_idx, coords in enumerate(models):\n                i = resid - 1\n                if i < len(coords):\n                    row_data[f'x_{m_idx+1}'] = round(float(coords[i, 0]), 3)\n                    row_data[f'y_{m_idx+1}'] = round(float(coords[i, 1]), 3)\n                    row_data[f'z_{m_idx+1}'] = round(float(coords[i, 2]), 3)\n                else:\n                    row_data[f'x_{m_idx+1}'] = 0.0\n                    row_data[f'y_{m_idx+1}'] = 0.0\n                    row_data[f'z_{m_idx+1}'] = 0.0\n    \n            rows.append(row_data)\n    \n    # Create submission\n    cols = ['ID', 'resname', 'resid']\n    for i in range(1, 6):\n        cols.extend([f'x_{i}', f'y_{i}', f'z_{i}'])\n    \n    df8 = pd.DataFrame(rows)[cols]\n    \n    df8.to_csv(\"/kaggle/working/subm_8.csv\", index=False)","metadata":{"trusted":true,"_kg_hide-output":true,"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2026-01-09T05:54:04.025079Z","iopub.execute_input":"2026-01-09T05:54:04.025399Z","iopub.status.idle":"2026-01-09T05:54:04.044938Z","shell.execute_reply.started":"2026-01-09T05:54:04.025374Z","shell.execute_reply":"2026-01-09T05:54:04.043668Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Helper functions","metadata":{}},{"cell_type":"code","source":"def sample_submission():\n    return pd.read_csv('/kaggle/input/stanford-rna-3d-folding-2/sample_submission.csv')\n\ndef weights(Ens_dict=ensemble_of_Solutions):\n    list_weight = list(Ens_dict.values())\n    w1 = list_weight[0] if len(list_weight) >= 1 else 0\n    w2 = list_weight[1] if len(list_weight) >= 2 else 0\n    w3 = list_weight[2] if len(list_weight) >= 3 else 0\n    w4 = list_weight[3] if len(list_weight) >= 4 else 0\n    w5 = list_weight[4] if len(list_weight) >= 5 else 0\n    return w1,w2,w3,w4,w5","metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2026-01-09T05:54:27.4832Z","iopub.execute_input":"2026-01-09T05:54:27.483591Z","iopub.status.idle":"2026-01-09T05:54:27.49031Z","shell.execute_reply.started":"2026-01-09T05:54:27.483565Z","shell.execute_reply":"2026-01-09T05:54:27.489377Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"d1,d2,d3,d4,d5 = df2,df3,df5,df8,sample_submission()\n\nw1,w2,w3,w4,w5 = weights()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:54:49.765772Z","iopub.execute_input":"2026-01-09T05:54:49.766066Z","iopub.status.idle":"2026-01-09T05:54:49.787194Z","shell.execute_reply.started":"2026-01-09T05:54:49.766046Z","shell.execute_reply":"2026-01-09T05:54:49.785886Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"x1,y1,z1 = 'x_1','y_1','z_1'\nx2,y2,z2 = 'x_2','y_2','z_2'\nx3,y3,z3 = 'x_3','y_3','z_3'\nx4,y4,z4 = 'x_4','y_4','z_5'\nx5,y5,z5 = 'x_5','y_4','z_5'\n\ndf = sample_submission()\n\ndf[x1] = d1[x1]*w1 + d2[x1]*w2 + d3[x1]*w3 + d4[x1]*w4 + d5[x1]*w5\ndf[y1] = d1[y1]*w1 + d2[y1]*w2 + d3[y1]*w3 + d4[y1]*w4 + d5[y1]*w5\ndf[z1] = d1[z1]*w1 + d2[z1]*w2 + d3[z1]*w3 + d4[z1]*w4 + d5[z1]*w5\n\ndf[x2] = d1[x2]*w1 + d2[x2]*w2 + d3[x2]*w3 + d4[x2]*w4 + d5[x2]*w5\ndf[y2] = d1[y2]*w1 + d2[y2]*w2 + d3[y2]*w3 + d4[y2]*w4 + d5[y2]*w5\ndf[z2] = d1[z2]*w1 + d2[z2]*w2 + d3[z2]*w3 + d4[z2]*w4 + d5[z2]*w5\n\ndf[x3] = d1[x3]*w1 + d2[x3]*w2 + d3[x3]*w3 + d4[x3]*w4 + d5[x3]*w5\ndf[y3] = d1[y3]*w1 + d2[y3]*w2 + d3[y3]*w3 + d4[y3]*w4 + d5[y3]*w5\ndf[z3] = d1[z3]*w1 + d2[z3]*w2 + d3[z3]*w3 + d4[z3]*w4 + d5[z3]*w5\n\ndf[x4] = d1[x4]*w1 + d2[x4]*w2 + d3[x4]*w3 + d4[x4]*w4 + d5[x4]*w5\ndf[y4] = d1[y4]*w1 + d2[y4]*w2 + d3[y4]*w3 + d4[y4]*w4 + d5[y4]*w5\ndf[z4] = d1[z4]*w1 + d2[z4]*w2 + d3[z4]*w3 + d4[z4]*w4 + d5[z4]*w5\n\ndf[x5] = d1[x5]*w1 + d2[x5]*w2 + d3[x5]*w3 + d4[x5]*w4 + d5[x5]*w5\ndf[y5] = d1[y5]*w1 + d2[y5]*w2 + d3[y5]*w3 + d4[y5]*w4 + d5[y5]*w5\ndf[z5] = d1[z5]*w1 + d2[z5]*w2 + d3[z5]*w3 + d4[z5]*w4 + d5[z5]*w5","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:54:52.644868Z","iopub.execute_input":"2026-01-09T05:54:52.645279Z","iopub.status.idle":"2026-01-09T05:54:52.692194Z","shell.execute_reply.started":"2026-01-09T05:54:52.645253Z","shell.execute_reply":"2026-01-09T05:54:52.691053Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if 'Model_1' in ensemble_of_Solutions: display(df1)\nif 'Model_2' in ensemble_of_Solutions: display(df2)\nif 'Model_3' in ensemble_of_Solutions: display(df3)\nif 'Model_4' in ensemble_of_Solutions: display(df4)\n\nif 'Model_5' in ensemble_of_Solutions: display(df5)\nif 'Model_6' in ensemble_of_Solutions: display(df6)\nif 'Model_7' in ensemble_of_Solutions: display(df7)\nif 'Model_8' in ensemble_of_Solutions: display(df8)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:54:58.396317Z","iopub.execute_input":"2026-01-09T05:54:58.396675Z","iopub.status.idle":"2026-01-09T05:54:58.48727Z","shell.execute_reply.started":"2026-01-09T05:54:58.39665Z","shell.execute_reply":"2026-01-09T05:54:58.486366Z"},"_kg_hide-input":true,"_kg_hide-output":true,"jupyter":{"outputs_hidden":true},"collapsed":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Submit","metadata":{}},{"cell_type":"code","source":"for name in [f'subm_{i}' for i in range(21)]:\n    file = f'/kaggle/working/{name}.csv'\n    if os.path.isfile(file): os.remove(file)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:56:28.703544Z","iopub.execute_input":"2026-01-09T05:56:28.703855Z","iopub.status.idle":"2026-01-09T05:56:28.710941Z","shell.execute_reply.started":"2026-01-09T05:56:28.703837Z","shell.execute_reply":"2026-01-09T05:56:28.709782Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.to_csv('submission.csv', index=False)\ndf","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T05:56:33.379654Z","iopub.execute_input":"2026-01-09T05:56:33.380038Z","iopub.status.idle":"2026-01-09T05:56:33.670859Z","shell.execute_reply.started":"2026-01-09T05:56:33.380011Z","shell.execute_reply":"2026-01-09T05:56:33.669633Z"}},"outputs":[],"execution_count":null}]}