{
  "id": 687180,
  "title": "7th Place Solution",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/687180",
  "author_name": "Samus",
  "post_date": "2026-04-02T15:00:42.632000",
  "votes": 12,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks to the organizers for a great competition!</p>\n<h2>Overview</h2>\n<p>3-phase ensemble pipeline combining Template-Based Modeling (TBM), Protenix (template-free) prediction, and RNAPro (NVIDIA's Part 1 winner).</p>\n<pre><code>test_sequences.csv\n      |\n      v\n+-----------+    pairwise alignment, top_n=50\n|  Phase 1  |--&gt; TBM Predictions (up to 5 slots)\n|   TBM     |\n+-----------+\n      |  targets with remaining slots\n      v\n+-----------+    2-5 samples (adaptive)\n|  Phase 2  |--&gt; Protenix (template-free)\n|  Protenix |\n+-----------+\n      |  combine TBM + Protenix into 5 slots\n      v\nconvert_templates_to_pt_files.py --&gt; templates.pt\n      |\n      v\n+-----------+\n|  Phase 3  |    RNAPro (NVIDIA, Part 1 winner)\n|  RNAPro   |    N_sample=1, N_step=200, max_len=1000\n+-----------+\n      |  merge (slots 1-4: RNAPro, slot 5: Phase 1-2 best)\n      v\nsubmission.csv (final)\n</code></pre>\n<h2>Motivation: Why Multi-Model Ensemble?</h2>\n<p>Public notebooks built on TBM + Protenix were scoring 0.43–0.44 on the public LB. When I forked and resubmitted the same notebooks, <strong>scores varied by ~0.02–0.03 across reruns</strong> due to Protenix's stochastic diffusion. This got me thinking — how much of these scores is real accuracy vs. lucky sampling?\nSince most participants seemed to be on TBM + Protenix, I figured I needed a different model to stand out. That led me to RNAPro.</p>\n<h2>Phase 1: Template-Based Modeling (TBM)</h2>\n<p>The core TBM (alignment parameters, constraints, slot perturbations) is mostly the same as public notebooks. The one addition is <strong>dynamic slot allocation</strong> — adjusting the TBM vs Protenix ratio based on template quality.\nFor each test target, find similar sequences from the template pool (train + validation) using pairwise alignment, ranked by normalized alignment score.</p>\n<h3>Slot Strategy</h3>\n<p>Instead of filling all 5 slots with TBM, allocate slots based on how good the best template is:\nAllocate 5 prediction slots based on the best template's quality. The table below shows the maximum TBM slots — the actual count may be lower if the template pool lacks enough sequences passing the similarity/identity thresholds. Remaining slots are filled by Protenix.</p>\n<pre><code>Template Quality Assessment\n------------------------------------------------------------\npct_identity ≥ 80%          --&gt;  HIGH    --&gt; up to 5 TBM\nAND match_rate ≥ 90%\npct_identity ≥ 50%          --&gt;  MEDIUM  --&gt; up to 3 TBM\notherwise                   --&gt;  LOW     --&gt; up to 1 TBM\n------------------------------------------------------------\n</code></pre>\n<p>In practice, most targets classified as HIGH still had fewer than 5 qualifying templates — e.g., only 1 out of 28 targets achieved a full 5 TBM slots in this run.\nEach TBM slot applies a different perturbation to diversify predictions:</p>\n<pre><code>Slot 1: Best template (direct adaptation)\nSlot 2: + Gaussian noise (scaled by dissimilarity)\nSlot 3: + Hinge motion (longest chain segment)\nSlot 4: + Per-chain jitter\nSlot 5: + Smooth wiggle deformation\n        |\n        v\nadaptive_rna_constraints()  &lt;-- stereochemical regularization\n</code></pre>\n<p>Key parameters: <code>top_n=50</code>, <code>MIN_PERCENT_IDENTITY=50%</code></p>\n<h2>Phase 2: Protenix (Template-Free)</h2>\n<p>For targets with remaining empty slots (medium/low quality templates), run <strong>Protenix</strong> (qiweiyin's adjusted weights) for template-free structure prediction, with MSA and template disabled. The number of Protenix samples is adaptive per target — 2 for medium-quality templates, 4 for low-quality (or 5 if no qualifying TBM template exists). TBM and Protenix predictions are combined into 5 slots per target:</p>\n<pre><code>Per target:\n+-----------------------------------------------------+\n| Slot 1  | Slot 2  | Slot 3  | Slot 4  | Slot 5    |\n|---------|---------|---------|---------|-----------|\n|  TBM    |  TBM    |  TBM    | Protenix| Protenix  |  &lt;- medium quality\n|  TBM    |  TBM    |  TBM    |  TBM    |  TBM      |  &lt;- high quality\n|  TBM    | Protenix| Protenix| Protenix| Protenix  |  &lt;- low quality\n+-----------------------------------------------------+\n</code></pre>\n<h2>Phase 3: RNAPro Enhancement (Key Differentiator)</h2>\n<p>Run <strong>RNAPro</strong> (NVIDIA, Part 1 winner, score 0.640) using Phase 1-2 predictions as template input. The RNAPro inference pipeline is based on <a href=\"https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference\" target=\"_blank\">theoviel's public notebook</a>, adapted to use our Phase 1-2 output as templates. For targets where RNAPro succeeds (≤1000nt), slots 1-4 are replaced with RNAPro predictions while slot 5 retains the Phase 1-2 best prediction. Targets exceeding 1000nt keep Phase 1-2 predictions entirely.</p>\n<pre><code>Phase 1-2 submission.csv\n        |\n        v\nconvert_templates_to_pt_files.py  --&gt;  templates.pt\n                                             |\n                                             v\n                              +-------------------------+\n                              |        RNAPro           |\n                              |  checkpoint: 500M       |\n                              |  + RibonanzaNet2 (MSA)  |\n                              |                         |\n                              |  N_sample=1             |\n                              |  N_step=200, N_cycle=10 |\n                              |  max_len=1000           |\n                              +------------+------------+\n                                           |\n                              +------------v------------+\n                              |   Per-target merge:     |\n                              |                         |\n                              |   RNAPro OK (≤1000nt):  |\n                              |     Slots 1-4 &lt;- RNAPro |\n                              |     Slot 5   &lt;- Phase1-2|\n                              |                  best   |\n                              |                         |\n                              |   RNAPro skip (&gt;1000nt) |\n                              |   or fail/timeout:      |\n                              |     -&gt; KEEP Phase 1-2    |\n                              +------------+------------+\n                                           |\n                                           v\n                                  submission.csv (final)\n</code></pre>\n<p>RNAPro runs sequentially on a single GPU in the submitted notebook.</p>\n<h2>What I Tried</h2>\n<h3>Protenix Parameter Tuning</h3>\n<p>Before pursuing multi-model ensembles, I tried many variations on the Protenix + TBM pipeline. None of the changes showed clear improvement:</p>\n<ul>\n<li>Protenix settings: SDPA/Flash Attention, MSA, multi-chain input, ligand SMILES, seed changes</li>\n<li>Inference scaling: multi-seed strategies, N_model_seed=2, ranking score-based sample selection</li>\n<li>Protenix version upgrade: v1.0.9 backport (SDPA, mol_type fix, chain reindex)</li>\n<li>Template pool expansion: Ribonanza templates (+167k), post-cutoff PDB, full PDB-RNA (+23k)</li>\n</ul>\n<h3>Alternative Models Explored</h3>\n<p>After running out of ideas on Protenix tuning, I tried other models:</p>\n<ul>\n<li><strong>Chai-1</strong> — AF3-family; slow and weak on validation; rejected</li>\n<li><strong>RhoFold+</strong> — MSA-based; similar results to Chai-1; rejected</li>\n<li><strong>DRfold2</strong> — RNA language model; fast but weak standalone; rejected</li>\n<li><strong>Boltz</strong> — AF3-family; considered but deprioritized</li>\n<li><strong>RNAPro</strong> — Part 1 winner (0.640); the only model that reached competitive scores on public LB</li>\n</ul>\n<h3>Limitations After RNAPro Introduction</h3>\n<p>After integrating RNAPro into the pipeline, I was unable to conduct thorough ablation studies or hyperparameter tuning. The Kaggle GPU weekly quota was nearly exhausted from the preceding experiments, and the competition deadline left little room for systematic exploration of ensemble strategies (e.g., weighted merging, per-target model selection, or increasing RNAPro's <code>N_sample</code>).</p>\n<h2>Compute</h2>\n<pre><code>+-----------------------------------------------------+\n| Hardware: Kaggle P100 x1    Time limit: 8 hours     |\n+-------------+--------------+------------------------+\n| Phase       | Time         | Notes                  |\n+-------------+--------------+------------------------+\n| Phase 1 TBM | ~11 min      | CPU-only               |\n| Phase 2 Ptx | ~79 min      | GPU (diffusion model)  |\n| Phase 3 RNAPro | ~282 min  | GPU (sequential)       |\n+-------------+--------------+------------------------+\n| Total       | ~6.2 hours   | within 8h limit        |\n+-------------+--------------+------------------------+\n</code></pre>\n<h2>Acknowledgments</h2>\n<p>Thank you to the Stanford and Kaggle organizers for hosting this competition, and to all participants who shared their code publicly.</p>\n<h2>References</h2>\n<ul>\n<li><strong><a href=\"https://github.com/NVIDIA-Digital-Bio/RNAPro\" target=\"_blank\">RNAPro</a></strong> (NVIDIA / Theo Viel) — Part 1 winning model; weights and inference code</li>\n<li><strong><a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">Protenix</a></strong> (ByteDance) — Diffusion-based structure prediction; adjusted weights by qiweiyin</li>\n<li><strong><a href=\"https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference\" target=\"_blank\">theoviel's RNAPro inference notebook</a></strong> — Reference for Phase 3 (RNAPro) pipeline</li>\n</ul>",
  "messages": [
    {
      "id": 3434081,
      "postDate": "2026-04-02T15:00:42.633Z",
      "content": "<p>Thanks to the organizers for a great competition!</p>\n<h2>Overview</h2>\n<p>3-phase ensemble pipeline combining Template-Based Modeling (TBM), Protenix (template-free) prediction, and RNAPro (NVIDIA's Part 1 winner).</p>\n<pre><code>test_sequences.csv\n      |\n      v\n+-----------+    pairwise alignment, top_n=50\n|  Phase 1  |--&gt; TBM Predictions (up to 5 slots)\n|   TBM     |\n+-----------+\n      |  targets with remaining slots\n      v\n+-----------+    2-5 samples (adaptive)\n|  Phase 2  |--&gt; Protenix (template-free)\n|  Protenix |\n+-----------+\n      |  combine TBM + Protenix into 5 slots\n      v\nconvert_templates_to_pt_files.py --&gt; templates.pt\n      |\n      v\n+-----------+\n|  Phase 3  |    RNAPro (NVIDIA, Part 1 winner)\n|  RNAPro   |    N_sample=1, N_step=200, max_len=1000\n+-----------+\n      |  merge (slots 1-4: RNAPro, slot 5: Phase 1-2 best)\n      v\nsubmission.csv (final)\n</code></pre>\n<h2>Motivation: Why Multi-Model Ensemble?</h2>\n<p>Public notebooks built on TBM + Protenix were scoring 0.43–0.44 on the public LB. When I forked and resubmitted the same notebooks, <strong>scores varied by ~0.02–0.03 across reruns</strong> due to Protenix's stochastic diffusion. This got me thinking — how much of these scores is real accuracy vs. lucky sampling?\nSince most participants seemed to be on TBM + Protenix, I figured I needed a different model to stand out. That led me to RNAPro.</p>\n<h2>Phase 1: Template-Based Modeling (TBM)</h2>\n<p>The core TBM (alignment parameters, constraints, slot perturbations) is mostly the same as public notebooks. The one addition is <strong>dynamic slot allocation</strong> — adjusting the TBM vs Protenix ratio based on template quality.\nFor each test target, find similar sequences from the template pool (train + validation) using pairwise alignment, ranked by normalized alignment score.</p>\n<h3>Slot Strategy</h3>\n<p>Instead of filling all 5 slots with TBM, allocate slots based on how good the best template is:\nAllocate 5 prediction slots based on the best template's quality. The table below shows the maximum TBM slots — the actual count may be lower if the template pool lacks enough sequences passing the similarity/identity thresholds. Remaining slots are filled by Protenix.</p>\n<pre><code>Template Quality Assessment\n------------------------------------------------------------\npct_identity ≥ 80%          --&gt;  HIGH    --&gt; up to 5 TBM\nAND match_rate ≥ 90%\npct_identity ≥ 50%          --&gt;  MEDIUM  --&gt; up to 3 TBM\notherwise                   --&gt;  LOW     --&gt; up to 1 TBM\n------------------------------------------------------------\n</code></pre>\n<p>In practice, most targets classified as HIGH still had fewer than 5 qualifying templates — e.g., only 1 out of 28 targets achieved a full 5 TBM slots in this run.\nEach TBM slot applies a different perturbation to diversify predictions:</p>\n<pre><code>Slot 1: Best template (direct adaptation)\nSlot 2: + Gaussian noise (scaled by dissimilarity)\nSlot 3: + Hinge motion (longest chain segment)\nSlot 4: + Per-chain jitter\nSlot 5: + Smooth wiggle deformation\n        |\n        v\nadaptive_rna_constraints()  &lt;-- stereochemical regularization\n</code></pre>\n<p>Key parameters: <code>top_n=50</code>, <code>MIN_PERCENT_IDENTITY=50%</code></p>\n<h2>Phase 2: Protenix (Template-Free)</h2>\n<p>For targets with remaining empty slots (medium/low quality templates), run <strong>Protenix</strong> (qiweiyin's adjusted weights) for template-free structure prediction, with MSA and template disabled. The number of Protenix samples is adaptive per target — 2 for medium-quality templates, 4 for low-quality (or 5 if no qualifying TBM template exists). TBM and Protenix predictions are combined into 5 slots per target:</p>\n<pre><code>Per target:\n+-----------------------------------------------------+\n| Slot 1  | Slot 2  | Slot 3  | Slot 4  | Slot 5    |\n|---------|---------|---------|---------|-----------|\n|  TBM    |  TBM    |  TBM    | Protenix| Protenix  |  &lt;- medium quality\n|  TBM    |  TBM    |  TBM    |  TBM    |  TBM      |  &lt;- high quality\n|  TBM    | Protenix| Protenix| Protenix| Protenix  |  &lt;- low quality\n+-----------------------------------------------------+\n</code></pre>\n<h2>Phase 3: RNAPro Enhancement (Key Differentiator)</h2>\n<p>Run <strong>RNAPro</strong> (NVIDIA, Part 1 winner, score 0.640) using Phase 1-2 predictions as template input. The RNAPro inference pipeline is based on <a href=\"https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference\" target=\"_blank\">theoviel's public notebook</a>, adapted to use our Phase 1-2 output as templates. For targets where RNAPro succeeds (≤1000nt), slots 1-4 are replaced with RNAPro predictions while slot 5 retains the Phase 1-2 best prediction. Targets exceeding 1000nt keep Phase 1-2 predictions entirely.</p>\n<pre><code>Phase 1-2 submission.csv\n        |\n        v\nconvert_templates_to_pt_files.py  --&gt;  templates.pt\n                                             |\n                                             v\n                              +-------------------------+\n                              |        RNAPro           |\n                              |  checkpoint: 500M       |\n                              |  + RibonanzaNet2 (MSA)  |\n                              |                         |\n                              |  N_sample=1             |\n                              |  N_step=200, N_cycle=10 |\n                              |  max_len=1000           |\n                              +------------+------------+\n                                           |\n                              +------------v------------+\n                              |   Per-target merge:     |\n                              |                         |\n                              |   RNAPro OK (≤1000nt):  |\n                              |     Slots 1-4 &lt;- RNAPro |\n                              |     Slot 5   &lt;- Phase1-2|\n                              |                  best   |\n                              |                         |\n                              |   RNAPro skip (&gt;1000nt) |\n                              |   or fail/timeout:      |\n                              |     -&gt; KEEP Phase 1-2    |\n                              +------------+------------+\n                                           |\n                                           v\n                                  submission.csv (final)\n</code></pre>\n<p>RNAPro runs sequentially on a single GPU in the submitted notebook.</p>\n<h2>What I Tried</h2>\n<h3>Protenix Parameter Tuning</h3>\n<p>Before pursuing multi-model ensembles, I tried many variations on the Protenix + TBM pipeline. None of the changes showed clear improvement:</p>\n<ul>\n<li>Protenix settings: SDPA/Flash Attention, MSA, multi-chain input, ligand SMILES, seed changes</li>\n<li>Inference scaling: multi-seed strategies, N_model_seed=2, ranking score-based sample selection</li>\n<li>Protenix version upgrade: v1.0.9 backport (SDPA, mol_type fix, chain reindex)</li>\n<li>Template pool expansion: Ribonanza templates (+167k), post-cutoff PDB, full PDB-RNA (+23k)</li>\n</ul>\n<h3>Alternative Models Explored</h3>\n<p>After running out of ideas on Protenix tuning, I tried other models:</p>\n<ul>\n<li><strong>Chai-1</strong> — AF3-family; slow and weak on validation; rejected</li>\n<li><strong>RhoFold+</strong> — MSA-based; similar results to Chai-1; rejected</li>\n<li><strong>DRfold2</strong> — RNA language model; fast but weak standalone; rejected</li>\n<li><strong>Boltz</strong> — AF3-family; considered but deprioritized</li>\n<li><strong>RNAPro</strong> — Part 1 winner (0.640); the only model that reached competitive scores on public LB</li>\n</ul>\n<h3>Limitations After RNAPro Introduction</h3>\n<p>After integrating RNAPro into the pipeline, I was unable to conduct thorough ablation studies or hyperparameter tuning. The Kaggle GPU weekly quota was nearly exhausted from the preceding experiments, and the competition deadline left little room for systematic exploration of ensemble strategies (e.g., weighted merging, per-target model selection, or increasing RNAPro's <code>N_sample</code>).</p>\n<h2>Compute</h2>\n<pre><code>+-----------------------------------------------------+\n| Hardware: Kaggle P100 x1    Time limit: 8 hours     |\n+-------------+--------------+------------------------+\n| Phase       | Time         | Notes                  |\n+-------------+--------------+------------------------+\n| Phase 1 TBM | ~11 min      | CPU-only               |\n| Phase 2 Ptx | ~79 min      | GPU (diffusion model)  |\n| Phase 3 RNAPro | ~282 min  | GPU (sequential)       |\n+-------------+--------------+------------------------+\n| Total       | ~6.2 hours   | within 8h limit        |\n+-------------+--------------+------------------------+\n</code></pre>\n<h2>Acknowledgments</h2>\n<p>Thank you to the Stanford and Kaggle organizers for hosting this competition, and to all participants who shared their code publicly.</p>\n<h2>References</h2>\n<ul>\n<li><strong><a href=\"https://github.com/NVIDIA-Digital-Bio/RNAPro\" target=\"_blank\">RNAPro</a></strong> (NVIDIA / Theo Viel) — Part 1 winning model; weights and inference code</li>\n<li><strong><a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">Protenix</a></strong> (ByteDance) — Diffusion-based structure prediction; adjusted weights by qiweiyin</li>\n<li><strong><a href=\"https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference\" target=\"_blank\">theoviel's RNAPro inference notebook</a></strong> — Reference for Phase 3 (RNAPro) pipeline</li>\n</ul>",
      "rawMarkdown": "\nThanks to the organizers for a great competition!\n\n## Overview\n\n3-phase ensemble pipeline combining Template-Based Modeling (TBM), Protenix (template-free) prediction, and RNAPro (NVIDIA's Part 1 winner).\n\n```\n test_sequences.csv\n       |\n       v\n +-----------+    pairwise alignment, top_n=50\n |  Phase 1  |--> TBM Predictions (up to 5 slots)\n |   TBM     |\n +-----------+\n       |  targets with remaining slots\n       v\n +-----------+    2-5 samples (adaptive)\n |  Phase 2  |--> Protenix (template-free)\n |  Protenix |\n +-----------+\n       |  combine TBM + Protenix into 5 slots\n       v\n convert_templates_to_pt_files.py --> templates.pt\n       |\n       v\n +-----------+\n |  Phase 3  |    RNAPro (NVIDIA, Part 1 winner)\n |  RNAPro   |    N_sample=1, N_step=200, max_len=1000\n +-----------+\n       |  merge (slots 1-4: RNAPro, slot 5: Phase 1-2 best)\n       v\n submission.csv (final)\n```\n\n\n\n## Motivation: Why Multi-Model Ensemble?\n\nPublic notebooks built on TBM + Protenix were scoring 0.43–0.44 on the public LB. When I forked and resubmitted the same notebooks, **scores varied by ~0.02–0.03 across reruns** due to Protenix's stochastic diffusion. This got me thinking — how much of these scores is real accuracy vs. lucky sampling?\n\nSince most participants seemed to be on TBM + Protenix, I figured I needed a different model to stand out. That led me to RNAPro.\n\n\n\n## Phase 1: Template-Based Modeling (TBM)\n\nThe core TBM (alignment parameters, constraints, slot perturbations) is mostly the same as public notebooks. The one addition is **dynamic slot allocation** — adjusting the TBM vs Protenix ratio based on template quality.\n\nFor each test target, find similar sequences from the template pool (train + validation) using pairwise alignment, ranked by normalized alignment score.\n\n### Slot Strategy\n\nInstead of filling all 5 slots with TBM, allocate slots based on how good the best template is:\n\nAllocate 5 prediction slots based on the best template's quality. The table below shows the maximum TBM slots — the actual count may be lower if the template pool lacks enough sequences passing the similarity/identity thresholds. Remaining slots are filled by Protenix.\n\n```\n Template Quality Assessment\n ------------------------------------------------------------\n pct_identity ≥ 80%          -->  HIGH    --> up to 5 TBM\n AND match_rate ≥ 90%\n\n pct_identity ≥ 50%          -->  MEDIUM  --> up to 3 TBM\n\n otherwise                   -->  LOW     --> up to 1 TBM\n ------------------------------------------------------------\n```\n\nIn practice, most targets classified as HIGH still had fewer than 5 qualifying templates — e.g., only 1 out of 28 targets achieved a full 5 TBM slots in this run.\n\nEach TBM slot applies a different perturbation to diversify predictions:\n\n```\n Slot 1: Best template (direct adaptation)\n Slot 2: + Gaussian noise (scaled by dissimilarity)\n Slot 3: + Hinge motion (longest chain segment)\n Slot 4: + Per-chain jitter\n Slot 5: + Smooth wiggle deformation\n         |\n         v\n adaptive_rna_constraints()  <-- stereochemical regularization\n```\n\nKey parameters: `top_n=50`, `MIN_PERCENT_IDENTITY=50%`\n\n\n## Phase 2: Protenix (Template-Free)\n\nFor targets with remaining empty slots (medium/low quality templates), run **Protenix** (qiweiyin's adjusted weights) for template-free structure prediction, with MSA and template disabled. The number of Protenix samples is adaptive per target — 2 for medium-quality templates, 4 for low-quality (or 5 if no qualifying TBM template exists). TBM and Protenix predictions are combined into 5 slots per target:\n\n```\n Per target:\n +-----------------------------------------------------+\n | Slot 1  | Slot 2  | Slot 3  | Slot 4  | Slot 5    |\n |---------|---------|---------|---------|-----------|\n |  TBM    |  TBM    |  TBM    | Protenix| Protenix  |  <- medium quality\n |  TBM    |  TBM    |  TBM    |  TBM    |  TBM      |  <- high quality\n |  TBM    | Protenix| Protenix| Protenix| Protenix  |  <- low quality\n +-----------------------------------------------------+\n```\n\n\n## Phase 3: RNAPro Enhancement (Key Differentiator)\n\nRun **RNAPro** (NVIDIA, Part 1 winner, score 0.640) using Phase 1-2 predictions as template input. The RNAPro inference pipeline is based on [theoviel's public notebook](https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference), adapted to use our Phase 1-2 output as templates. For targets where RNAPro succeeds (≤1000nt), slots 1-4 are replaced with RNAPro predictions while slot 5 retains the Phase 1-2 best prediction. Targets exceeding 1000nt keep Phase 1-2 predictions entirely.\n\n```\n Phase 1-2 submission.csv\n         |\n         v\n convert_templates_to_pt_files.py  -->  templates.pt\n                                              |\n                                              v\n                               +-------------------------+\n                               |        RNAPro           |\n                               |  checkpoint: 500M       |\n                               |  + RibonanzaNet2 (MSA)  |\n                               |                         |\n                               |  N_sample=1             |\n                               |  N_step=200, N_cycle=10 |\n                               |  max_len=1000           |\n                               +------------+------------+\n                                            |\n                               +------------v------------+\n                               |   Per-target merge:     |\n                               |                         |\n                               |   RNAPro OK (≤1000nt):  |\n                               |     Slots 1-4 <- RNAPro |\n                               |     Slot 5   <- Phase1-2|\n                               |                  best   |\n                               |                         |\n                               |   RNAPro skip (>1000nt) |\n                               |   or fail/timeout:      |\n                               |     -> KEEP Phase 1-2    |\n                               +------------+------------+\n                                            |\n                                            v\n                                   submission.csv (final)\n```\n\nRNAPro runs sequentially on a single GPU in the submitted notebook.\n\n\n## What I Tried\n\n### Protenix Parameter Tuning\n\nBefore pursuing multi-model ensembles, I tried many variations on the Protenix + TBM pipeline. None of the changes showed clear improvement:\n\n- Protenix settings: SDPA/Flash Attention, MSA, multi-chain input, ligand SMILES, seed changes\n- Inference scaling: multi-seed strategies, N_model_seed=2, ranking score-based sample selection\n- Protenix version upgrade: v1.0.9 backport (SDPA, mol_type fix, chain reindex)\n- Template pool expansion: Ribonanza templates (+167k), post-cutoff PDB, full PDB-RNA (+23k)\n\n### Alternative Models Explored\n\nAfter running out of ideas on Protenix tuning, I tried other models:\n\n- **Chai-1** — AF3-family; slow and weak on validation; rejected\n- **RhoFold+** — MSA-based; similar results to Chai-1; rejected\n- **DRfold2** — RNA language model; fast but weak standalone; rejected\n- **Boltz** — AF3-family; considered but deprioritized\n- **RNAPro** — Part 1 winner (0.640); the only model that reached competitive scores on public LB\n\n### Limitations After RNAPro Introduction\n\nAfter integrating RNAPro into the pipeline, I was unable to conduct thorough ablation studies or hyperparameter tuning. The Kaggle GPU weekly quota was nearly exhausted from the preceding experiments, and the competition deadline left little room for systematic exploration of ensemble strategies (e.g., weighted merging, per-target model selection, or increasing RNAPro's `N_sample`).\n\n\n## Compute\n\n```\n +-----------------------------------------------------+\n | Hardware: Kaggle P100 x1    Time limit: 8 hours     |\n +-------------+--------------+------------------------+\n | Phase       | Time         | Notes                  |\n +-------------+--------------+------------------------+\n | Phase 1 TBM | ~11 min      | CPU-only               |\n | Phase 2 Ptx | ~79 min      | GPU (diffusion model)  |\n | Phase 3 RNAPro | ~282 min  | GPU (sequential)       |\n +-------------+--------------+------------------------+\n | Total       | ~6.2 hours   | within 8h limit        |\n +-------------+--------------+------------------------+\n```\n\n\n## Acknowledgments\n\nThank you to the Stanford and Kaggle organizers for hosting this competition, and to all participants who shared their code publicly.\n\n\n\n## References\n\n- **[RNAPro](https://github.com/NVIDIA-Digital-Bio/RNAPro)** (NVIDIA / Theo Viel) — Part 1 winning model; weights and inference code\n- **[Protenix](https://github.com/bytedance/Protenix)** (ByteDance) — Diffusion-based structure prediction; adjusted weights by qiweiyin\n- **[theoviel's RNAPro inference notebook](https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference)** — Reference for Phase 3 (RNAPro) pipeline",
      "votes": 12
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3434081": "\nThanks to the organizers for a great competition!\n\n## Overview\n\n3-phase ensemble pipeline combining Template-Based Modeling (TBM), Protenix (template-free) prediction, and RNAPro (NVIDIA's Part 1 winner).\n\n```\n test_sequences.csv\n       |\n       v\n +-----------+    pairwise alignment, top_n=50\n |  Phase 1  |--> TBM Predictions (up to 5 slots)\n |   TBM     |\n +-----------+\n       |  targets with remaining slots\n       v\n +-----------+    2-5 samples (adaptive)\n |  Phase 2  |--> Protenix (template-free)\n |  Protenix |\n +-----------+\n       |  combine TBM + Protenix into 5 slots\n       v\n convert_templates_to_pt_files.py --> templates.pt\n       |\n       v\n +-----------+\n |  Phase 3  |    RNAPro (NVIDIA, Part 1 winner)\n |  RNAPro   |    N_sample=1, N_step=200, max_len=1000\n +-----------+\n       |  merge (slots 1-4: RNAPro, slot 5: Phase 1-2 best)\n       v\n submission.csv (final)\n```\n\n\n\n## Motivation: Why Multi-Model Ensemble?\n\nPublic notebooks built on TBM + Protenix were scoring 0.43–0.44 on the public LB. When I forked and resubmitted the same notebooks, **scores varied by ~0.02–0.03 across reruns** due to Protenix's stochastic diffusion. This got me thinking — how much of these scores is real accuracy vs. lucky sampling?\n\nSince most participants seemed to be on TBM + Protenix, I figured I needed a different model to stand out. That led me to RNAPro.\n\n\n\n## Phase 1: Template-Based Modeling (TBM)\n\nThe core TBM (alignment parameters, constraints, slot perturbations) is mostly the same as public notebooks. The one addition is **dynamic slot allocation** — adjusting the TBM vs Protenix ratio based on template quality.\n\nFor each test target, find similar sequences from the template pool (train + validation) using pairwise alignment, ranked by normalized alignment score.\n\n### Slot Strategy\n\nInstead of filling all 5 slots with TBM, allocate slots based on how good the best template is:\n\nAllocate 5 prediction slots based on the best template's quality. The table below shows the maximum TBM slots — the actual count may be lower if the template pool lacks enough sequences passing the similarity/identity thresholds. Remaining slots are filled by Protenix.\n\n```\n Template Quality Assessment\n ------------------------------------------------------------\n pct_identity ≥ 80%          -->  HIGH    --> up to 5 TBM\n AND match_rate ≥ 90%\n\n pct_identity ≥ 50%          -->  MEDIUM  --> up to 3 TBM\n\n otherwise                   -->  LOW     --> up to 1 TBM\n ------------------------------------------------------------\n```\n\nIn practice, most targets classified as HIGH still had fewer than 5 qualifying templates — e.g., only 1 out of 28 targets achieved a full 5 TBM slots in this run.\n\nEach TBM slot applies a different perturbation to diversify predictions:\n\n```\n Slot 1: Best template (direct adaptation)\n Slot 2: + Gaussian noise (scaled by dissimilarity)\n Slot 3: + Hinge motion (longest chain segment)\n Slot 4: + Per-chain jitter\n Slot 5: + Smooth wiggle deformation\n         |\n         v\n adaptive_rna_constraints()  <-- stereochemical regularization\n```\n\nKey parameters: `top_n=50`, `MIN_PERCENT_IDENTITY=50%`\n\n\n## Phase 2: Protenix (Template-Free)\n\nFor targets with remaining empty slots (medium/low quality templates), run **Protenix** (qiweiyin's adjusted weights) for template-free structure prediction, with MSA and template disabled. The number of Protenix samples is adaptive per target — 2 for medium-quality templates, 4 for low-quality (or 5 if no qualifying TBM template exists). TBM and Protenix predictions are combined into 5 slots per target:\n\n```\n Per target:\n +-----------------------------------------------------+\n | Slot 1  | Slot 2  | Slot 3  | Slot 4  | Slot 5    |\n |---------|---------|---------|---------|-----------|\n |  TBM    |  TBM    |  TBM    | Protenix| Protenix  |  <- medium quality\n |  TBM    |  TBM    |  TBM    |  TBM    |  TBM      |  <- high quality\n |  TBM    | Protenix| Protenix| Protenix| Protenix  |  <- low quality\n +-----------------------------------------------------+\n```\n\n\n## Phase 3: RNAPro Enhancement (Key Differentiator)\n\nRun **RNAPro** (NVIDIA, Part 1 winner, score 0.640) using Phase 1-2 predictions as template input. The RNAPro inference pipeline is based on [theoviel's public notebook](https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference), adapted to use our Phase 1-2 output as templates. For targets where RNAPro succeeds (≤1000nt), slots 1-4 are replaced with RNAPro predictions while slot 5 retains the Phase 1-2 best prediction. Targets exceeding 1000nt keep Phase 1-2 predictions entirely.\n\n```\n Phase 1-2 submission.csv\n         |\n         v\n convert_templates_to_pt_files.py  -->  templates.pt\n                                              |\n                                              v\n                               +-------------------------+\n                               |        RNAPro           |\n                               |  checkpoint: 500M       |\n                               |  + RibonanzaNet2 (MSA)  |\n                               |                         |\n                               |  N_sample=1             |\n                               |  N_step=200, N_cycle=10 |\n                               |  max_len=1000           |\n                               +------------+------------+\n                                            |\n                               +------------v------------+\n                               |   Per-target merge:     |\n                               |                         |\n                               |   RNAPro OK (≤1000nt):  |\n                               |     Slots 1-4 <- RNAPro |\n                               |     Slot 5   <- Phase1-2|\n                               |                  best   |\n                               |                         |\n                               |   RNAPro skip (>1000nt) |\n                               |   or fail/timeout:      |\n                               |     -> KEEP Phase 1-2    |\n                               +------------+------------+\n                                            |\n                                            v\n                                   submission.csv (final)\n```\n\nRNAPro runs sequentially on a single GPU in the submitted notebook.\n\n\n## What I Tried\n\n### Protenix Parameter Tuning\n\nBefore pursuing multi-model ensembles, I tried many variations on the Protenix + TBM pipeline. None of the changes showed clear improvement:\n\n- Protenix settings: SDPA/Flash Attention, MSA, multi-chain input, ligand SMILES, seed changes\n- Inference scaling: multi-seed strategies, N_model_seed=2, ranking score-based sample selection\n- Protenix version upgrade: v1.0.9 backport (SDPA, mol_type fix, chain reindex)\n- Template pool expansion: Ribonanza templates (+167k), post-cutoff PDB, full PDB-RNA (+23k)\n\n### Alternative Models Explored\n\nAfter running out of ideas on Protenix tuning, I tried other models:\n\n- **Chai-1** — AF3-family; slow and weak on validation; rejected\n- **RhoFold+** — MSA-based; similar results to Chai-1; rejected\n- **DRfold2** — RNA language model; fast but weak standalone; rejected\n- **Boltz** — AF3-family; considered but deprioritized\n- **RNAPro** — Part 1 winner (0.640); the only model that reached competitive scores on public LB\n\n### Limitations After RNAPro Introduction\n\nAfter integrating RNAPro into the pipeline, I was unable to conduct thorough ablation studies or hyperparameter tuning. The Kaggle GPU weekly quota was nearly exhausted from the preceding experiments, and the competition deadline left little room for systematic exploration of ensemble strategies (e.g., weighted merging, per-target model selection, or increasing RNAPro's `N_sample`).\n\n\n## Compute\n\n```\n +-----------------------------------------------------+\n | Hardware: Kaggle P100 x1    Time limit: 8 hours     |\n +-------------+--------------+------------------------+\n | Phase       | Time         | Notes                  |\n +-------------+--------------+------------------------+\n | Phase 1 TBM | ~11 min      | CPU-only               |\n | Phase 2 Ptx | ~79 min      | GPU (diffusion model)  |\n | Phase 3 RNAPro | ~282 min  | GPU (sequential)       |\n +-------------+--------------+------------------------+\n | Total       | ~6.2 hours   | within 8h limit        |\n +-------------+--------------+------------------------+\n```\n\n\n## Acknowledgments\n\nThank you to the Stanford and Kaggle organizers for hosting this competition, and to all participants who shared their code publicly.\n\n\n\n## References\n\n- **[RNAPro](https://github.com/NVIDIA-Digital-Bio/RNAPro)** (NVIDIA / Theo Viel) — Part 1 winning model; weights and inference code\n- **[Protenix](https://github.com/bytedance/Protenix)** (ByteDance) — Diffusion-based structure prediction; adjusted weights by qiweiyin\n- **[theoviel's RNAPro inference notebook](https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference)** — Reference for Phase 3 (RNAPro) pipeline"
  }
}