{
  "id": 687234,
  "title": "16th Place Solution ",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/687234",
  "author_name": "Shivam_Bhujbal",
  "post_date": "2026-04-02T17:44:30.518000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thank you to the organizers at Stanford, HHMI, NVIDIA, and Kaggle for this incredible competition. As someone with no formal biology or bioinformatics background, this was my first encounter with RNA structure prediction — and it turned into one of the most rewarding learning experiences I've had. Congratulations to all participants </p>\n<h2>The Core Idea</h2>\n<p>Use Protenix as a <strong>template generator</strong>, not as the final predictor. Protenix produces one structural prediction per target, which gets fed into RNAPro as template conditioning. RNAPro then refines it using RibonanzaNet2 embeddings and its Pairformer architecture. TBM fills the remaining template slots to give RNAPro structural diversity to work with.</p>\n<p>The time constraint shaped the design: running Protenix with N_SAMPLE=5 was too slow, so I used N_SAMPLE=1 and let RNAPro do the heavy lifting for diversity.</p>\n<h2>Pipeline Overview</h2>\n<pre><code>                    ┌──────────────┐\n                    │   Test Seqs  │\n                    └──────┬───────┘\n                           │\n              ┌────────────┴────────────┐\n              ▼                         ▼\n      ┌──────────────┐         ┌──────────────┐\n      │     TBM      │         │   Protenix   │\n      │  (top 5 by   │         │  (N_SAMPLE=1 │\n      │  alignment)  │         │  per target) │\n      └──────┬───────┘         └──────┬───────┘\n             │                        │\n        TBM[1..4]              Protenix[0]\n             │                        │\n             └───────────┬────────────┘\n                         ▼\n              ┌─────────────────────┐\n              │   Template CSV      │\n              │  Slot 1: Protenix   │\n              │  Slot 2-5: TBM     │\n              └─────────┬───────────┘\n                        ▼\n              ┌─────────────────────┐\n              │      RNAPro         │\n              │  (template_idx=0,   │\n              │   RibonanzaNet2,    │\n              │   N_SAMPLE=3)       │\n              └─────────┬───────────┘\n                        ▼\n              ┌─────────────────────┐\n              │   5 Predictions     │\n              ├─────────────────────┤\n              │ P1-P3: RNAPro       │\n              │ P4-P5: Protenix/TBM │\n              │        fallback     │\n              └─────────────────────┘\n</code></pre>\n<h2>Phase 1 — Template-Based Modeling</h2>\n<p>TBM serves two purposes: backup predictions and template slots 2-5 for RNAPro input.</p>\n<p><strong>Template search:</strong> BioPython <code>PairwiseAligner</code> in global mode with strong gap penalties (<code>open: -8, extend: -0.4</code>). Searched ~5744 training+validation structures. Length ratio filter skips candidates with &gt;30% length difference.</p>\n<p><strong>Coordinate transfer:</strong> C1' coordinates mapped from template to query via alignment. Gaps filled by linear interpolation between nearest aligned positions, or extrapolated at ends with 3.8Å steps.</p>\n<p><strong>No quality threshold filtering</strong> — all templates passed through to the CSV. RNAPro handles template quality internally through its confidence mechanism.</p>\n<h2>Phase 2 — Protenix (Template Generation)</h2>\n<p>Protenix ran in minimal mode — the goal was one decent structural prediction per target, not five polished ones.</p>\n<ul>\n<li><strong>N_SAMPLE=1</strong> (single prediction per target)</li>\n<li><strong>Single seed</strong> (42)</li>\n<li><strong>No MSA, RNA MSA enabled</strong></li>\n<li><strong>Chunking</strong> for sequences &gt;512nt: overlapping windows with Kabsch alignment and cosine blending at boundaries (MAX_SEQ_LEN=512, CHUNK_OVERLAP=128)</li>\n<li>C1' extraction via <code>centre_atom_mask</code> with <code>atom_to_tokatom_idx</code> fallback</li>\n</ul>\n<p>A Protenix-only backup <code>submission.csv</code> was saved before RNAPro ran — if RNAPro timed out, the notebook still produced a valid submission.</p>\n<h2>Phase 3 — Template CSV Construction</h2>\n<p>For each target, the 5-slot template CSV was built as:</p>\n<table>\n<thead>\n<tr>\n<th>Slot</th>\n<th>Source</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>Protenix prediction</td>\n<td>Primary template — RNAPro refines this</td>\n</tr>\n<tr>\n<td>2-5</td>\n<td>TBM templates</td>\n<td>Structural diversity (or noise-augmented copies if &lt;4 templates)</td>\n</tr>\n</tbody>\n</table>\n<p>If Protenix failed for a target, slot 1 fell back to the best TBM template.</p>\n<p>The CSV was converted to <code>.pt</code> format via RNAPro's <code>convert_templates_to_pt_files.py</code>.</p>\n<h2>Phase 4 — RNAPro Inference</h2>\n<p>RNAPro conditioned on the Protenix template via <code>--template_idx 0</code> (uses only the top template slot — our Protenix prediction). Its architecture combines:</p>\n<ul>\n<li><strong>RibonanzaNet2</strong> as frozen encoder — captures RNA-specific co-evolutionary features</li>\n<li><strong>Template conditioning</strong> — Protenix structural prior guides the diffusion</li>\n<li><strong>Pairformer</strong> — RNA-adapted structure module</li>\n</ul>\n<p>Params were aggressive to fit the 8-hour time budget:</p>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>N_SAMPLE</td>\n<td>3</td>\n</tr>\n<tr>\n<td>N_STEP</td>\n<td>100</td>\n</tr>\n<tr>\n<td>N_CYCLE</td>\n<td>4</td>\n</tr>\n<tr>\n<td>MAX_LEN</td>\n<td>500</td>\n</tr>\n<tr>\n<td>MSA</td>\n<td>disabled</td>\n</tr>\n<tr>\n<td>Seed</td>\n<td>42</td>\n</tr>\n<tr>\n<td>template_idx</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<h2>Phase 5 — Submission Assembly</h2>\n<p>Simple priority filling:</p>\n<ol>\n<li>RNAPro CIF outputs → slots 1-3 (if available)</li>\n<li>Protenix/TBM from backup CSV → remaining slots</li>\n<li>Helix fallback → last resort (never triggered on public test)</li>\n</ol>\n<p>No ensemble selection or post-processing — just straight slot filling.</p>\n<h2>What Made the Difference</h2>\n<p><strong>Protenix as template generator, not final predictor.</strong> The Protenix→RNAPro pipeline scored 0.464 private vs 0.410 for standalone Protenix. RNAPro's RibonanzaNet2 embeddings added RNA-specific knowledge that base Protenix lacks, especially on novel targets without close homologs.</p>\n<p><strong>Time-budget-driven design.</strong> Using Protenix N_SAMPLE=1 instead of 5 saved time, which was exactly the time needed for RNAPro to run. The constraint that felt limiting actually forced a better architecture.</p>\n<h2>What Didn't Work</h2>\n<ul>\n<li><strong>MSA in RNAPro</strong> — disabled to save time; didn't noticeably help when tested</li>\n<li><strong>Protenix N_SAMPLE=5 as templates</strong> — too slow, didn't finish within 8h</li>\n<li><strong>Running Protenix with MSA</strong> — no improvement, extra time cost</li>\n<li><strong>Post-processing constraints on neural net output</strong> — slightly degraded predictions; Protenix/RNAPro geometry is already valid</li>\n<li><strong>RBSSA (embedding-based template search)</strong> — RibonanzaNet2 embeddings + Smith-Waterman. Hurt when it replaced Protenix (0.33), marginal when additive(0.425) on public data</li>\n<li><strong>Finetuned Protenix (RNA3DB checkpoint)</strong> — scored 0.392; base model was better</li>\n<li><strong>DRfold2 / Boltz ensemble</strong> — architecturally different but not better on any target; selector correctly ignored them</li>\n</ul>\n<h2>Key Lessons</h2>\n<ol>\n<li><strong>Template quality is the foundation.</strong> Every DL model performed better with good structural starting points. Protenix-generated templates were better than pure TBM for RNAPro conditioning.</li>\n<li><strong>DL models are template refiners on Kaggle hardware.</strong> On P100 (16GB), no model reliably folded RNA from scratch. They performed best refining structural starting points.</li>\n<li><strong>Time budget shapes architecture.</strong> The 8h limit forced Protenix N_SAMPLE=1, which turned out to be the right call — one good template + RNAPro refinement beat five mediocre Protenix predictions.</li>\n</ol>\n<hr>\n<h2>Notebook 2: Protenix + TBM (Private: 0.410)</h2>\n<p>Brief summary of the second selected notebook:</p>\n<ul>\n<li>TBM with 50% identity threshold, diversity transforms (hinge, jitter, wiggle)</li>\n<li>Protenix multi-seed [42, 137] for targets below threshold</li>\n<li><strong>Consistency-based ensemble selection</strong>: anchor on consensus Protenix prediction (lowest avg RMSD to others), guarantee one TBM slot, fill remaining by RMSD consistency. This was the single biggest improvement in the Protenix pipeline (0.428→0.440 on public).</li>\n</ul>\n<h2>Acknowledgments</h2>\n<p>Thanks to the community for shared resources — Protenix notebooks, RNAPro from NVIDIA, and TBM approaches from Part 1 winners. Excited about the CASP17 collaboration opportunity.</p>\n<p><strong>Kaggle:</strong> <a href=\"https://www.kaggle.com/error1249x\" target=\"_blank\">@error1249x</a></p>",
  "messages": [
    {
      "id": 3434197,
      "postDate": "2026-04-02T17:44:30.520Z",
      "content": "<p>Thank you to the organizers at Stanford, HHMI, NVIDIA, and Kaggle for this incredible competition. As someone with no formal biology or bioinformatics background, this was my first encounter with RNA structure prediction — and it turned into one of the most rewarding learning experiences I've had. Congratulations to all participants </p>\n<h2>The Core Idea</h2>\n<p>Use Protenix as a <strong>template generator</strong>, not as the final predictor. Protenix produces one structural prediction per target, which gets fed into RNAPro as template conditioning. RNAPro then refines it using RibonanzaNet2 embeddings and its Pairformer architecture. TBM fills the remaining template slots to give RNAPro structural diversity to work with.</p>\n<p>The time constraint shaped the design: running Protenix with N_SAMPLE=5 was too slow, so I used N_SAMPLE=1 and let RNAPro do the heavy lifting for diversity.</p>\n<h2>Pipeline Overview</h2>\n<pre><code>                    ┌──────────────┐\n                    │   Test Seqs  │\n                    └──────┬───────┘\n                           │\n              ┌────────────┴────────────┐\n              ▼                         ▼\n      ┌──────────────┐         ┌──────────────┐\n      │     TBM      │         │   Protenix   │\n      │  (top 5 by   │         │  (N_SAMPLE=1 │\n      │  alignment)  │         │  per target) │\n      └──────┬───────┘         └──────┬───────┘\n             │                        │\n        TBM[1..4]              Protenix[0]\n             │                        │\n             └───────────┬────────────┘\n                         ▼\n              ┌─────────────────────┐\n              │   Template CSV      │\n              │  Slot 1: Protenix   │\n              │  Slot 2-5: TBM     │\n              └─────────┬───────────┘\n                        ▼\n              ┌─────────────────────┐\n              │      RNAPro         │\n              │  (template_idx=0,   │\n              │   RibonanzaNet2,    │\n              │   N_SAMPLE=3)       │\n              └─────────┬───────────┘\n                        ▼\n              ┌─────────────────────┐\n              │   5 Predictions     │\n              ├─────────────────────┤\n              │ P1-P3: RNAPro       │\n              │ P4-P5: Protenix/TBM │\n              │        fallback     │\n              └─────────────────────┘\n</code></pre>\n<h2>Phase 1 — Template-Based Modeling</h2>\n<p>TBM serves two purposes: backup predictions and template slots 2-5 for RNAPro input.</p>\n<p><strong>Template search:</strong> BioPython <code>PairwiseAligner</code> in global mode with strong gap penalties (<code>open: -8, extend: -0.4</code>). Searched ~5744 training+validation structures. Length ratio filter skips candidates with &gt;30% length difference.</p>\n<p><strong>Coordinate transfer:</strong> C1' coordinates mapped from template to query via alignment. Gaps filled by linear interpolation between nearest aligned positions, or extrapolated at ends with 3.8Å steps.</p>\n<p><strong>No quality threshold filtering</strong> — all templates passed through to the CSV. RNAPro handles template quality internally through its confidence mechanism.</p>\n<h2>Phase 2 — Protenix (Template Generation)</h2>\n<p>Protenix ran in minimal mode — the goal was one decent structural prediction per target, not five polished ones.</p>\n<ul>\n<li><strong>N_SAMPLE=1</strong> (single prediction per target)</li>\n<li><strong>Single seed</strong> (42)</li>\n<li><strong>No MSA, RNA MSA enabled</strong></li>\n<li><strong>Chunking</strong> for sequences &gt;512nt: overlapping windows with Kabsch alignment and cosine blending at boundaries (MAX_SEQ_LEN=512, CHUNK_OVERLAP=128)</li>\n<li>C1' extraction via <code>centre_atom_mask</code> with <code>atom_to_tokatom_idx</code> fallback</li>\n</ul>\n<p>A Protenix-only backup <code>submission.csv</code> was saved before RNAPro ran — if RNAPro timed out, the notebook still produced a valid submission.</p>\n<h2>Phase 3 — Template CSV Construction</h2>\n<p>For each target, the 5-slot template CSV was built as:</p>\n<table>\n<thead>\n<tr>\n<th>Slot</th>\n<th>Source</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>Protenix prediction</td>\n<td>Primary template — RNAPro refines this</td>\n</tr>\n<tr>\n<td>2-5</td>\n<td>TBM templates</td>\n<td>Structural diversity (or noise-augmented copies if &lt;4 templates)</td>\n</tr>\n</tbody>\n</table>\n<p>If Protenix failed for a target, slot 1 fell back to the best TBM template.</p>\n<p>The CSV was converted to <code>.pt</code> format via RNAPro's <code>convert_templates_to_pt_files.py</code>.</p>\n<h2>Phase 4 — RNAPro Inference</h2>\n<p>RNAPro conditioned on the Protenix template via <code>--template_idx 0</code> (uses only the top template slot — our Protenix prediction). Its architecture combines:</p>\n<ul>\n<li><strong>RibonanzaNet2</strong> as frozen encoder — captures RNA-specific co-evolutionary features</li>\n<li><strong>Template conditioning</strong> — Protenix structural prior guides the diffusion</li>\n<li><strong>Pairformer</strong> — RNA-adapted structure module</li>\n</ul>\n<p>Params were aggressive to fit the 8-hour time budget:</p>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>N_SAMPLE</td>\n<td>3</td>\n</tr>\n<tr>\n<td>N_STEP</td>\n<td>100</td>\n</tr>\n<tr>\n<td>N_CYCLE</td>\n<td>4</td>\n</tr>\n<tr>\n<td>MAX_LEN</td>\n<td>500</td>\n</tr>\n<tr>\n<td>MSA</td>\n<td>disabled</td>\n</tr>\n<tr>\n<td>Seed</td>\n<td>42</td>\n</tr>\n<tr>\n<td>template_idx</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<h2>Phase 5 — Submission Assembly</h2>\n<p>Simple priority filling:</p>\n<ol>\n<li>RNAPro CIF outputs → slots 1-3 (if available)</li>\n<li>Protenix/TBM from backup CSV → remaining slots</li>\n<li>Helix fallback → last resort (never triggered on public test)</li>\n</ol>\n<p>No ensemble selection or post-processing — just straight slot filling.</p>\n<h2>What Made the Difference</h2>\n<p><strong>Protenix as template generator, not final predictor.</strong> The Protenix→RNAPro pipeline scored 0.464 private vs 0.410 for standalone Protenix. RNAPro's RibonanzaNet2 embeddings added RNA-specific knowledge that base Protenix lacks, especially on novel targets without close homologs.</p>\n<p><strong>Time-budget-driven design.</strong> Using Protenix N_SAMPLE=1 instead of 5 saved time, which was exactly the time needed for RNAPro to run. The constraint that felt limiting actually forced a better architecture.</p>\n<h2>What Didn't Work</h2>\n<ul>\n<li><strong>MSA in RNAPro</strong> — disabled to save time; didn't noticeably help when tested</li>\n<li><strong>Protenix N_SAMPLE=5 as templates</strong> — too slow, didn't finish within 8h</li>\n<li><strong>Running Protenix with MSA</strong> — no improvement, extra time cost</li>\n<li><strong>Post-processing constraints on neural net output</strong> — slightly degraded predictions; Protenix/RNAPro geometry is already valid</li>\n<li><strong>RBSSA (embedding-based template search)</strong> — RibonanzaNet2 embeddings + Smith-Waterman. Hurt when it replaced Protenix (0.33), marginal when additive(0.425) on public data</li>\n<li><strong>Finetuned Protenix (RNA3DB checkpoint)</strong> — scored 0.392; base model was better</li>\n<li><strong>DRfold2 / Boltz ensemble</strong> — architecturally different but not better on any target; selector correctly ignored them</li>\n</ul>\n<h2>Key Lessons</h2>\n<ol>\n<li><strong>Template quality is the foundation.</strong> Every DL model performed better with good structural starting points. Protenix-generated templates were better than pure TBM for RNAPro conditioning.</li>\n<li><strong>DL models are template refiners on Kaggle hardware.</strong> On P100 (16GB), no model reliably folded RNA from scratch. They performed best refining structural starting points.</li>\n<li><strong>Time budget shapes architecture.</strong> The 8h limit forced Protenix N_SAMPLE=1, which turned out to be the right call — one good template + RNAPro refinement beat five mediocre Protenix predictions.</li>\n</ol>\n<hr>\n<h2>Notebook 2: Protenix + TBM (Private: 0.410)</h2>\n<p>Brief summary of the second selected notebook:</p>\n<ul>\n<li>TBM with 50% identity threshold, diversity transforms (hinge, jitter, wiggle)</li>\n<li>Protenix multi-seed [42, 137] for targets below threshold</li>\n<li><strong>Consistency-based ensemble selection</strong>: anchor on consensus Protenix prediction (lowest avg RMSD to others), guarantee one TBM slot, fill remaining by RMSD consistency. This was the single biggest improvement in the Protenix pipeline (0.428→0.440 on public).</li>\n</ul>\n<h2>Acknowledgments</h2>\n<p>Thanks to the community for shared resources — Protenix notebooks, RNAPro from NVIDIA, and TBM approaches from Part 1 winners. Excited about the CASP17 collaboration opportunity.</p>\n<p><strong>Kaggle:</strong> <a href=\"https://www.kaggle.com/error1249x\" target=\"_blank\">@error1249x</a></p>",
      "rawMarkdown": "Thank you to the organizers at Stanford, HHMI, NVIDIA, and Kaggle for this incredible competition. As someone with no formal biology or bioinformatics background, this was my first encounter with RNA structure prediction — and it turned into one of the most rewarding learning experiences I've had. Congratulations to all participants \n\n## The Core Idea\nUse Protenix as a **template generator**, not as the final predictor. Protenix produces one structural prediction per target, which gets fed into RNAPro as template conditioning. RNAPro then refines it using RibonanzaNet2 embeddings and its Pairformer architecture. TBM fills the remaining template slots to give RNAPro structural diversity to work with.\n\nThe time constraint shaped the design: running Protenix with N_SAMPLE=5 was too slow, so I used N_SAMPLE=1 and let RNAPro do the heavy lifting for diversity.\n\n## Pipeline Overview\n\n```\n                    ┌──────────────┐\n                    │   Test Seqs  │\n                    └──────┬───────┘\n                           │\n              ┌────────────┴────────────┐\n              ▼                         ▼\n      ┌──────────────┐         ┌──────────────┐\n      │     TBM      │         │   Protenix   │\n      │  (top 5 by   │         │  (N_SAMPLE=1 │\n      │  alignment)  │         │  per target) │\n      └──────┬───────┘         └──────┬───────┘\n             │                        │\n        TBM[1..4]              Protenix[0]\n             │                        │\n             └───────────┬────────────┘\n                         ▼\n              ┌─────────────────────┐\n              │   Template CSV      │\n              │  Slot 1: Protenix   │\n              │  Slot 2-5: TBM     │\n              └─────────┬───────────┘\n                        ▼\n              ┌─────────────────────┐\n              │      RNAPro         │\n              │  (template_idx=0,   │\n              │   RibonanzaNet2,    │\n              │   N_SAMPLE=3)       │\n              └─────────┬───────────┘\n                        ▼\n              ┌─────────────────────┐\n              │   5 Predictions     │\n              ├─────────────────────┤\n              │ P1-P3: RNAPro       │\n              │ P4-P5: Protenix/TBM │\n              │        fallback     │\n              └─────────────────────┘\n```\n\n## Phase 1 — Template-Based Modeling\n\nTBM serves two purposes: backup predictions and template slots 2-5 for RNAPro input.\n\n**Template search:** BioPython `PairwiseAligner` in global mode with strong gap penalties (`open: -8, extend: -0.4`). Searched ~5744 training+validation structures. Length ratio filter skips candidates with >30% length difference.\n\n**Coordinate transfer:** C1' coordinates mapped from template to query via alignment. Gaps filled by linear interpolation between nearest aligned positions, or extrapolated at ends with 3.8Å steps.\n\n**No quality threshold filtering** — all templates passed through to the CSV. RNAPro handles template quality internally through its confidence mechanism.\n\n## Phase 2 — Protenix (Template Generation)\n\nProtenix ran in minimal mode — the goal was one decent structural prediction per target, not five polished ones.\n\n- **N_SAMPLE=1** (single prediction per target)\n- **Single seed** (42)\n- **No MSA, RNA MSA enabled**\n- **Chunking** for sequences >512nt: overlapping windows with Kabsch alignment and cosine blending at boundaries (MAX_SEQ_LEN=512, CHUNK_OVERLAP=128)\n- C1' extraction via `centre_atom_mask` with `atom_to_tokatom_idx` fallback\n\nA Protenix-only backup `submission.csv` was saved before RNAPro ran — if RNAPro timed out, the notebook still produced a valid submission.\n\n## Phase 3 — Template CSV Construction\n\nFor each target, the 5-slot template CSV was built as:\n\n| Slot | Source | Purpose |\n|------|--------|---------|\n| 1 | Protenix prediction | Primary template — RNAPro refines this |\n| 2-5 | TBM templates | Structural diversity (or noise-augmented copies if <4 templates) |\n\nIf Protenix failed for a target, slot 1 fell back to the best TBM template.\n\nThe CSV was converted to `.pt` format via RNAPro's `convert_templates_to_pt_files.py`.\n\n## Phase 4 — RNAPro Inference\n\nRNAPro conditioned on the Protenix template via `--template_idx 0` (uses only the top template slot — our Protenix prediction). Its architecture combines:\n\n- **RibonanzaNet2** as frozen encoder — captures RNA-specific co-evolutionary features\n- **Template conditioning** — Protenix structural prior guides the diffusion\n- **Pairformer** — RNA-adapted structure module\n\nParams were aggressive to fit the 8-hour time budget:\n\n| Parameter | Value |\n|-----------|-------|\n| N_SAMPLE | 3 |\n| N_STEP | 100 |\n| N_CYCLE | 4 |\n| MAX_LEN | 500 |\n| MSA | disabled |\n| Seed | 42 |\n| template_idx | 0 |\n\n## Phase 5 — Submission Assembly\n\nSimple priority filling:\n1. RNAPro CIF outputs → slots 1-3 (if available)\n2. Protenix/TBM from backup CSV → remaining slots\n3. Helix fallback → last resort (never triggered on public test)\n\nNo ensemble selection or post-processing — just straight slot filling.\n\n## What Made the Difference\n\n**Protenix as template generator, not final predictor.** The Protenix→RNAPro pipeline scored 0.464 private vs 0.410 for standalone Protenix. RNAPro's RibonanzaNet2 embeddings added RNA-specific knowledge that base Protenix lacks, especially on novel targets without close homologs.\n\n**Time-budget-driven design.** Using Protenix N_SAMPLE=1 instead of 5 saved time, which was exactly the time needed for RNAPro to run. The constraint that felt limiting actually forced a better architecture.\n\n## What Didn't Work\n\n- **MSA in RNAPro** — disabled to save time; didn't noticeably help when tested\n- **Protenix N_SAMPLE=5 as templates** — too slow, didn't finish within 8h\n- **Running Protenix with MSA** — no improvement, extra time cost\n- **Post-processing constraints on neural net output** — slightly degraded predictions; Protenix/RNAPro geometry is already valid\n- **RBSSA (embedding-based template search)** — RibonanzaNet2 embeddings + Smith-Waterman. Hurt when it replaced Protenix (0.33), marginal when additive(0.425) on public data\n- **Finetuned Protenix (RNA3DB checkpoint)** — scored 0.392; base model was better\n- **DRfold2 / Boltz ensemble** — architecturally different but not better on any target; selector correctly ignored them\n\n## Key Lessons\n\n1. **Template quality is the foundation.** Every DL model performed better with good structural starting points. Protenix-generated templates were better than pure TBM for RNAPro conditioning.\n2. **DL models are template refiners on Kaggle hardware.** On P100 (16GB), no model reliably folded RNA from scratch. They performed best refining structural starting points.\n3. **Time budget shapes architecture.** The 8h limit forced Protenix N_SAMPLE=1, which turned out to be the right call — one good template + RNAPro refinement beat five mediocre Protenix predictions.\n\n---\n\n## Notebook 2: Protenix + TBM (Private: 0.410)\n\nBrief summary of the second selected notebook:\n- TBM with 50% identity threshold, diversity transforms (hinge, jitter, wiggle)\n- Protenix multi-seed [42, 137] for targets below threshold\n- **Consistency-based ensemble selection**: anchor on consensus Protenix prediction (lowest avg RMSD to others), guarantee one TBM slot, fill remaining by RMSD consistency. This was the single biggest improvement in the Protenix pipeline (0.428→0.440 on public).\n\n## Acknowledgments\n\nThanks to the community for shared resources — Protenix notebooks, RNAPro from NVIDIA, and TBM approaches from Part 1 winners. Excited about the CASP17 collaboration opportunity.\n\n**Kaggle:** [@error1249x](https://www.kaggle.com/error1249x)",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3434197": "Thank you to the organizers at Stanford, HHMI, NVIDIA, and Kaggle for this incredible competition. As someone with no formal biology or bioinformatics background, this was my first encounter with RNA structure prediction — and it turned into one of the most rewarding learning experiences I've had. Congratulations to all participants \n\n## The Core Idea\nUse Protenix as a **template generator**, not as the final predictor. Protenix produces one structural prediction per target, which gets fed into RNAPro as template conditioning. RNAPro then refines it using RibonanzaNet2 embeddings and its Pairformer architecture. TBM fills the remaining template slots to give RNAPro structural diversity to work with.\n\nThe time constraint shaped the design: running Protenix with N_SAMPLE=5 was too slow, so I used N_SAMPLE=1 and let RNAPro do the heavy lifting for diversity.\n\n## Pipeline Overview\n\n```\n                    ┌──────────────┐\n                    │   Test Seqs  │\n                    └──────┬───────┘\n                           │\n              ┌────────────┴────────────┐\n              ▼                         ▼\n      ┌──────────────┐         ┌──────────────┐\n      │     TBM      │         │   Protenix   │\n      │  (top 5 by   │         │  (N_SAMPLE=1 │\n      │  alignment)  │         │  per target) │\n      └──────┬───────┘         └──────┬───────┘\n             │                        │\n        TBM[1..4]              Protenix[0]\n             │                        │\n             └───────────┬────────────┘\n                         ▼\n              ┌─────────────────────┐\n              │   Template CSV      │\n              │  Slot 1: Protenix   │\n              │  Slot 2-5: TBM     │\n              └─────────┬───────────┘\n                        ▼\n              ┌─────────────────────┐\n              │      RNAPro         │\n              │  (template_idx=0,   │\n              │   RibonanzaNet2,    │\n              │   N_SAMPLE=3)       │\n              └─────────┬───────────┘\n                        ▼\n              ┌─────────────────────┐\n              │   5 Predictions     │\n              ├─────────────────────┤\n              │ P1-P3: RNAPro       │\n              │ P4-P5: Protenix/TBM │\n              │        fallback     │\n              └─────────────────────┘\n```\n\n## Phase 1 — Template-Based Modeling\n\nTBM serves two purposes: backup predictions and template slots 2-5 for RNAPro input.\n\n**Template search:** BioPython `PairwiseAligner` in global mode with strong gap penalties (`open: -8, extend: -0.4`). Searched ~5744 training+validation structures. Length ratio filter skips candidates with >30% length difference.\n\n**Coordinate transfer:** C1' coordinates mapped from template to query via alignment. Gaps filled by linear interpolation between nearest aligned positions, or extrapolated at ends with 3.8Å steps.\n\n**No quality threshold filtering** — all templates passed through to the CSV. RNAPro handles template quality internally through its confidence mechanism.\n\n## Phase 2 — Protenix (Template Generation)\n\nProtenix ran in minimal mode — the goal was one decent structural prediction per target, not five polished ones.\n\n- **N_SAMPLE=1** (single prediction per target)\n- **Single seed** (42)\n- **No MSA, RNA MSA enabled**\n- **Chunking** for sequences >512nt: overlapping windows with Kabsch alignment and cosine blending at boundaries (MAX_SEQ_LEN=512, CHUNK_OVERLAP=128)\n- C1' extraction via `centre_atom_mask` with `atom_to_tokatom_idx` fallback\n\nA Protenix-only backup `submission.csv` was saved before RNAPro ran — if RNAPro timed out, the notebook still produced a valid submission.\n\n## Phase 3 — Template CSV Construction\n\nFor each target, the 5-slot template CSV was built as:\n\n| Slot | Source | Purpose |\n|------|--------|---------|\n| 1 | Protenix prediction | Primary template — RNAPro refines this |\n| 2-5 | TBM templates | Structural diversity (or noise-augmented copies if <4 templates) |\n\nIf Protenix failed for a target, slot 1 fell back to the best TBM template.\n\nThe CSV was converted to `.pt` format via RNAPro's `convert_templates_to_pt_files.py`.\n\n## Phase 4 — RNAPro Inference\n\nRNAPro conditioned on the Protenix template via `--template_idx 0` (uses only the top template slot — our Protenix prediction). Its architecture combines:\n\n- **RibonanzaNet2** as frozen encoder — captures RNA-specific co-evolutionary features\n- **Template conditioning** — Protenix structural prior guides the diffusion\n- **Pairformer** — RNA-adapted structure module\n\nParams were aggressive to fit the 8-hour time budget:\n\n| Parameter | Value |\n|-----------|-------|\n| N_SAMPLE | 3 |\n| N_STEP | 100 |\n| N_CYCLE | 4 |\n| MAX_LEN | 500 |\n| MSA | disabled |\n| Seed | 42 |\n| template_idx | 0 |\n\n## Phase 5 — Submission Assembly\n\nSimple priority filling:\n1. RNAPro CIF outputs → slots 1-3 (if available)\n2. Protenix/TBM from backup CSV → remaining slots\n3. Helix fallback → last resort (never triggered on public test)\n\nNo ensemble selection or post-processing — just straight slot filling.\n\n## What Made the Difference\n\n**Protenix as template generator, not final predictor.** The Protenix→RNAPro pipeline scored 0.464 private vs 0.410 for standalone Protenix. RNAPro's RibonanzaNet2 embeddings added RNA-specific knowledge that base Protenix lacks, especially on novel targets without close homologs.\n\n**Time-budget-driven design.** Using Protenix N_SAMPLE=1 instead of 5 saved time, which was exactly the time needed for RNAPro to run. The constraint that felt limiting actually forced a better architecture.\n\n## What Didn't Work\n\n- **MSA in RNAPro** — disabled to save time; didn't noticeably help when tested\n- **Protenix N_SAMPLE=5 as templates** — too slow, didn't finish within 8h\n- **Running Protenix with MSA** — no improvement, extra time cost\n- **Post-processing constraints on neural net output** — slightly degraded predictions; Protenix/RNAPro geometry is already valid\n- **RBSSA (embedding-based template search)** — RibonanzaNet2 embeddings + Smith-Waterman. Hurt when it replaced Protenix (0.33), marginal when additive(0.425) on public data\n- **Finetuned Protenix (RNA3DB checkpoint)** — scored 0.392; base model was better\n- **DRfold2 / Boltz ensemble** — architecturally different but not better on any target; selector correctly ignored them\n\n## Key Lessons\n\n1. **Template quality is the foundation.** Every DL model performed better with good structural starting points. Protenix-generated templates were better than pure TBM for RNAPro conditioning.\n2. **DL models are template refiners on Kaggle hardware.** On P100 (16GB), no model reliably folded RNA from scratch. They performed best refining structural starting points.\n3. **Time budget shapes architecture.** The 8h limit forced Protenix N_SAMPLE=1, which turned out to be the right call — one good template + RNAPro refinement beat five mediocre Protenix predictions.\n\n---\n\n## Notebook 2: Protenix + TBM (Private: 0.410)\n\nBrief summary of the second selected notebook:\n- TBM with 50% identity threshold, diversity transforms (hinge, jitter, wiggle)\n- Protenix multi-seed [42, 137] for targets below threshold\n- **Consistency-based ensemble selection**: anchor on consensus Protenix prediction (lowest avg RMSD to others), guarantee one TBM slot, fill remaining by RMSD consistency. This was the single biggest improvement in the Protenix pipeline (0.428→0.440 on public).\n\n## Acknowledgments\n\nThanks to the community for shared resources — Protenix notebooks, RNAPro from NVIDIA, and TBM approaches from Part 1 winners. Excited about the CASP17 collaboration opportunity.\n\n**Kaggle:** [@error1249x](https://www.kaggle.com/error1249x)"
  }
}