{
  "id": 686777,
  "title": "6th Place Solution",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/686777",
  "author_name": "Shivam Shinde",
  "post_date": "2026-04-01T05:52:38.321000",
  "votes": 23,
  "comment_count": 3,
  "views": 0,
  "content": "<p>` Thanks everyone for a great competition. Congrats to all the winners!</p>\n<hr>\n<h2>The Core Idea</h2>\n<p>The competition scores the best of 5 predictions, meaning diversity matters as much as accuracy. My strategy was to first generate diverse structural hypotheses cheaply using Template-Based Modeling (TBM). I then used Protenix in no-MSA, no-template mode to produce independent neural predictions. These outputs were later fed into RNAPro as templates, allowing the model to explore different regions of structure space. This cross-pollination between models turned out to be the key.</p>\n<hr>\n<h2>Pipeline Overview</h2>\n<pre><code>                         ┌──────────────────┐\n                         │    Test Seqs     │\n                         └────────┬─────────┘\n                                  │\n                  ┌───────────────┴───────────────┐\n                  │                               │\n           seq &gt; 1000 nt                   seq ≤ 1000 nt\n                  │                               │\n                  │                ┌──────────────┴──────────────┐\n                  │                ▼                             ▼\n                  │         ┌─────────────┐             ┌─────────────┐\n                  │         │     TBM     │             │   Protenix  │\n                  │         │  (fast, 5   │             │  (no TBM,   │\n                  │         │   diverse)  │             │  no MSA)    │\n                  │         └──────┬──────┘             └──────┬──────┘\n                  │                │                           │\n                  │         TBM[0..4]              Protenix[0..1] as templates\n                  │                │                           │\n                  │                └─────────────┬─────────────┘\n                  │                              ▼\n                  │                      ┌──────────────┐\n                  │                      │    RNAPro    │\n                  │                      │  (refine +   │\n                  │                      │   MSA +      │\n                  │                      │ RibonanzaNet)│\n                  │                      └──────┬───────┘\n                  │                             │\n                  └──────────────┬──────────────┘\n                                 ▼\n          ┌──────────────────────────────────────────┐\n          │              5 Predictions               │\n          ├──────────────────────────────────────────┤\n          │  P1 : RNAPro  +  TBM template [0]        │\n          │  P2 : RNAPro  +  TBM template [1]        │\n          │  P3 : RNAPro  +  Protenix template [0]   │\n          │  P4 : RNAPro  +  Protenix template [1]   │\n          │  P5 : Pure Protenix [0]  (no RNAPro)     │\n          ├──────────────────────────────────────────┤\n          │  ⚠️  seq &gt; 1000 nt → TBM × 5 (all slots) │\n          └──────────────────────────────────────────┘\n</code></pre>\n<hr>\n<h2>Phase 1 — Template-Based Modeling (TBM)</h2>\n<p>This runs first and serves two roles: a standalone fallback for long sequences, and a source of structural templates for RNAPro.</p>\n<p><strong>Template search:</strong> I use BioPython's <code>PairwiseAligner</code> in global mode. The  strong gap penalties (<code>open: -8, extend: -0.4</code>) including at the terminals. </p>\n<p><strong>Coordinate transfer:</strong> Once aligned, C1' coordinates from the training structure are mapped onto the query residues directly. Gaps are filled by linear interpolation between neighboring residues, or extrapolated at the ends with a fixed 3Å step.</p>\n<p><strong>Geometry refinement:</strong> After transfer I apply <code>adaptive_rna_constraints</code> which runs within each chain segment independently (important for multi-chain targets — you don't want fake bonds across chain breaks):</p>\n<ul>\n<li>i↔i+1 bond → ~5.95 Å</li>\n<li>i↔i+2 soft angle → ~10.20 Å</li>\n<li>Laplacian smoothing to kill kinks</li>\n<li>Light steric self-avoidance for chains ≥ 25 residues</li>\n</ul>\n<p>Correction strength scales with <code>(1 - similarity)</code> — if the template is a close match, barely touch it. If it's a rough match, correct more aggressively.</p>\n<p><strong>Getting 5 diverse predictions from one template:</strong></p>\n<table>\n<thead>\n<tr>\n<th>Pred</th>\n<th>Transform</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>Best template, untouched</td>\n</tr>\n<tr>\n<td>1</td>\n<td>Mild Gaussian noise scaled by (1 − similarity)</td>\n</tr>\n<tr>\n<td>2</td>\n<td>Hinge rotation on the longest chain segment</td>\n</tr>\n<tr>\n<td>3</td>\n<td>Independent rigid-body jitter per chain</td>\n</tr>\n<tr>\n<td>4</td>\n<td>Smooth low-frequency wiggle via interpolated control points</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>Phase 2 — Protenix</h2>\n<p>I used <a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">Protenix</a> (ByteDance's AlphaFold3-style model) in <strong>no-MSA, no-template</strong> mode. The goal here wasn't to get the best individual prediction — it was to get something structurally <em>different</em> from TBM that RNAPro could use as an alternative starting point.</p>\n<ul>\n<li>Applied to sequences ≤ 1000 nt (truncated to 512 for featurization)</li>\n<li>5 samples generated; first 2 are passed forward as templates</li>\n<li>C1' atoms extracted via <code>centre_atom_mask</code> with fallback to <code>atom_to_tokatom_idx</code></li>\n</ul>\n<hr>\n<h2>Phase 3 — RNAPro</h2>\n<p>This is where the cross-pollination happens. I precompute a 4-slot template file combining TBM and Protenix outputs, then run RNAPro once per slot:</p>\n<table>\n<thead>\n<tr>\n<th>Slot</th>\n<th>Template Source</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>TBM prediction 1</td>\n</tr>\n<tr>\n<td>1</td>\n<td>TBM prediction 2</td>\n</tr>\n<tr>\n<td>2</td>\n<td>Protenix prediction 1</td>\n</tr>\n<tr>\n<td>3</td>\n<td>Protenix prediction 2</td>\n</tr>\n</tbody>\n</table>\n<p>RNAPro runs with MSA + RibonanzaNet2 embeddings, 10 recycling cycles, 200 diffusion steps. Each run is conditioned on a different template, so you get 4 structurally distinct refined outputs.</p>\n<p>Prediction 5 is just raw Protenix[0] — no RNAPro. I kept it because on some targets (well-studied motifs with close training analogs) the direct Protenix output was hard to beat.</p>\n<hr>\n<h2>What Actually Made the Difference</h2>\n<p><strong>Feeding Protenix outputs into RNAPro as templates.</strong> Protenix and RNAPro have genuinely different inductive biases. RNAPro conditioned on a Protenix template explores a different part of structure space than RNAPro conditioned on a TBM template. This gave predictions 3 and 4 real independence from predictions 1 and 2.</p>\n<p><strong>Not throwing away raw Protenix.</strong> Prediction 5 being unrefined saved several targets where RNAPro overcorrected.</p>\n<hr>\n<h2>What Didn't Work</h2>\n<ul>\n<li>Running Protenix with MSA — didn't help and cost time</li>\n<li>More than 2 TBM template slots in RNAPro — performance plateaued after slot 2</li>\n<li>Ensemble averaging of coordinates — worse than best-of-5 strategy for this metric</li>\n</ul>\n<hr>\n<h2>Hardware</h2>\n<ul>\n<li>GPU P100 (Kaggle)\n`</li>\n</ul>",
  "messages": [
    {
      "id": 3433045,
      "postDate": "2026-04-01T05:52:38.320Z",
      "content": "<p>` Thanks everyone for a great competition. Congrats to all the winners!</p>\n<hr>\n<h2>The Core Idea</h2>\n<p>The competition scores the best of 5 predictions, meaning diversity matters as much as accuracy. My strategy was to first generate diverse structural hypotheses cheaply using Template-Based Modeling (TBM). I then used Protenix in no-MSA, no-template mode to produce independent neural predictions. These outputs were later fed into RNAPro as templates, allowing the model to explore different regions of structure space. This cross-pollination between models turned out to be the key.</p>\n<hr>\n<h2>Pipeline Overview</h2>\n<pre><code>                         ┌──────────────────┐\n                         │    Test Seqs     │\n                         └────────┬─────────┘\n                                  │\n                  ┌───────────────┴───────────────┐\n                  │                               │\n           seq &gt; 1000 nt                   seq ≤ 1000 nt\n                  │                               │\n                  │                ┌──────────────┴──────────────┐\n                  │                ▼                             ▼\n                  │         ┌─────────────┐             ┌─────────────┐\n                  │         │     TBM     │             │   Protenix  │\n                  │         │  (fast, 5   │             │  (no TBM,   │\n                  │         │   diverse)  │             │  no MSA)    │\n                  │         └──────┬──────┘             └──────┬──────┘\n                  │                │                           │\n                  │         TBM[0..4]              Protenix[0..1] as templates\n                  │                │                           │\n                  │                └─────────────┬─────────────┘\n                  │                              ▼\n                  │                      ┌──────────────┐\n                  │                      │    RNAPro    │\n                  │                      │  (refine +   │\n                  │                      │   MSA +      │\n                  │                      │ RibonanzaNet)│\n                  │                      └──────┬───────┘\n                  │                             │\n                  └──────────────┬──────────────┘\n                                 ▼\n          ┌──────────────────────────────────────────┐\n          │              5 Predictions               │\n          ├──────────────────────────────────────────┤\n          │  P1 : RNAPro  +  TBM template [0]        │\n          │  P2 : RNAPro  +  TBM template [1]        │\n          │  P3 : RNAPro  +  Protenix template [0]   │\n          │  P4 : RNAPro  +  Protenix template [1]   │\n          │  P5 : Pure Protenix [0]  (no RNAPro)     │\n          ├──────────────────────────────────────────┤\n          │  ⚠️  seq &gt; 1000 nt → TBM × 5 (all slots) │\n          └──────────────────────────────────────────┘\n</code></pre>\n<hr>\n<h2>Phase 1 — Template-Based Modeling (TBM)</h2>\n<p>This runs first and serves two roles: a standalone fallback for long sequences, and a source of structural templates for RNAPro.</p>\n<p><strong>Template search:</strong> I use BioPython's <code>PairwiseAligner</code> in global mode. The  strong gap penalties (<code>open: -8, extend: -0.4</code>) including at the terminals. </p>\n<p><strong>Coordinate transfer:</strong> Once aligned, C1' coordinates from the training structure are mapped onto the query residues directly. Gaps are filled by linear interpolation between neighboring residues, or extrapolated at the ends with a fixed 3Å step.</p>\n<p><strong>Geometry refinement:</strong> After transfer I apply <code>adaptive_rna_constraints</code> which runs within each chain segment independently (important for multi-chain targets — you don't want fake bonds across chain breaks):</p>\n<ul>\n<li>i↔i+1 bond → ~5.95 Å</li>\n<li>i↔i+2 soft angle → ~10.20 Å</li>\n<li>Laplacian smoothing to kill kinks</li>\n<li>Light steric self-avoidance for chains ≥ 25 residues</li>\n</ul>\n<p>Correction strength scales with <code>(1 - similarity)</code> — if the template is a close match, barely touch it. If it's a rough match, correct more aggressively.</p>\n<p><strong>Getting 5 diverse predictions from one template:</strong></p>\n<table>\n<thead>\n<tr>\n<th>Pred</th>\n<th>Transform</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>Best template, untouched</td>\n</tr>\n<tr>\n<td>1</td>\n<td>Mild Gaussian noise scaled by (1 − similarity)</td>\n</tr>\n<tr>\n<td>2</td>\n<td>Hinge rotation on the longest chain segment</td>\n</tr>\n<tr>\n<td>3</td>\n<td>Independent rigid-body jitter per chain</td>\n</tr>\n<tr>\n<td>4</td>\n<td>Smooth low-frequency wiggle via interpolated control points</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>Phase 2 — Protenix</h2>\n<p>I used <a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">Protenix</a> (ByteDance's AlphaFold3-style model) in <strong>no-MSA, no-template</strong> mode. The goal here wasn't to get the best individual prediction — it was to get something structurally <em>different</em> from TBM that RNAPro could use as an alternative starting point.</p>\n<ul>\n<li>Applied to sequences ≤ 1000 nt (truncated to 512 for featurization)</li>\n<li>5 samples generated; first 2 are passed forward as templates</li>\n<li>C1' atoms extracted via <code>centre_atom_mask</code> with fallback to <code>atom_to_tokatom_idx</code></li>\n</ul>\n<hr>\n<h2>Phase 3 — RNAPro</h2>\n<p>This is where the cross-pollination happens. I precompute a 4-slot template file combining TBM and Protenix outputs, then run RNAPro once per slot:</p>\n<table>\n<thead>\n<tr>\n<th>Slot</th>\n<th>Template Source</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>TBM prediction 1</td>\n</tr>\n<tr>\n<td>1</td>\n<td>TBM prediction 2</td>\n</tr>\n<tr>\n<td>2</td>\n<td>Protenix prediction 1</td>\n</tr>\n<tr>\n<td>3</td>\n<td>Protenix prediction 2</td>\n</tr>\n</tbody>\n</table>\n<p>RNAPro runs with MSA + RibonanzaNet2 embeddings, 10 recycling cycles, 200 diffusion steps. Each run is conditioned on a different template, so you get 4 structurally distinct refined outputs.</p>\n<p>Prediction 5 is just raw Protenix[0] — no RNAPro. I kept it because on some targets (well-studied motifs with close training analogs) the direct Protenix output was hard to beat.</p>\n<hr>\n<h2>What Actually Made the Difference</h2>\n<p><strong>Feeding Protenix outputs into RNAPro as templates.</strong> Protenix and RNAPro have genuinely different inductive biases. RNAPro conditioned on a Protenix template explores a different part of structure space than RNAPro conditioned on a TBM template. This gave predictions 3 and 4 real independence from predictions 1 and 2.</p>\n<p><strong>Not throwing away raw Protenix.</strong> Prediction 5 being unrefined saved several targets where RNAPro overcorrected.</p>\n<hr>\n<h2>What Didn't Work</h2>\n<ul>\n<li>Running Protenix with MSA — didn't help and cost time</li>\n<li>More than 2 TBM template slots in RNAPro — performance plateaued after slot 2</li>\n<li>Ensemble averaging of coordinates — worse than best-of-5 strategy for this metric</li>\n</ul>\n<hr>\n<h2>Hardware</h2>\n<ul>\n<li>GPU P100 (Kaggle)\n`</li>\n</ul>",
      "rawMarkdown": "` Thanks everyone for a great competition. Congrats to all the winners!\n\n---\n\n## The Core Idea\n\nThe competition scores the best of 5 predictions, meaning diversity matters as much as accuracy. My strategy was to first generate diverse structural hypotheses cheaply using Template-Based Modeling (TBM). I then used Protenix in no-MSA, no-template mode to produce independent neural predictions. These outputs were later fed into RNAPro as templates, allowing the model to explore different regions of structure space. This cross-pollination between models turned out to be the key.\n\n---\n\n## Pipeline Overview\n\n```\n                         ┌──────────────────┐\n                         │    Test Seqs     │\n                         └────────┬─────────┘\n                                  │\n                  ┌───────────────┴───────────────┐\n                  │                               │\n           seq > 1000 nt                   seq ≤ 1000 nt\n                  │                               │\n                  │                ┌──────────────┴──────────────┐\n                  │                ▼                             ▼\n                  │         ┌─────────────┐             ┌─────────────┐\n                  │         │     TBM     │             │   Protenix  │\n                  │         │  (fast, 5   │             │  (no TBM,   │\n                  │         │   diverse)  │             │  no MSA)    │\n                  │         └──────┬──────┘             └──────┬──────┘\n                  │                │                           │\n                  │         TBM[0..4]              Protenix[0..1] as templates\n                  │                │                           │\n                  │                └─────────────┬─────────────┘\n                  │                              ▼\n                  │                      ┌──────────────┐\n                  │                      │    RNAPro    │\n                  │                      │  (refine +   │\n                  │                      │   MSA +      │\n                  │                      │ RibonanzaNet)│\n                  │                      └──────┬───────┘\n                  │                             │\n                  └──────────────┬──────────────┘\n                                 ▼\n          ┌──────────────────────────────────────────┐\n          │              5 Predictions               │\n          ├──────────────────────────────────────────┤\n          │  P1 : RNAPro  +  TBM template [0]        │\n          │  P2 : RNAPro  +  TBM template [1]        │\n          │  P3 : RNAPro  +  Protenix template [0]   │\n          │  P4 : RNAPro  +  Protenix template [1]   │\n          │  P5 : Pure Protenix [0]  (no RNAPro)     │\n          ├──────────────────────────────────────────┤\n          │  ⚠️  seq > 1000 nt → TBM × 5 (all slots) │\n          └──────────────────────────────────────────┘\n```\n\n---\n\n## Phase 1 — Template-Based Modeling (TBM)\n\nThis runs first and serves two roles: a standalone fallback for long sequences, and a source of structural templates for RNAPro.\n\n**Template search:** I use BioPython's `PairwiseAligner` in global mode. The  strong gap penalties (`open: -8, extend: -0.4`) including at the terminals. \n\n**Coordinate transfer:** Once aligned, C1' coordinates from the training structure are mapped onto the query residues directly. Gaps are filled by linear interpolation between neighboring residues, or extrapolated at the ends with a fixed 3Å step.\n\n**Geometry refinement:** After transfer I apply `adaptive_rna_constraints` which runs within each chain segment independently (important for multi-chain targets — you don't want fake bonds across chain breaks):\n- i↔i+1 bond → ~5.95 Å\n- i↔i+2 soft angle → ~10.20 Å\n- Laplacian smoothing to kill kinks\n- Light steric self-avoidance for chains ≥ 25 residues\n\nCorrection strength scales with `(1 - similarity)` — if the template is a close match, barely touch it. If it's a rough match, correct more aggressively.\n\n**Getting 5 diverse predictions from one template:**\n\n| Pred | Transform |\n|------|-----------|\n| 0 | Best template, untouched |\n| 1 | Mild Gaussian noise scaled by (1 − similarity) |\n| 2 | Hinge rotation on the longest chain segment |\n| 3 | Independent rigid-body jitter per chain |\n| 4 | Smooth low-frequency wiggle via interpolated control points |\n\n---\n\n## Phase 2 — Protenix\n\nI used [Protenix](https://github.com/bytedance/Protenix) (ByteDance's AlphaFold3-style model) in **no-MSA, no-template** mode. The goal here wasn't to get the best individual prediction — it was to get something structurally *different* from TBM that RNAPro could use as an alternative starting point.\n\n- Applied to sequences ≤ 1000 nt (truncated to 512 for featurization)\n- 5 samples generated; first 2 are passed forward as templates\n- C1' atoms extracted via `centre_atom_mask` with fallback to `atom_to_tokatom_idx`\n\n---\n\n## Phase 3 — RNAPro\n\nThis is where the cross-pollination happens. I precompute a 4-slot template file combining TBM and Protenix outputs, then run RNAPro once per slot:\n\n| Slot | Template Source |\n|------|----------------|\n| 0 | TBM prediction 1 |\n| 1 | TBM prediction 2 |\n| 2 | Protenix prediction 1 |\n| 3 | Protenix prediction 2 |\n\nRNAPro runs with MSA + RibonanzaNet2 embeddings, 10 recycling cycles, 200 diffusion steps. Each run is conditioned on a different template, so you get 4 structurally distinct refined outputs.\n\nPrediction 5 is just raw Protenix[0] — no RNAPro. I kept it because on some targets (well-studied motifs with close training analogs) the direct Protenix output was hard to beat.\n\n---\n\n## What Actually Made the Difference\n\n\n**Feeding Protenix outputs into RNAPro as templates.** Protenix and RNAPro have genuinely different inductive biases. RNAPro conditioned on a Protenix template explores a different part of structure space than RNAPro conditioned on a TBM template. This gave predictions 3 and 4 real independence from predictions 1 and 2.\n\n**Not throwing away raw Protenix.** Prediction 5 being unrefined saved several targets where RNAPro overcorrected.\n\n---\n\n## What Didn't Work\n\n- Running Protenix with MSA — didn't help and cost time\n- More than 2 TBM template slots in RNAPro — performance plateaued after slot 2\n- Ensemble averaging of coordinates — worse than best-of-5 strategy for this metric\n\n---\n\n## Hardware \n\n-  GPU P100 (Kaggle)\n`",
      "votes": 23
    },
    {
      "id": 3433065,
      "postDate": "2026-04-01T06:22:04.070Z",
      "content": "<p>The same notebook has scored 0.583 on temporary private evaluation set (during competition)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F27044890%2Fa828efd3251faf1af2811d6912cf0214%2FScreenshot%202026-04-01%20at%2011.45.49AM.png?generation=1775024494284908&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "The same notebook has scored 0.583 on temporary private evaluation set (during competition)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F27044890%2Fa828efd3251faf1af2811d6912cf0214%2FScreenshot%202026-04-01%20at%2011.45.49AM.png?generation=1775024494284908&alt=media)",
      "votes": 1
    },
    {
      "id": 3433484,
      "postDate": "2026-04-01T18:32:48.497Z",
      "content": "<p>Great work!</p>",
      "rawMarkdown": "Great work!",
      "votes": 1
    },
    {
      "id": 3433075,
      "postDate": "2026-04-01T06:41:08.107Z",
      "content": "<p>Thanks for sharing. Congrats!</p>",
      "rawMarkdown": "Thanks for sharing. Congrats!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 3433065,
      "author_name": "Shivam Shinde",
      "author_url": "",
      "post_date": "2026-04-01T06:22:04.070000",
      "content": "<p>The same notebook has scored 0.583 on temporary private evaluation set (during competition)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F27044890%2Fa828efd3251faf1af2811d6912cf0214%2FScreenshot%202026-04-01%20at%2011.45.49AM.png?generation=1775024494284908&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3433484,
      "author_name": "Ethan Kershner",
      "author_url": "",
      "post_date": "2026-04-01T18:32:48.497000",
      "content": "<p>Great work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3433075,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2026-04-01T06:41:08.107000",
      "content": "<p>Thanks for sharing. Congrats!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3433045": "` Thanks everyone for a great competition. Congrats to all the winners!\n\n---\n\n## The Core Idea\n\nThe competition scores the best of 5 predictions, meaning diversity matters as much as accuracy. My strategy was to first generate diverse structural hypotheses cheaply using Template-Based Modeling (TBM). I then used Protenix in no-MSA, no-template mode to produce independent neural predictions. These outputs were later fed into RNAPro as templates, allowing the model to explore different regions of structure space. This cross-pollination between models turned out to be the key.\n\n---\n\n## Pipeline Overview\n\n```\n                         ┌──────────────────┐\n                         │    Test Seqs     │\n                         └────────┬─────────┘\n                                  │\n                  ┌───────────────┴───────────────┐\n                  │                               │\n           seq > 1000 nt                   seq ≤ 1000 nt\n                  │                               │\n                  │                ┌──────────────┴──────────────┐\n                  │                ▼                             ▼\n                  │         ┌─────────────┐             ┌─────────────┐\n                  │         │     TBM     │             │   Protenix  │\n                  │         │  (fast, 5   │             │  (no TBM,   │\n                  │         │   diverse)  │             │  no MSA)    │\n                  │         └──────┬──────┘             └──────┬──────┘\n                  │                │                           │\n                  │         TBM[0..4]              Protenix[0..1] as templates\n                  │                │                           │\n                  │                └─────────────┬─────────────┘\n                  │                              ▼\n                  │                      ┌──────────────┐\n                  │                      │    RNAPro    │\n                  │                      │  (refine +   │\n                  │                      │   MSA +      │\n                  │                      │ RibonanzaNet)│\n                  │                      └──────┬───────┘\n                  │                             │\n                  └──────────────┬──────────────┘\n                                 ▼\n          ┌──────────────────────────────────────────┐\n          │              5 Predictions               │\n          ├──────────────────────────────────────────┤\n          │  P1 : RNAPro  +  TBM template [0]        │\n          │  P2 : RNAPro  +  TBM template [1]        │\n          │  P3 : RNAPro  +  Protenix template [0]   │\n          │  P4 : RNAPro  +  Protenix template [1]   │\n          │  P5 : Pure Protenix [0]  (no RNAPro)     │\n          ├──────────────────────────────────────────┤\n          │  ⚠️  seq > 1000 nt → TBM × 5 (all slots) │\n          └──────────────────────────────────────────┘\n```\n\n---\n\n## Phase 1 — Template-Based Modeling (TBM)\n\nThis runs first and serves two roles: a standalone fallback for long sequences, and a source of structural templates for RNAPro.\n\n**Template search:** I use BioPython's `PairwiseAligner` in global mode. The  strong gap penalties (`open: -8, extend: -0.4`) including at the terminals. \n\n**Coordinate transfer:** Once aligned, C1' coordinates from the training structure are mapped onto the query residues directly. Gaps are filled by linear interpolation between neighboring residues, or extrapolated at the ends with a fixed 3Å step.\n\n**Geometry refinement:** After transfer I apply `adaptive_rna_constraints` which runs within each chain segment independently (important for multi-chain targets — you don't want fake bonds across chain breaks):\n- i↔i+1 bond → ~5.95 Å\n- i↔i+2 soft angle → ~10.20 Å\n- Laplacian smoothing to kill kinks\n- Light steric self-avoidance for chains ≥ 25 residues\n\nCorrection strength scales with `(1 - similarity)` — if the template is a close match, barely touch it. If it's a rough match, correct more aggressively.\n\n**Getting 5 diverse predictions from one template:**\n\n| Pred | Transform |\n|------|-----------|\n| 0 | Best template, untouched |\n| 1 | Mild Gaussian noise scaled by (1 − similarity) |\n| 2 | Hinge rotation on the longest chain segment |\n| 3 | Independent rigid-body jitter per chain |\n| 4 | Smooth low-frequency wiggle via interpolated control points |\n\n---\n\n## Phase 2 — Protenix\n\nI used [Protenix](https://github.com/bytedance/Protenix) (ByteDance's AlphaFold3-style model) in **no-MSA, no-template** mode. The goal here wasn't to get the best individual prediction — it was to get something structurally *different* from TBM that RNAPro could use as an alternative starting point.\n\n- Applied to sequences ≤ 1000 nt (truncated to 512 for featurization)\n- 5 samples generated; first 2 are passed forward as templates\n- C1' atoms extracted via `centre_atom_mask` with fallback to `atom_to_tokatom_idx`\n\n---\n\n## Phase 3 — RNAPro\n\nThis is where the cross-pollination happens. I precompute a 4-slot template file combining TBM and Protenix outputs, then run RNAPro once per slot:\n\n| Slot | Template Source |\n|------|----------------|\n| 0 | TBM prediction 1 |\n| 1 | TBM prediction 2 |\n| 2 | Protenix prediction 1 |\n| 3 | Protenix prediction 2 |\n\nRNAPro runs with MSA + RibonanzaNet2 embeddings, 10 recycling cycles, 200 diffusion steps. Each run is conditioned on a different template, so you get 4 structurally distinct refined outputs.\n\nPrediction 5 is just raw Protenix[0] — no RNAPro. I kept it because on some targets (well-studied motifs with close training analogs) the direct Protenix output was hard to beat.\n\n---\n\n## What Actually Made the Difference\n\n\n**Feeding Protenix outputs into RNAPro as templates.** Protenix and RNAPro have genuinely different inductive biases. RNAPro conditioned on a Protenix template explores a different part of structure space than RNAPro conditioned on a TBM template. This gave predictions 3 and 4 real independence from predictions 1 and 2.\n\n**Not throwing away raw Protenix.** Prediction 5 being unrefined saved several targets where RNAPro overcorrected.\n\n---\n\n## What Didn't Work\n\n- Running Protenix with MSA — didn't help and cost time\n- More than 2 TBM template slots in RNAPro — performance plateaued after slot 2\n- Ensemble averaging of coordinates — worse than best-of-5 strategy for this metric\n\n---\n\n## Hardware \n\n-  GPU P100 (Kaggle)\n`",
    "3433065": "The same notebook has scored 0.583 on temporary private evaluation set (during competition)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F27044890%2Fa828efd3251faf1af2811d6912cf0214%2FScreenshot%202026-04-01%20at%2011.45.49AM.png?generation=1775024494284908&alt=media)",
    "3433484": "Great work!",
    "3433075": "Thanks for sharing. Congrats!"
  }
}