{
  "id": 687756,
  "title": "9th place solution: Optimize diversity when representation is limited",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/687756",
  "author_name": "hongan",
  "post_date": "2026-04-03T15:43:39.693000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First, I would like to thank Kaggle and the host for this interesting challenge. Second, I want to give credits to authors of these incredible public notebooks, on which my solution was built: <a href=\"https://www.kaggle.com/code/lbugnon/boltz2-baseline\" target=\"_blank\">Boltz2_baseline</a>, <a href=\"https://www.kaggle.com/code/llkh0a/stanford-rna-3d-folding-part-2-protenix-tbm\" target=\"_blank\">Protenix+TBM</a>, <a href=\"https://www.kaggle.com/code/jaejohn/rnapro-inference-with-tbm\" target=\"_blank\">RNAPro inference with TBM\n</a>, <a href=\"https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference\" target=\"_blank\">RNAPro Inference</a></p>\n<h2>Templates</h2>\n<p>The representation of RNA that is available is its sequence. That's the only thing I can work with really. So, to find the best templates, I need to first understand what templates are \"good\" in terms of their sequence representation. Thus, I did a templates audit on <code>validation_sequences</code>, trying to find their best templates from <code>train_sequences</code>. Here is an example of what I found: for sequence <code>8ZNQ</code> (30nt) the best templates seems to be <code>1AKX</code>, which have a TM score of 0.27864. However, using biopython sequence aligner, the sequence similarity will be -0.175, which means there is no chance to find this template only using sequence aligner. This is what referred to as short \"template-free\" targets discussed by the host. For longer sequences, aligner seems to work better.</p>\n<p>The conclusion here is clear, given representation of RNA only as sequences, there is a limit to the quality of templates I can find. I did experiment other TBM methods (primary chain + side chain search). Some worked well in my local validation but performs poorly on the LB when combined with DL models so I did not end up using. Because \"hard\" sequences cannot be found by aligner and \"easy\" sequences can almost be found by any aligner, the TBM used in my final submission are public TBM methods with very little modification.</p>\n<table>\n<thead>\n<tr>\n<th>method</th>\n<th>local validation</th>\n<th>Public LB (before rerun)</th>\n<th>Private LB (before rerun)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>primary chain (selected)</td>\n<td>0.33840</td>\n<td>0.32006</td>\n<td>0.46044</td>\n</tr>\n<tr>\n<td>primary chain john ver. (selected)</td>\n<td>0.31873</td>\n<td>0.34932</td>\n<td>0.45131</td>\n</tr>\n<tr>\n<td>primary chain + side chain</td>\n<td>0.40150</td>\n<td>0.32790</td>\n<td>0.45089</td>\n</tr>\n</tbody>\n</table>\n<h2>Overall prediction pipeline</h2>\n<p>With a standard TBM method, my model then pivots to create as much diversity in my predictions as possible. The final model is very similar to other shared solutions (TBM+protenix+boltz2+rnapro refinement). The difference is that for TBM and Protenix generated solutions, I select the ones that are most different from each other, thus maximizing diversity in my 5 predictions. More details can be found in the flowchart below.</p>\n<pre><code>┌─────────────────────────────────────────────────────────┐\n│              📥  Input: Test RNA Sequences               │\n└──────────────────────────┬──────────────────────────────┘\n                           │\n                           ▼\n┌─────────────────────────────────────────────────────────┐\n│                  PHASE 1: Template-Based Modeling (TBM)  │\n│                                                         │\n│  1. Align each test sequence against 5,744 training     │\n│     structures using global pairwise alignment          │\n│                                                         │\n│  2. Filter by similarity &amp; percent identity (≥50%)      │\n│                                                         │\n│  3. Adapt template coordinates → query via alignment    │\n│     (interpolate/extrapolate gaps)                      │\n│                                                         │\n│  4. Apply diversity transforms to generate variants:    │\n│     • Hinge bending at random pivot points              │\n│     • Chain jittering (rotate + translate per chain)    │\n│     • Smooth wiggle (spline-based displacement)         │\n│                                                         │\n│  5. Greedy max-diversity selection                      │\n│     (quality × diversity trade-off)                     │\n│                                                         │\n│  Slot allocation by length:                             │\n│    Short (≤512 nt) → max 2 TBM, min 3 Protenix        │\n│    Long  (&gt;512 nt) → max 5 TBM, 0 min Protenix        │\n└──────────────────────────┬──────────────────────────────┘\n                           │\n                           ▼\n              ┌────────────────────────┐\n              │  Remaining slots =     │\n              │  5 − TBM count         │\n              └─────┬─────────┬────────┘\n                    │         │\n          ┌─────────┘         └──────────┐\n          ▼                              ▼\n┌───────────────────────┐  ┌──────────────────────────────┐\n│  PHASE 2: Protenix    │  │  PHASE 2.5: Boltz2           │\n│  (2× T4 GPUs)         │  │  (2× T4 GPUs, after Protenix)│\n│                       │  │                              │\n│  IF n ≥ 3 slots:      │  │  Eligible: seq &lt; 900 nt     │\n│  ┌──────────────────┐ │  │  &amp; slots still available     │\n│  │ Phase A: Ensemble │ │  │                              │\n│  │ GPU 0 → MSA run  │ │  │  1. Write multi-chain FASTA  │\n│  │ GPU 1 → noMSA run│ │  │                              │\n│  │       ↓          │ │  │  2. Persistent GPU workers   │\n│  │ Kabsch-align +   │ │  │     (model loaded once,      │\n│  │ weighted average  │ │  │      jobs run sequentially)  │\n│  └──────────────────┘ │  │                              │\n│                       │  │  3. Extract C1' coords       │\n│  IF n &lt; 3 slots:      │  │     from PDB output          │\n│  ┌──────────────────┐ │  │                              │\n│  │ Phase B: Split   │ │  │  ⚠ No chunking — full       │\n│  │ Round-robin      │ │  │    sequence per predict call  │\n│  │ across GPUs      │ │  └──────────────────────────────┘\n│  │ (both use MSA)   │ │\n│  └──────────────────┘ │\n│                       │\n│  Long seq handling:   │\n│  ┌──────────────────┐ │\n│  │ Chunk (512 nt,   │ │\n│  │ 64 overlap)      │ │\n│  │      ↓           │ │\n│  │ Predict chunks   │ │\n│  │      ↓           │ │\n│  │ Kabsch-align     │ │\n│  │ overlaps + blend │ │\n│  │ → reassemble     │ │\n│  └──────────────────┘ │\n└───────────┬───────────┘\n            │\n            ▼\n┌─────────────────────────────────────────────────────────┐\n│           🔗  Build Combined Predictions (5 slots)       │\n│                                                         │\n│   Slot priority:                                        │\n│     1. TBM predictions                                  │\n│     2. Protenix predictions (w/ RNA constraints)        │\n│     3. Boltz2 predictions  (w/ RNA constraints)         │\n│     4. De novo fallback    (idealized A-form helix)     │\n│                                                         │\n│   All predictions pass through adaptive_rna_constraints │\n│   (bond lengths, angles, smoothing, self-avoidance)     │\n└──────────────────────────┬──────────────────────────────┘\n                           │\n                           ▼\n┌─────────────────────────────────────────────────────────┐\n│         PHASE 3: RNAPro Refinement (2× T4 GPUs)         │\n│                                                         │\n│  Skip if sequence &gt; 1,000 nt                            │\n│                                                         │\n│  1. Export all 5 combined predictions as .pt templates  │\n│                                                         │\n│  2. Split targets across GPUs                           │\n│                                                         │\n│  3. Run RNAPro inference in parallel                    │\n│     (RibonanzaNet2 embeddings + precomputed templates)  │\n│                                                         │\n│  4. Replace combined predictions with RNAPro output     │\n│     (keep originals for any failed targets)             │\n└──────────────────────────┬──────────────────────────────┘\n                           │\n                           ▼\n┌─────────────────────────────────────────────────────────┐\n│               PHASE 4: Save Submission                   │\n│                                                         │\n│  📤  submission.csv                                      │\n│  Format: 5 samples × (x, y, z) per residue C1' atom    │\n│  Coordinates clipped to [−999.999, 9999.999]            │\n└─────────────────────────────────────────────────────────┘\n</code></pre>\n<h2>Key Design Decisions</h2>\n<ul>\n<li><strong>5 prediction slots</strong> per target — filled greedily by TBM → Protenix → Boltz2 → de novo</li>\n<li><strong>Length-adaptive strategy</strong> — short sequences favor model diversity (Protenix + Boltz2); long sequences favor template coverage</li>\n<li><strong>MSA/noMSA ensemble</strong> — for targets needing ≥3 Protenix slots, two independent runs are Kabsch-aligned and averaged for better accuracy</li>\n<li><strong>Diversity selection</strong> — TBM candidates are filtered by a greedy algorithm balancing structural diversity (TM-score) against template quality</li>\n<li><strong>Multi-GPU parallelism</strong> — every phase distributes work across both T4 GPUs</li>\n<li><strong>RNAPro as final refinement</strong> — treats all 5 combined predictions as templates and re-predicts, improving structural quality</li>\n</ul>\n<h2>What I wished to do if have more time:</h2>\n<ul>\n<li>Explore embedding based search</li>\n</ul>",
  "messages": [
    {
      "id": 3435037,
      "postDate": "2026-04-03T15:43:39.693Z",
      "content": "<p>First, I would like to thank Kaggle and the host for this interesting challenge. Second, I want to give credits to authors of these incredible public notebooks, on which my solution was built: <a href=\"https://www.kaggle.com/code/lbugnon/boltz2-baseline\" target=\"_blank\">Boltz2_baseline</a>, <a href=\"https://www.kaggle.com/code/llkh0a/stanford-rna-3d-folding-part-2-protenix-tbm\" target=\"_blank\">Protenix+TBM</a>, <a href=\"https://www.kaggle.com/code/jaejohn/rnapro-inference-with-tbm\" target=\"_blank\">RNAPro inference with TBM\n</a>, <a href=\"https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference\" target=\"_blank\">RNAPro Inference</a></p>\n<h2>Templates</h2>\n<p>The representation of RNA that is available is its sequence. That's the only thing I can work with really. So, to find the best templates, I need to first understand what templates are \"good\" in terms of their sequence representation. Thus, I did a templates audit on <code>validation_sequences</code>, trying to find their best templates from <code>train_sequences</code>. Here is an example of what I found: for sequence <code>8ZNQ</code> (30nt) the best templates seems to be <code>1AKX</code>, which have a TM score of 0.27864. However, using biopython sequence aligner, the sequence similarity will be -0.175, which means there is no chance to find this template only using sequence aligner. This is what referred to as short \"template-free\" targets discussed by the host. For longer sequences, aligner seems to work better.</p>\n<p>The conclusion here is clear, given representation of RNA only as sequences, there is a limit to the quality of templates I can find. I did experiment other TBM methods (primary chain + side chain search). Some worked well in my local validation but performs poorly on the LB when combined with DL models so I did not end up using. Because \"hard\" sequences cannot be found by aligner and \"easy\" sequences can almost be found by any aligner, the TBM used in my final submission are public TBM methods with very little modification.</p>\n<table>\n<thead>\n<tr>\n<th>method</th>\n<th>local validation</th>\n<th>Public LB (before rerun)</th>\n<th>Private LB (before rerun)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>primary chain (selected)</td>\n<td>0.33840</td>\n<td>0.32006</td>\n<td>0.46044</td>\n</tr>\n<tr>\n<td>primary chain john ver. (selected)</td>\n<td>0.31873</td>\n<td>0.34932</td>\n<td>0.45131</td>\n</tr>\n<tr>\n<td>primary chain + side chain</td>\n<td>0.40150</td>\n<td>0.32790</td>\n<td>0.45089</td>\n</tr>\n</tbody>\n</table>\n<h2>Overall prediction pipeline</h2>\n<p>With a standard TBM method, my model then pivots to create as much diversity in my predictions as possible. The final model is very similar to other shared solutions (TBM+protenix+boltz2+rnapro refinement). The difference is that for TBM and Protenix generated solutions, I select the ones that are most different from each other, thus maximizing diversity in my 5 predictions. More details can be found in the flowchart below.</p>\n<pre><code>┌─────────────────────────────────────────────────────────┐\n│              📥  Input: Test RNA Sequences               │\n└──────────────────────────┬──────────────────────────────┘\n                           │\n                           ▼\n┌─────────────────────────────────────────────────────────┐\n│                  PHASE 1: Template-Based Modeling (TBM)  │\n│                                                         │\n│  1. Align each test sequence against 5,744 training     │\n│     structures using global pairwise alignment          │\n│                                                         │\n│  2. Filter by similarity &amp; percent identity (≥50%)      │\n│                                                         │\n│  3. Adapt template coordinates → query via alignment    │\n│     (interpolate/extrapolate gaps)                      │\n│                                                         │\n│  4. Apply diversity transforms to generate variants:    │\n│     • Hinge bending at random pivot points              │\n│     • Chain jittering (rotate + translate per chain)    │\n│     • Smooth wiggle (spline-based displacement)         │\n│                                                         │\n│  5. Greedy max-diversity selection                      │\n│     (quality × diversity trade-off)                     │\n│                                                         │\n│  Slot allocation by length:                             │\n│    Short (≤512 nt) → max 2 TBM, min 3 Protenix        │\n│    Long  (&gt;512 nt) → max 5 TBM, 0 min Protenix        │\n└──────────────────────────┬──────────────────────────────┘\n                           │\n                           ▼\n              ┌────────────────────────┐\n              │  Remaining slots =     │\n              │  5 − TBM count         │\n              └─────┬─────────┬────────┘\n                    │         │\n          ┌─────────┘         └──────────┐\n          ▼                              ▼\n┌───────────────────────┐  ┌──────────────────────────────┐\n│  PHASE 2: Protenix    │  │  PHASE 2.5: Boltz2           │\n│  (2× T4 GPUs)         │  │  (2× T4 GPUs, after Protenix)│\n│                       │  │                              │\n│  IF n ≥ 3 slots:      │  │  Eligible: seq &lt; 900 nt     │\n│  ┌──────────────────┐ │  │  &amp; slots still available     │\n│  │ Phase A: Ensemble │ │  │                              │\n│  │ GPU 0 → MSA run  │ │  │  1. Write multi-chain FASTA  │\n│  │ GPU 1 → noMSA run│ │  │                              │\n│  │       ↓          │ │  │  2. Persistent GPU workers   │\n│  │ Kabsch-align +   │ │  │     (model loaded once,      │\n│  │ weighted average  │ │  │      jobs run sequentially)  │\n│  └──────────────────┘ │  │                              │\n│                       │  │  3. Extract C1' coords       │\n│  IF n &lt; 3 slots:      │  │     from PDB output          │\n│  ┌──────────────────┐ │  │                              │\n│  │ Phase B: Split   │ │  │  ⚠ No chunking — full       │\n│  │ Round-robin      │ │  │    sequence per predict call  │\n│  │ across GPUs      │ │  └──────────────────────────────┘\n│  │ (both use MSA)   │ │\n│  └──────────────────┘ │\n│                       │\n│  Long seq handling:   │\n│  ┌──────────────────┐ │\n│  │ Chunk (512 nt,   │ │\n│  │ 64 overlap)      │ │\n│  │      ↓           │ │\n│  │ Predict chunks   │ │\n│  │      ↓           │ │\n│  │ Kabsch-align     │ │\n│  │ overlaps + blend │ │\n│  │ → reassemble     │ │\n│  └──────────────────┘ │\n└───────────┬───────────┘\n            │\n            ▼\n┌─────────────────────────────────────────────────────────┐\n│           🔗  Build Combined Predictions (5 slots)       │\n│                                                         │\n│   Slot priority:                                        │\n│     1. TBM predictions                                  │\n│     2. Protenix predictions (w/ RNA constraints)        │\n│     3. Boltz2 predictions  (w/ RNA constraints)         │\n│     4. De novo fallback    (idealized A-form helix)     │\n│                                                         │\n│   All predictions pass through adaptive_rna_constraints │\n│   (bond lengths, angles, smoothing, self-avoidance)     │\n└──────────────────────────┬──────────────────────────────┘\n                           │\n                           ▼\n┌─────────────────────────────────────────────────────────┐\n│         PHASE 3: RNAPro Refinement (2× T4 GPUs)         │\n│                                                         │\n│  Skip if sequence &gt; 1,000 nt                            │\n│                                                         │\n│  1. Export all 5 combined predictions as .pt templates  │\n│                                                         │\n│  2. Split targets across GPUs                           │\n│                                                         │\n│  3. Run RNAPro inference in parallel                    │\n│     (RibonanzaNet2 embeddings + precomputed templates)  │\n│                                                         │\n│  4. Replace combined predictions with RNAPro output     │\n│     (keep originals for any failed targets)             │\n└──────────────────────────┬──────────────────────────────┘\n                           │\n                           ▼\n┌─────────────────────────────────────────────────────────┐\n│               PHASE 4: Save Submission                   │\n│                                                         │\n│  📤  submission.csv                                      │\n│  Format: 5 samples × (x, y, z) per residue C1' atom    │\n│  Coordinates clipped to [−999.999, 9999.999]            │\n└─────────────────────────────────────────────────────────┘\n</code></pre>\n<h2>Key Design Decisions</h2>\n<ul>\n<li><strong>5 prediction slots</strong> per target — filled greedily by TBM → Protenix → Boltz2 → de novo</li>\n<li><strong>Length-adaptive strategy</strong> — short sequences favor model diversity (Protenix + Boltz2); long sequences favor template coverage</li>\n<li><strong>MSA/noMSA ensemble</strong> — for targets needing ≥3 Protenix slots, two independent runs are Kabsch-aligned and averaged for better accuracy</li>\n<li><strong>Diversity selection</strong> — TBM candidates are filtered by a greedy algorithm balancing structural diversity (TM-score) against template quality</li>\n<li><strong>Multi-GPU parallelism</strong> — every phase distributes work across both T4 GPUs</li>\n<li><strong>RNAPro as final refinement</strong> — treats all 5 combined predictions as templates and re-predicts, improving structural quality</li>\n</ul>\n<h2>What I wished to do if have more time:</h2>\n<ul>\n<li>Explore embedding based search</li>\n</ul>",
      "rawMarkdown": "First, I would like to thank Kaggle and the host for this interesting challenge. Second, I want to give credits to authors of these incredible public notebooks, on which my solution was built: [Boltz2_baseline](https://www.kaggle.com/code/lbugnon/boltz2-baseline), [Protenix+TBM](https://www.kaggle.com/code/llkh0a/stanford-rna-3d-folding-part-2-protenix-tbm), [RNAPro inference with TBM\n](https://www.kaggle.com/code/jaejohn/rnapro-inference-with-tbm), [RNAPro Inference](https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference)\n\n## Templates\nThe representation of RNA that is available is its sequence. That's the only thing I can work with really. So, to find the best templates, I need to first understand what templates are \"good\" in terms of their sequence representation. Thus, I did a templates audit on `validation_sequences`, trying to find their best templates from `train_sequences`. Here is an example of what I found: for sequence `8ZNQ` (30nt) the best templates seems to be `1AKX`, which have a TM score of 0.27864. However, using biopython sequence aligner, the sequence similarity will be -0.175, which means there is no chance to find this template only using sequence aligner. This is what referred to as short \"template-free\" targets discussed by the host. For longer sequences, aligner seems to work better.\n\nThe conclusion here is clear, given representation of RNA only as sequences, there is a limit to the quality of templates I can find. I did experiment other TBM methods (primary chain + side chain search). Some worked well in my local validation but performs poorly on the LB when combined with DL models so I did not end up using. Because \"hard\" sequences cannot be found by aligner and \"easy\" sequences can almost be found by any aligner, the TBM used in my final submission are public TBM methods with very little modification.\n\n| method | local validation | Public LB (before rerun)  | Private LB (before rerun)\n| --- | --- | --- | --- |\n| primary chain (selected) | 0.33840 | 0.32006 | 0.46044 |\n| primary chain john ver. (selected) | 0.31873 | 0.34932 | 0.45131 |\n| primary chain + side chain | 0.40150 | 0.32790 | 0.45089 |\n\n## Overall prediction pipeline\nWith a standard TBM method, my model then pivots to create as much diversity in my predictions as possible. The final model is very similar to other shared solutions (TBM+protenix+boltz2+rnapro refinement). The difference is that for TBM and Protenix generated solutions, I select the ones that are most different from each other, thus maximizing diversity in my 5 predictions. More details can be found in the flowchart below.\n\n\n```\n┌─────────────────────────────────────────────────────────┐\n│              📥  Input: Test RNA Sequences               │\n└──────────────────────────┬──────────────────────────────┘\n                           │\n                           ▼\n┌─────────────────────────────────────────────────────────┐\n│                  PHASE 1: Template-Based Modeling (TBM)  │\n│                                                         │\n│  1. Align each test sequence against 5,744 training     │\n│     structures using global pairwise alignment          │\n│                                                         │\n│  2. Filter by similarity & percent identity (≥50%)      │\n│                                                         │\n│  3. Adapt template coordinates → query via alignment    │\n│     (interpolate/extrapolate gaps)                      │\n│                                                         │\n│  4. Apply diversity transforms to generate variants:    │\n│     • Hinge bending at random pivot points              │\n│     • Chain jittering (rotate + translate per chain)    │\n│     • Smooth wiggle (spline-based displacement)         │\n│                                                         │\n│  5. Greedy max-diversity selection                      │\n│     (quality × diversity trade-off)                     │\n│                                                         │\n│  Slot allocation by length:                             │\n│    Short (≤512 nt) → max 2 TBM, min 3 Protenix        │\n│    Long  (>512 nt) → max 5 TBM, 0 min Protenix        │\n└──────────────────────────┬──────────────────────────────┘\n                           │\n                           ▼\n              ┌────────────────────────┐\n              │  Remaining slots =     │\n              │  5 − TBM count         │\n              └─────┬─────────┬────────┘\n                    │         │\n          ┌─────────┘         └──────────┐\n          ▼                              ▼\n┌───────────────────────┐  ┌──────────────────────────────┐\n│  PHASE 2: Protenix    │  │  PHASE 2.5: Boltz2           │\n│  (2× T4 GPUs)         │  │  (2× T4 GPUs, after Protenix)│\n│                       │  │                              │\n│  IF n ≥ 3 slots:      │  │  Eligible: seq < 900 nt     │\n│  ┌──────────────────┐ │  │  & slots still available     │\n│  │ Phase A: Ensemble │ │  │                              │\n│  │ GPU 0 → MSA run  │ │  │  1. Write multi-chain FASTA  │\n│  │ GPU 1 → noMSA run│ │  │                              │\n│  │       ↓          │ │  │  2. Persistent GPU workers   │\n│  │ Kabsch-align +   │ │  │     (model loaded once,      │\n│  │ weighted average  │ │  │      jobs run sequentially)  │\n│  └──────────────────┘ │  │                              │\n│                       │  │  3. Extract C1' coords       │\n│  IF n < 3 slots:      │  │     from PDB output          │\n│  ┌──────────────────┐ │  │                              │\n│  │ Phase B: Split   │ │  │  ⚠ No chunking — full       │\n│  │ Round-robin      │ │  │    sequence per predict call  │\n│  │ across GPUs      │ │  └──────────────────────────────┘\n│  │ (both use MSA)   │ │\n│  └──────────────────┘ │\n│                       │\n│  Long seq handling:   │\n│  ┌──────────────────┐ │\n│  │ Chunk (512 nt,   │ │\n│  │ 64 overlap)      │ │\n│  │      ↓           │ │\n│  │ Predict chunks   │ │\n│  │      ↓           │ │\n│  │ Kabsch-align     │ │\n│  │ overlaps + blend │ │\n│  │ → reassemble     │ │\n│  └──────────────────┘ │\n└───────────┬───────────┘\n            │\n            ▼\n┌─────────────────────────────────────────────────────────┐\n│           🔗  Build Combined Predictions (5 slots)       │\n│                                                         │\n│   Slot priority:                                        │\n│     1. TBM predictions                                  │\n│     2. Protenix predictions (w/ RNA constraints)        │\n│     3. Boltz2 predictions  (w/ RNA constraints)         │\n│     4. De novo fallback    (idealized A-form helix)     │\n│                                                         │\n│   All predictions pass through adaptive_rna_constraints │\n│   (bond lengths, angles, smoothing, self-avoidance)     │\n└──────────────────────────┬──────────────────────────────┘\n                           │\n                           ▼\n┌─────────────────────────────────────────────────────────┐\n│         PHASE 3: RNAPro Refinement (2× T4 GPUs)         │\n│                                                         │\n│  Skip if sequence > 1,000 nt                            │\n│                                                         │\n│  1. Export all 5 combined predictions as .pt templates  │\n│                                                         │\n│  2. Split targets across GPUs                           │\n│                                                         │\n│  3. Run RNAPro inference in parallel                    │\n│     (RibonanzaNet2 embeddings + precomputed templates)  │\n│                                                         │\n│  4. Replace combined predictions with RNAPro output     │\n│     (keep originals for any failed targets)             │\n└──────────────────────────┬──────────────────────────────┘\n                           │\n                           ▼\n┌─────────────────────────────────────────────────────────┐\n│               PHASE 4: Save Submission                   │\n│                                                         │\n│  📤  submission.csv                                      │\n│  Format: 5 samples × (x, y, z) per residue C1' atom    │\n│  Coordinates clipped to [−999.999, 9999.999]            │\n└─────────────────────────────────────────────────────────┘\n```\n\n## Key Design Decisions\n\n- **5 prediction slots** per target — filled greedily by TBM → Protenix → Boltz2 → de novo\n- **Length-adaptive strategy** — short sequences favor model diversity (Protenix + Boltz2); long sequences favor template coverage\n- **MSA/noMSA ensemble** — for targets needing ≥3 Protenix slots, two independent runs are Kabsch-aligned and averaged for better accuracy\n- **Diversity selection** — TBM candidates are filtered by a greedy algorithm balancing structural diversity (TM-score) against template quality\n- **Multi-GPU parallelism** — every phase distributes work across both T4 GPUs\n- **RNAPro as final refinement** — treats all 5 combined predictions as templates and re-predicts, improving structural quality\n\n## What I wished to do if have more time:\n- Explore embedding based search",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3435037": "First, I would like to thank Kaggle and the host for this interesting challenge. Second, I want to give credits to authors of these incredible public notebooks, on which my solution was built: [Boltz2_baseline](https://www.kaggle.com/code/lbugnon/boltz2-baseline), [Protenix+TBM](https://www.kaggle.com/code/llkh0a/stanford-rna-3d-folding-part-2-protenix-tbm), [RNAPro inference with TBM\n](https://www.kaggle.com/code/jaejohn/rnapro-inference-with-tbm), [RNAPro Inference](https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference)\n\n## Templates\nThe representation of RNA that is available is its sequence. That's the only thing I can work with really. So, to find the best templates, I need to first understand what templates are \"good\" in terms of their sequence representation. Thus, I did a templates audit on `validation_sequences`, trying to find their best templates from `train_sequences`. Here is an example of what I found: for sequence `8ZNQ` (30nt) the best templates seems to be `1AKX`, which have a TM score of 0.27864. However, using biopython sequence aligner, the sequence similarity will be -0.175, which means there is no chance to find this template only using sequence aligner. This is what referred to as short \"template-free\" targets discussed by the host. For longer sequences, aligner seems to work better.\n\nThe conclusion here is clear, given representation of RNA only as sequences, there is a limit to the quality of templates I can find. I did experiment other TBM methods (primary chain + side chain search). Some worked well in my local validation but performs poorly on the LB when combined with DL models so I did not end up using. Because \"hard\" sequences cannot be found by aligner and \"easy\" sequences can almost be found by any aligner, the TBM used in my final submission are public TBM methods with very little modification.\n\n| method | local validation | Public LB (before rerun)  | Private LB (before rerun)\n| --- | --- | --- | --- |\n| primary chain (selected) | 0.33840 | 0.32006 | 0.46044 |\n| primary chain john ver. (selected) | 0.31873 | 0.34932 | 0.45131 |\n| primary chain + side chain | 0.40150 | 0.32790 | 0.45089 |\n\n## Overall prediction pipeline\nWith a standard TBM method, my model then pivots to create as much diversity in my predictions as possible. The final model is very similar to other shared solutions (TBM+protenix+boltz2+rnapro refinement). The difference is that for TBM and Protenix generated solutions, I select the ones that are most different from each other, thus maximizing diversity in my 5 predictions. More details can be found in the flowchart below.\n\n\n```\n┌─────────────────────────────────────────────────────────┐\n│              📥  Input: Test RNA Sequences               │\n└──────────────────────────┬──────────────────────────────┘\n                           │\n                           ▼\n┌─────────────────────────────────────────────────────────┐\n│                  PHASE 1: Template-Based Modeling (TBM)  │\n│                                                         │\n│  1. Align each test sequence against 5,744 training     │\n│     structures using global pairwise alignment          │\n│                                                         │\n│  2. Filter by similarity & percent identity (≥50%)      │\n│                                                         │\n│  3. Adapt template coordinates → query via alignment    │\n│     (interpolate/extrapolate gaps)                      │\n│                                                         │\n│  4. Apply diversity transforms to generate variants:    │\n│     • Hinge bending at random pivot points              │\n│     • Chain jittering (rotate + translate per chain)    │\n│     • Smooth wiggle (spline-based displacement)         │\n│                                                         │\n│  5. Greedy max-diversity selection                      │\n│     (quality × diversity trade-off)                     │\n│                                                         │\n│  Slot allocation by length:                             │\n│    Short (≤512 nt) → max 2 TBM, min 3 Protenix        │\n│    Long  (>512 nt) → max 5 TBM, 0 min Protenix        │\n└──────────────────────────┬──────────────────────────────┘\n                           │\n                           ▼\n              ┌────────────────────────┐\n              │  Remaining slots =     │\n              │  5 − TBM count         │\n              └─────┬─────────┬────────┘\n                    │         │\n          ┌─────────┘         └──────────┐\n          ▼                              ▼\n┌───────────────────────┐  ┌──────────────────────────────┐\n│  PHASE 2: Protenix    │  │  PHASE 2.5: Boltz2           │\n│  (2× T4 GPUs)         │  │  (2× T4 GPUs, after Protenix)│\n│                       │  │                              │\n│  IF n ≥ 3 slots:      │  │  Eligible: seq < 900 nt     │\n│  ┌──────────────────┐ │  │  & slots still available     │\n│  │ Phase A: Ensemble │ │  │                              │\n│  │ GPU 0 → MSA run  │ │  │  1. Write multi-chain FASTA  │\n│  │ GPU 1 → noMSA run│ │  │                              │\n│  │       ↓          │ │  │  2. Persistent GPU workers   │\n│  │ Kabsch-align +   │ │  │     (model loaded once,      │\n│  │ weighted average  │ │  │      jobs run sequentially)  │\n│  └──────────────────┘ │  │                              │\n│                       │  │  3. Extract C1' coords       │\n│  IF n < 3 slots:      │  │     from PDB output          │\n│  ┌──────────────────┐ │  │                              │\n│  │ Phase B: Split   │ │  │  ⚠ No chunking — full       │\n│  │ Round-robin      │ │  │    sequence per predict call  │\n│  │ across GPUs      │ │  └──────────────────────────────┘\n│  │ (both use MSA)   │ │\n│  └──────────────────┘ │\n│                       │\n│  Long seq handling:   │\n│  ┌──────────────────┐ │\n│  │ Chunk (512 nt,   │ │\n│  │ 64 overlap)      │ │\n│  │      ↓           │ │\n│  │ Predict chunks   │ │\n│  │      ↓           │ │\n│  │ Kabsch-align     │ │\n│  │ overlaps + blend │ │\n│  │ → reassemble     │ │\n│  └──────────────────┘ │\n└───────────┬───────────┘\n            │\n            ▼\n┌─────────────────────────────────────────────────────────┐\n│           🔗  Build Combined Predictions (5 slots)       │\n│                                                         │\n│   Slot priority:                                        │\n│     1. TBM predictions                                  │\n│     2. Protenix predictions (w/ RNA constraints)        │\n│     3. Boltz2 predictions  (w/ RNA constraints)         │\n│     4. De novo fallback    (idealized A-form helix)     │\n│                                                         │\n│   All predictions pass through adaptive_rna_constraints │\n│   (bond lengths, angles, smoothing, self-avoidance)     │\n└──────────────────────────┬──────────────────────────────┘\n                           │\n                           ▼\n┌─────────────────────────────────────────────────────────┐\n│         PHASE 3: RNAPro Refinement (2× T4 GPUs)         │\n│                                                         │\n│  Skip if sequence > 1,000 nt                            │\n│                                                         │\n│  1. Export all 5 combined predictions as .pt templates  │\n│                                                         │\n│  2. Split targets across GPUs                           │\n│                                                         │\n│  3. Run RNAPro inference in parallel                    │\n│     (RibonanzaNet2 embeddings + precomputed templates)  │\n│                                                         │\n│  4. Replace combined predictions with RNAPro output     │\n│     (keep originals for any failed targets)             │\n└──────────────────────────┬──────────────────────────────┘\n                           │\n                           ▼\n┌─────────────────────────────────────────────────────────┐\n│               PHASE 4: Save Submission                   │\n│                                                         │\n│  📤  submission.csv                                      │\n│  Format: 5 samples × (x, y, z) per residue C1' atom    │\n│  Coordinates clipped to [−999.999, 9999.999]            │\n└─────────────────────────────────────────────────────────┘\n```\n\n## Key Design Decisions\n\n- **5 prediction slots** per target — filled greedily by TBM → Protenix → Boltz2 → de novo\n- **Length-adaptive strategy** — short sequences favor model diversity (Protenix + Boltz2); long sequences favor template coverage\n- **MSA/noMSA ensemble** — for targets needing ≥3 Protenix slots, two independent runs are Kabsch-aligned and averaged for better accuracy\n- **Diversity selection** — TBM candidates are filtered by a greedy algorithm balancing structural diversity (TM-score) against template quality\n- **Multi-GPU parallelism** — every phase distributes work across both T4 GPUs\n- **RNAPro as final refinement** — treats all 5 combined predictions as templates and re-predicts, improving structural quality\n\n## What I wished to do if have more time:\n- Explore embedding based search"
  }
}