{
  "id": 692863,
  "title": "4th Place Solution — Cross-Attention Template Reranker + Protenix + RNAPro",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/692863",
  "author_name": "aleksandr3312",
  "post_date": "2026-04-18T09:28:32.398000",
  "votes": 9,
  "comment_count": 8,
  "views": 0,
  "content": "<h1>4th Place Solution — Cross-Attention Template Reranker + Protenix + RNAPro</h1>\n<h2>Summary</h2>\n<p>This was my first bioinformatics competition, and more generally my first serious Kaggle competition overall. I was curious to see how far I could go without domain expertise, and with AI tools as my main leverage. Big thanks to the organizers for the huge amount of work they put into Part 2, to the community for the excellent public datasets and kernels, and congratulations to everyone on the prize list!</p>\n<p>The pipeline is <strong>Protenix + RNAPro + template search reranked by a Cross-Attention TM predictor</strong>. The core idea everything else hangs off is a <strong>CUDA port of US-align</strong> that made it cheap to label TM-scores for a very large number of template–template pairs (~1.42M). That labeled set is what the Cross-Attention model is trained on, which turns \"take top-N by sequence identity\" into a ranking that reorders candidates by <strong>predicted structural TM</strong>.</p>\n<h2>Slot allocation (5 predictions per target)</h2>\n<table>\n<thead>\n<tr>\n<th>Slot</th>\n<th>Source</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>TBM — <strong>Bio top-1</strong> (<code>Bio.Align.PairwiseAligner</code> in global mode, competition-tuned scores)</td>\n</tr>\n<tr>\n<td>2</td>\n<td>TBM — <strong>Model top-1</strong> (Cross-Attn predicted TM, mmseqs-cluster-diverse from slot 1)</td>\n</tr>\n<tr>\n<td>3</td>\n<td>TBM — <strong>Combined top-1</strong> (<code>sim * model_score</code>, cluster-diverse from slots 1-2)</td>\n</tr>\n<tr>\n<td>4</td>\n<td><strong>Protenix v1</strong> (co-fold with protein/DNA/ligands if they fit)</td>\n</tr>\n<tr>\n<td>5</td>\n<td><strong>RNAPro</strong></td>\n</tr>\n</tbody>\n</table>\n<p>For targets longer than 512 nt the Cross-Attn model isn't run (it was trained up to 512) and slots 2-3 fall back to Bio ranks 2 and 3 with cluster diversity.</p>\n<h2>GPU US-align (the labeling engine)</h2>\n<p><a href=\"https://www.kaggle.com/code/gapchenko/gpu-accelerated-usalign/notebook\" target=\"_blank\">https://www.kaggle.com/code/gapchenko/gpu-accelerated-usalign/notebook</a></p>\n<p>I re-implemented <code>TMalign_main</code> from US-align in CUDA — faithful to the CPU source. All five phases (initial alignment, SS alignment, local fragment superposition, SS + distance matrix, fragment gapless threading) run on GPU; <code>DP_iter_gpu</code>, <code>gpu_TMscore8_search</code>, and <code>get_initial5_gpu_fast</code> are the heavy kernels. A binary stdin batch mode takes coordinate pairs directly via <code>struct.pack</code> — no PDB I/O. A separate PDB-batch mode uses a CPU thread pool with one thread per CUDA stream, so the GPU stays saturated while files are parsed.</p>\n<p>On P100-class hardware this gets to hundreds of milliseconds per RNA pair, vs. seconds on the CPU binary — with the biggest speedup on longer molecules (&gt;2k nt). That turned a \"days on CPU\" labeling problem into a \"hours on one GPU\" one, and is the prerequisite for the Cross-Attention predictor below.</p>\n<p><strong>What got labeled</strong>: every pair of available templates in the DB with length ratio in [0.75, 1.25] (~1.42M pairs). These pairs together with their C1' ground-truth coordinates are the supervised signal — TM-score is what I want the network to predict.</p>\n<h2>Cross-Attention TM predictor</h2>\n<p>The model takes <strong>two nucleotide sequences</strong> (query and template) as independent inputs. Each passes through a shared encoder (embed -&gt; 5 self-attention layers); the two encoded sequences then interact through one bidirectional cross-attention block. The query is pooled, fused with a small context vector, and mapped to a single logit for <code>P(TM &gt;= 0.5)</code>.</p>\n<pre><code>        query seq                  template seq\n      (N &lt;= 512 nt)              (M &lt;= 512 nt)\n           |                           |\n           |   shared encoder applied  |\n           |   to each independently   |\n           v                           v\n    embed(vocab=5, d=128) + learned pos(max_len=512, d=128)\n           |                           |\n           v                           v\n    5 x [ self-attention (h=4, d=128) + FFN ]\n           |                           |\n           v                           v\n       q_enc (N, 128)           t_enc (M, 128)\n           \\                         /\n            \\                       /\n             v                     v\n      +---------------------------------+\n      |   bidirectional cross-attention |\n      |     q &lt;- q + MHA(Q=q, K=t, V=t) |   # q attends to t\n      |     t &lt;- t + MHA(Q=t, K=q, V=q) |   # t attends to q\n      +-----------------+---------------+\n                        |   (t is then discarded)\n                        v\n              FFN + LayerNorm on q\n                        |\n                        v\n            masked-mean pool over N\n                        |\n                        v\n                pooled in R^128\n                        |\n                        |     context vector (10d):\n                        |       4 common feats:\n                        |         log(q_len)/6, log(t_len)/6, q_len/(t_len+1), Levenshtein similarity\n                        |       6 auxiliary legacy feats (filled with zeros):\n                        |         \n                        |                       | \n                        |                       |\n                        |                Linear(10 -&gt; 128)\n                        |                       |\n                        +------- concat &lt;-------+\n                                 (256d)\n                                   |\n                                   v\n                        Linear(256 -&gt; 128) -&gt; GELU\n                                   |\n                                   v\n                           Linear(128 -&gt; 1)\n                                   |\n                                   v\n                                sigmoid\n                                   |\n                                   v\n                             P(TM &gt;= 0.5)\n</code></pre>\n<p>d = 128, 4 heads, 5 self-attention layers, 1 bidirectional cross-attention layer. Embedding is plain learned (vocab of 5 for <code>AUGC</code> + pad), positions are learned up to 512.</p>\n<p><strong>Training.</strong></p>\n<ul>\n<li>Labels: 1.42M template-template pairs from GPU US-align, binarized at <code>TM &gt;= 0.5</code>.</li>\n<li>Split: <code>mmseqs_0.300</code> clusters — 15% of clusters held out for validation; train pairs must cross clusters, val pairs must involve at least one val-cluster target.</li>\n<li>Filters: length ratio in [0.75, 1.25], TM in [0.12, 0.92] (drop trivially similar and trivially different), negatives downsampled to 1/3 of positives.</li>\n<li>Loss: <code>BCEWithLogitsLoss(pos_weight)</code> to compensate for residual class imbalance.</li>\n<li>Optimizer: AdamW (lr = 3e-4, wd = 1e-4), cosine schedule, early stop on val Spearman.</li>\n<li>This config (d=128, heads=4, layers=5) was the winner of a small architecture sweep over depth (1-8), width (24-192), and heads (2-8).</li>\n</ul>\n<p><strong>Inference.</strong> For each test target take top-500 templates by sequence similarity, score each one with\n      the CA predictor in batches of 64, sigmoid the logits — that's the <code>model_score</code>. Slot 2 picks <code>argma\n      x(model_score)</code> in a fresh <code>mmseqs_0.300</code> cluster; slot 3 picks <code>argmax(sim × model_score)</code> in yet ano\n      ther fresh cluster. The three TBM slots thus cover three different reasons to like a template: high sequence identity, high predicted TM, and both agreeing at once.</p>\n<h2>Protenix &amp; RNAPro integration</h2>\n<p>Both run in-memory — model loaded once, configs updated per target, no subprocess restart. One caveat up front: <strong>according to my comparisons, the marginal benefit of explicit MSA usage turned out to be minimal for both Protenix and RNAPro on this eval setup</strong> — so a lot of the MSA wiring below is really about <em>not breaking</em> the models rather than squeezing more out of MSA.</p>\n<p>On top of the public inference recipes:</p>\n<ul>\n<li><strong>Co-fold.</strong> Non-RNA chains go in if they fit the token budget; if they miss by ≤100 tokens the largest protein ≥500 aa is trimmed symmetrically on both ends (and its MSA is column-trimmed to match).</li>\n<li><strong>Ligand filtering.</strong> Buffer and crystallization artifacts (<code>GOL</code>, <code>EDO</code>, polyethylene glycols <code>PEG</code>/<code>PG4</code>/<code>P6G</code>/<code>1PE</code>/<code>PE4</code>/<code>EPE</code>, <code>SO4</code>, <code>PO4</code>, salts and buffers <code>CIT</code>/<code>ACT</code>/<code>FMT</code>/<code>ACY</code>/<code>TRS</code>/<code>MES</code>, solvents <code>MPD</code>/<code>IPA</code>/<code>DMS</code>/<code>BME</code>, polyamines <code>SPM</code>/<code>SPD</code>, crystallographic heavy-atom probes <code>NCO</code>/<code>IRI</code>/<code>RHD</code>, <code>HEZ</code>) are blacklisted. Everything else is injected as a <code>CCD_*</code> entity.</li>\n<li><strong>Heterodimer handling.</strong> Multi-chain RNA targets use a separate <code>rnaSequence</code> entity per distinct chain, with the competition MSA column-sliced per chain.</li>\n<li><strong>Long sequences.</strong> Overlapping chunks, per-chunk MSA column slicing, Kabsch alignment on the full overlap, and a narrow ±6-residue linear blend around each handover midpoint (hard switch outside that window — averaging over long overlaps hurt on my val).</li>\n<li><strong>Time budgets.</strong> Both phases have wall-clock budgets with adaptive per-target limits so that a single slow target can't starve the rest.</li>\n</ul>\n<h2>Findings which worked</h2>\n<ul>\n<li><strong>GPU US-align → CA predictor.</strong> The main lever. Without cheap TM labeling there's no supervised signal to rank templates <em>structurally</em>. Without the rerank, slot 2 is \"take rank-2 by sequence\", which is noisy.</li>\n<li><strong>Diverse TBM slot filling (bio / model / combined). I believe this was the biggest single source of Bo5 gain.</strong> Strictly better than any single ranking on my held-out targets — the bio rank anchors, the model rank finds structurally similar but sequence-diverse templates, the combined rank catches the cases where both agree.</li>\n<li><strong>mmseqs-cluster diversity inside TBM.</strong> Requiring the three TBM slots to come from different <code>mmseqs_0.300</code> clusters was bigger than any scoring tweak — it eliminates near-duplicate templates that all fail the same way.</li>\n<li><strong>Keeping PTX and RNAPro as independent single-slot de-novo predictions</strong> beat letting either model take 2 slots on any consistent subset of targets in my evals. It would probably be more efficient with a conditional slot allocator, but I didn't have enough time to tune such a selector to the point where it consistently earned its keep.</li>\n</ul>\n<h2>What didn't work (or helped only a little)</h2>\n<ul>\n<li><strong>Coordinate-space blending.</strong> Whether uniform, learned, or anchor-based, averaging Kabsch-aligned templates moved scores a little on close templates and hurt on diverse ones. Pick-one-template was simpler and better.</li>\n<li><strong>Explicit secondary-structure features.</strong> I tried folding SS in via common tools like ViennaRNA — both as extra inputs to the Cross-Attention selector and as constraints when stitching chunk-wise de-novo predictions. Neither meaningfully moved the score. Entirely possible this is a skill issue on my end: I'm not confident I was using SS in the right way.</li>\n<li><strong>Better chunk-to-chunk blending.</strong> Tried various stitching schemes — linear vs cubic-spline blends, wider overlap windows, even an L-BFGS fit over the overlap region — but the underlying de-novo chunks were probably just not accurate enough for this to show up in the end score.</li>\n<li><strong>Metadata-based routing of PTX vs. RNAPro slots.</strong> Code computes <code>ptx_heavy</code> / <code>rnapro_heavy</code> / <code>balanced</code> labels from length, protein/DNA presence, and keywords, but the final allocation is flat (1 PTX + 1 RNAPro for everyone). The routing labels are reported for statistics only. Attempts to actually route to 2-of-one model lost consistency on the val set.</li>\n</ul>\n<h2>Attachments</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/datasets/gapchenko/rna3d2-pairwise-tm-labels\" target=\"_blank\">https://www.kaggle.com/datasets/gapchenko/rna3d2-pairwise-tm-labels</a> — 1.42M ground-truth TM-scores for pairs of RNA\nstructures from the competition training set</li>\n<li><a href=\"https://www.kaggle.com/code/gapchenko/gpu-accelerated-usalign/notebook\" target=\"_blank\">https://www.kaggle.com/code/gapchenko/gpu-accelerated-usalign/notebook</a> —CUDA port of TMalign_main from\nUS-align. Includes performance and accuracy comparison against CPU original version.</li>\n<li>train_tm_cross_attn.py — training script for the Cross-Attention TM predictor. Reference\nimplementation showing the full pipeline end-to-end (data loading, feature construction, model, training loop, eval by\nSpearman).</li>\n</ul>\n<h2>Acknowledgements</h2>\n<ul>\n<li>Organizers for running Part 2 so soon after Part 1, and for the clean data release.</li>\n<li><code>@theoviel</code> and <code>@jaejohn</code> for the baseline TBM and RNAPro notebooks that everyone built on.</li>\n<li>Protenix (ByteDance) and RNAPro (NVIDIA Digital Bio) for the open checkpoints.</li>\n<li>Claude (Anthropic) was my pair-programmer for most of this, especially for the CUDA port.</li>\n</ul>",
  "messages": [
    {
      "id": 3444714,
      "postDate": "2026-04-18T09:28:32.400Z",
      "content": "<h1>4th Place Solution — Cross-Attention Template Reranker + Protenix + RNAPro</h1>\n<h2>Summary</h2>\n<p>This was my first bioinformatics competition, and more generally my first serious Kaggle competition overall. I was curious to see how far I could go without domain expertise, and with AI tools as my main leverage. Big thanks to the organizers for the huge amount of work they put into Part 2, to the community for the excellent public datasets and kernels, and congratulations to everyone on the prize list!</p>\n<p>The pipeline is <strong>Protenix + RNAPro + template search reranked by a Cross-Attention TM predictor</strong>. The core idea everything else hangs off is a <strong>CUDA port of US-align</strong> that made it cheap to label TM-scores for a very large number of template–template pairs (~1.42M). That labeled set is what the Cross-Attention model is trained on, which turns \"take top-N by sequence identity\" into a ranking that reorders candidates by <strong>predicted structural TM</strong>.</p>\n<h2>Slot allocation (5 predictions per target)</h2>\n<table>\n<thead>\n<tr>\n<th>Slot</th>\n<th>Source</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>TBM — <strong>Bio top-1</strong> (<code>Bio.Align.PairwiseAligner</code> in global mode, competition-tuned scores)</td>\n</tr>\n<tr>\n<td>2</td>\n<td>TBM — <strong>Model top-1</strong> (Cross-Attn predicted TM, mmseqs-cluster-diverse from slot 1)</td>\n</tr>\n<tr>\n<td>3</td>\n<td>TBM — <strong>Combined top-1</strong> (<code>sim * model_score</code>, cluster-diverse from slots 1-2)</td>\n</tr>\n<tr>\n<td>4</td>\n<td><strong>Protenix v1</strong> (co-fold with protein/DNA/ligands if they fit)</td>\n</tr>\n<tr>\n<td>5</td>\n<td><strong>RNAPro</strong></td>\n</tr>\n</tbody>\n</table>\n<p>For targets longer than 512 nt the Cross-Attn model isn't run (it was trained up to 512) and slots 2-3 fall back to Bio ranks 2 and 3 with cluster diversity.</p>\n<h2>GPU US-align (the labeling engine)</h2>\n<p><a href=\"https://www.kaggle.com/code/gapchenko/gpu-accelerated-usalign/notebook\" target=\"_blank\">https://www.kaggle.com/code/gapchenko/gpu-accelerated-usalign/notebook</a></p>\n<p>I re-implemented <code>TMalign_main</code> from US-align in CUDA — faithful to the CPU source. All five phases (initial alignment, SS alignment, local fragment superposition, SS + distance matrix, fragment gapless threading) run on GPU; <code>DP_iter_gpu</code>, <code>gpu_TMscore8_search</code>, and <code>get_initial5_gpu_fast</code> are the heavy kernels. A binary stdin batch mode takes coordinate pairs directly via <code>struct.pack</code> — no PDB I/O. A separate PDB-batch mode uses a CPU thread pool with one thread per CUDA stream, so the GPU stays saturated while files are parsed.</p>\n<p>On P100-class hardware this gets to hundreds of milliseconds per RNA pair, vs. seconds on the CPU binary — with the biggest speedup on longer molecules (&gt;2k nt). That turned a \"days on CPU\" labeling problem into a \"hours on one GPU\" one, and is the prerequisite for the Cross-Attention predictor below.</p>\n<p><strong>What got labeled</strong>: every pair of available templates in the DB with length ratio in [0.75, 1.25] (~1.42M pairs). These pairs together with their C1' ground-truth coordinates are the supervised signal — TM-score is what I want the network to predict.</p>\n<h2>Cross-Attention TM predictor</h2>\n<p>The model takes <strong>two nucleotide sequences</strong> (query and template) as independent inputs. Each passes through a shared encoder (embed -&gt; 5 self-attention layers); the two encoded sequences then interact through one bidirectional cross-attention block. The query is pooled, fused with a small context vector, and mapped to a single logit for <code>P(TM &gt;= 0.5)</code>.</p>\n<pre><code>        query seq                  template seq\n      (N &lt;= 512 nt)              (M &lt;= 512 nt)\n           |                           |\n           |   shared encoder applied  |\n           |   to each independently   |\n           v                           v\n    embed(vocab=5, d=128) + learned pos(max_len=512, d=128)\n           |                           |\n           v                           v\n    5 x [ self-attention (h=4, d=128) + FFN ]\n           |                           |\n           v                           v\n       q_enc (N, 128)           t_enc (M, 128)\n           \\                         /\n            \\                       /\n             v                     v\n      +---------------------------------+\n      |   bidirectional cross-attention |\n      |     q &lt;- q + MHA(Q=q, K=t, V=t) |   # q attends to t\n      |     t &lt;- t + MHA(Q=t, K=q, V=q) |   # t attends to q\n      +-----------------+---------------+\n                        |   (t is then discarded)\n                        v\n              FFN + LayerNorm on q\n                        |\n                        v\n            masked-mean pool over N\n                        |\n                        v\n                pooled in R^128\n                        |\n                        |     context vector (10d):\n                        |       4 common feats:\n                        |         log(q_len)/6, log(t_len)/6, q_len/(t_len+1), Levenshtein similarity\n                        |       6 auxiliary legacy feats (filled with zeros):\n                        |         \n                        |                       | \n                        |                       |\n                        |                Linear(10 -&gt; 128)\n                        |                       |\n                        +------- concat &lt;-------+\n                                 (256d)\n                                   |\n                                   v\n                        Linear(256 -&gt; 128) -&gt; GELU\n                                   |\n                                   v\n                           Linear(128 -&gt; 1)\n                                   |\n                                   v\n                                sigmoid\n                                   |\n                                   v\n                             P(TM &gt;= 0.5)\n</code></pre>\n<p>d = 128, 4 heads, 5 self-attention layers, 1 bidirectional cross-attention layer. Embedding is plain learned (vocab of 5 for <code>AUGC</code> + pad), positions are learned up to 512.</p>\n<p><strong>Training.</strong></p>\n<ul>\n<li>Labels: 1.42M template-template pairs from GPU US-align, binarized at <code>TM &gt;= 0.5</code>.</li>\n<li>Split: <code>mmseqs_0.300</code> clusters — 15% of clusters held out for validation; train pairs must cross clusters, val pairs must involve at least one val-cluster target.</li>\n<li>Filters: length ratio in [0.75, 1.25], TM in [0.12, 0.92] (drop trivially similar and trivially different), negatives downsampled to 1/3 of positives.</li>\n<li>Loss: <code>BCEWithLogitsLoss(pos_weight)</code> to compensate for residual class imbalance.</li>\n<li>Optimizer: AdamW (lr = 3e-4, wd = 1e-4), cosine schedule, early stop on val Spearman.</li>\n<li>This config (d=128, heads=4, layers=5) was the winner of a small architecture sweep over depth (1-8), width (24-192), and heads (2-8).</li>\n</ul>\n<p><strong>Inference.</strong> For each test target take top-500 templates by sequence similarity, score each one with\n      the CA predictor in batches of 64, sigmoid the logits — that's the <code>model_score</code>. Slot 2 picks <code>argma\n      x(model_score)</code> in a fresh <code>mmseqs_0.300</code> cluster; slot 3 picks <code>argmax(sim × model_score)</code> in yet ano\n      ther fresh cluster. The three TBM slots thus cover three different reasons to like a template: high sequence identity, high predicted TM, and both agreeing at once.</p>\n<h2>Protenix &amp; RNAPro integration</h2>\n<p>Both run in-memory — model loaded once, configs updated per target, no subprocess restart. One caveat up front: <strong>according to my comparisons, the marginal benefit of explicit MSA usage turned out to be minimal for both Protenix and RNAPro on this eval setup</strong> — so a lot of the MSA wiring below is really about <em>not breaking</em> the models rather than squeezing more out of MSA.</p>\n<p>On top of the public inference recipes:</p>\n<ul>\n<li><strong>Co-fold.</strong> Non-RNA chains go in if they fit the token budget; if they miss by ≤100 tokens the largest protein ≥500 aa is trimmed symmetrically on both ends (and its MSA is column-trimmed to match).</li>\n<li><strong>Ligand filtering.</strong> Buffer and crystallization artifacts (<code>GOL</code>, <code>EDO</code>, polyethylene glycols <code>PEG</code>/<code>PG4</code>/<code>P6G</code>/<code>1PE</code>/<code>PE4</code>/<code>EPE</code>, <code>SO4</code>, <code>PO4</code>, salts and buffers <code>CIT</code>/<code>ACT</code>/<code>FMT</code>/<code>ACY</code>/<code>TRS</code>/<code>MES</code>, solvents <code>MPD</code>/<code>IPA</code>/<code>DMS</code>/<code>BME</code>, polyamines <code>SPM</code>/<code>SPD</code>, crystallographic heavy-atom probes <code>NCO</code>/<code>IRI</code>/<code>RHD</code>, <code>HEZ</code>) are blacklisted. Everything else is injected as a <code>CCD_*</code> entity.</li>\n<li><strong>Heterodimer handling.</strong> Multi-chain RNA targets use a separate <code>rnaSequence</code> entity per distinct chain, with the competition MSA column-sliced per chain.</li>\n<li><strong>Long sequences.</strong> Overlapping chunks, per-chunk MSA column slicing, Kabsch alignment on the full overlap, and a narrow ±6-residue linear blend around each handover midpoint (hard switch outside that window — averaging over long overlaps hurt on my val).</li>\n<li><strong>Time budgets.</strong> Both phases have wall-clock budgets with adaptive per-target limits so that a single slow target can't starve the rest.</li>\n</ul>\n<h2>Findings which worked</h2>\n<ul>\n<li><strong>GPU US-align → CA predictor.</strong> The main lever. Without cheap TM labeling there's no supervised signal to rank templates <em>structurally</em>. Without the rerank, slot 2 is \"take rank-2 by sequence\", which is noisy.</li>\n<li><strong>Diverse TBM slot filling (bio / model / combined). I believe this was the biggest single source of Bo5 gain.</strong> Strictly better than any single ranking on my held-out targets — the bio rank anchors, the model rank finds structurally similar but sequence-diverse templates, the combined rank catches the cases where both agree.</li>\n<li><strong>mmseqs-cluster diversity inside TBM.</strong> Requiring the three TBM slots to come from different <code>mmseqs_0.300</code> clusters was bigger than any scoring tweak — it eliminates near-duplicate templates that all fail the same way.</li>\n<li><strong>Keeping PTX and RNAPro as independent single-slot de-novo predictions</strong> beat letting either model take 2 slots on any consistent subset of targets in my evals. It would probably be more efficient with a conditional slot allocator, but I didn't have enough time to tune such a selector to the point where it consistently earned its keep.</li>\n</ul>\n<h2>What didn't work (or helped only a little)</h2>\n<ul>\n<li><strong>Coordinate-space blending.</strong> Whether uniform, learned, or anchor-based, averaging Kabsch-aligned templates moved scores a little on close templates and hurt on diverse ones. Pick-one-template was simpler and better.</li>\n<li><strong>Explicit secondary-structure features.</strong> I tried folding SS in via common tools like ViennaRNA — both as extra inputs to the Cross-Attention selector and as constraints when stitching chunk-wise de-novo predictions. Neither meaningfully moved the score. Entirely possible this is a skill issue on my end: I'm not confident I was using SS in the right way.</li>\n<li><strong>Better chunk-to-chunk blending.</strong> Tried various stitching schemes — linear vs cubic-spline blends, wider overlap windows, even an L-BFGS fit over the overlap region — but the underlying de-novo chunks were probably just not accurate enough for this to show up in the end score.</li>\n<li><strong>Metadata-based routing of PTX vs. RNAPro slots.</strong> Code computes <code>ptx_heavy</code> / <code>rnapro_heavy</code> / <code>balanced</code> labels from length, protein/DNA presence, and keywords, but the final allocation is flat (1 PTX + 1 RNAPro for everyone). The routing labels are reported for statistics only. Attempts to actually route to 2-of-one model lost consistency on the val set.</li>\n</ul>\n<h2>Attachments</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/datasets/gapchenko/rna3d2-pairwise-tm-labels\" target=\"_blank\">https://www.kaggle.com/datasets/gapchenko/rna3d2-pairwise-tm-labels</a> — 1.42M ground-truth TM-scores for pairs of RNA\nstructures from the competition training set</li>\n<li><a href=\"https://www.kaggle.com/code/gapchenko/gpu-accelerated-usalign/notebook\" target=\"_blank\">https://www.kaggle.com/code/gapchenko/gpu-accelerated-usalign/notebook</a> —CUDA port of TMalign_main from\nUS-align. Includes performance and accuracy comparison against CPU original version.</li>\n<li>train_tm_cross_attn.py — training script for the Cross-Attention TM predictor. Reference\nimplementation showing the full pipeline end-to-end (data loading, feature construction, model, training loop, eval by\nSpearman).</li>\n</ul>\n<h2>Acknowledgements</h2>\n<ul>\n<li>Organizers for running Part 2 so soon after Part 1, and for the clean data release.</li>\n<li><code>@theoviel</code> and <code>@jaejohn</code> for the baseline TBM and RNAPro notebooks that everyone built on.</li>\n<li>Protenix (ByteDance) and RNAPro (NVIDIA Digital Bio) for the open checkpoints.</li>\n<li>Claude (Anthropic) was my pair-programmer for most of this, especially for the CUDA port.</li>\n</ul>",
      "rawMarkdown": "# 4th Place Solution — Cross-Attention Template Reranker + Protenix + RNAPro\n\n## Summary\n\nThis was my first bioinformatics competition, and more generally my first serious Kaggle competition overall. I was curious to see how far I could go without domain expertise, and with AI tools as my main leverage. Big thanks to the organizers for the huge amount of work they put into Part 2, to the community for the excellent public datasets and kernels, and congratulations to everyone on the prize list!\n\nThe pipeline is **Protenix + RNAPro + template search reranked by a Cross-Attention TM predictor**. The core idea everything else hangs off is a **CUDA port of US-align** that made it cheap to label TM-scores for a very large number of template–template pairs (~1.42M). That labeled set is what the Cross-Attention model is trained on, which turns \"take top-N by sequence identity\" into a ranking that reorders candidates by **predicted structural TM**.\n\n## Slot allocation (5 predictions per target)\n\n| Slot | Source |\n|------|--------|\n| 1 | TBM — **Bio top-1** (`Bio.Align.PairwiseAligner` in global mode, competition-tuned scores) |\n| 2 | TBM — **Model top-1** (Cross-Attn predicted TM, mmseqs-cluster-diverse from slot 1) |\n| 3 | TBM — **Combined top-1** (`sim * model_score`, cluster-diverse from slots 1-2) |\n| 4 | **Protenix v1** (co-fold with protein/DNA/ligands if they fit) |\n| 5 | **RNAPro** |\n\n\nFor targets longer than 512 nt the Cross-Attn model isn't run (it was trained up to 512) and slots 2-3 fall back to Bio ranks 2 and 3 with cluster diversity.\n\n## GPU US-align (the labeling engine)\n\nhttps://www.kaggle.com/code/gapchenko/gpu-accelerated-usalign/notebook\n\nI re-implemented `TMalign_main` from US-align in CUDA — faithful to the CPU source. All five phases (initial alignment, SS alignment, local fragment superposition, SS + distance matrix, fragment gapless threading) run on GPU; `DP_iter_gpu`, `gpu_TMscore8_search`, and `get_initial5_gpu_fast` are the heavy kernels. A binary stdin batch mode takes coordinate pairs directly via `struct.pack` — no PDB I/O. A separate PDB-batch mode uses a CPU thread pool with one thread per CUDA stream, so the GPU stays saturated while files are parsed.\n\nOn P100-class hardware this gets to hundreds of milliseconds per RNA pair, vs. seconds on the CPU binary — with the biggest speedup on longer molecules (>2k nt). That turned a \"days on CPU\" labeling problem into a \"hours on one GPU\" one, and is the prerequisite for the Cross-Attention predictor below.\n\n**What got labeled**: every pair of available templates in the DB with length ratio in [0.75, 1.25] (~1.42M pairs). These pairs together with their C1' ground-truth coordinates are the supervised signal — TM-score is what I want the network to predict.\n\n## Cross-Attention TM predictor\n\nThe model takes **two nucleotide sequences** (query and template) as independent inputs. Each passes through a shared encoder (embed -> 5 self-attention layers); the two encoded sequences then interact through one bidirectional cross-attention block. The query is pooled, fused with a small context vector, and mapped to a single logit for `P(TM >= 0.5)`.\n\n```\n        query seq                  template seq\n      (N <= 512 nt)              (M <= 512 nt)\n           |                           |\n           |   shared encoder applied  |\n           |   to each independently   |\n           v                           v\n    embed(vocab=5, d=128) + learned pos(max_len=512, d=128)\n           |                           |\n           v                           v\n    5 x [ self-attention (h=4, d=128) + FFN ]\n           |                           |\n           v                           v\n       q_enc (N, 128)           t_enc (M, 128)\n           \\                         /\n            \\                       /\n             v                     v\n      +---------------------------------+\n      |   bidirectional cross-attention |\n      |     q <- q + MHA(Q=q, K=t, V=t) |   # q attends to t\n      |     t <- t + MHA(Q=t, K=q, V=q) |   # t attends to q\n      +-----------------+---------------+\n                        |   (t is then discarded)\n                        v\n              FFN + LayerNorm on q\n                        |\n                        v\n            masked-mean pool over N\n                        |\n                        v\n                pooled in R^128\n                        |\n                        |     context vector (10d):\n                        |       4 common feats:\n                        |         log(q_len)/6, log(t_len)/6, q_len/(t_len+1), Levenshtein similarity\n                        |       6 auxiliary legacy feats (filled with zeros):\n                        |         \n                        |                       | \n                        |                       |\n                        |                Linear(10 -> 128)\n                        |                       |\n                        +------- concat <-------+\n                                 (256d)\n                                   |\n                                   v\n                        Linear(256 -> 128) -> GELU\n                                   |\n                                   v\n                           Linear(128 -> 1)\n                                   |\n                                   v\n                                sigmoid\n                                   |\n                                   v\n                             P(TM >= 0.5)\n```\n\nd = 128, 4 heads, 5 self-attention layers, 1 bidirectional cross-attention layer. Embedding is plain learned (vocab of 5 for `AUGC` + pad), positions are learned up to 512.\n\n**Training.**\n- Labels: 1.42M template-template pairs from GPU US-align, binarized at `TM >= 0.5`.\n- Split: `mmseqs_0.300` clusters — 15% of clusters held out for validation; train pairs must cross clusters, val pairs must involve at least one val-cluster target.\n- Filters: length ratio in [0.75, 1.25], TM in [0.12, 0.92] (drop trivially similar and trivially different), negatives downsampled to 1/3 of positives.\n- Loss: `BCEWithLogitsLoss(pos_weight)` to compensate for residual class imbalance.\n- Optimizer: AdamW (lr = 3e-4, wd = 1e-4), cosine schedule, early stop on val Spearman.\n- This config (d=128, heads=4, layers=5) was the winner of a small architecture sweep over depth (1-8), width (24-192), and heads (2-8).\n\n**Inference.** For each test target take top-500 templates by sequence similarity, score each one with\n      the CA predictor in batches of 64, sigmoid the logits — that's the `model_score`. Slot 2 picks `argma\n      x(model_score)` in a fresh `mmseqs_0.300` cluster; slot 3 picks `argmax(sim × model_score)` in yet ano\n      ther fresh cluster. The three TBM slots thus cover three different reasons to like a template: high sequence identity, high predicted TM, and both agreeing at once.\n\n## Protenix & RNAPro integration\n\nBoth run in-memory — model loaded once, configs updated per target, no subprocess restart. One caveat up front: **according to my comparisons, the marginal benefit of explicit MSA usage turned out to be minimal for both Protenix and RNAPro on this eval setup** — so a lot of the MSA wiring below is really about *not breaking* the models rather than squeezing more out of MSA.\n\nOn top of the public inference recipes:\n\n- **Co-fold.** Non-RNA chains go in if they fit the token budget; if they miss by ≤100 tokens the largest protein ≥500 aa is trimmed symmetrically on both ends (and its MSA is column-trimmed to match).\n- **Ligand filtering.** Buffer and crystallization artifacts (`GOL`, `EDO`, polyethylene glycols `PEG`/`PG4`/`P6G`/`1PE`/`PE4`/`EPE`, `SO4`, `PO4`, salts and buffers `CIT`/`ACT`/`FMT`/`ACY`/`TRS`/`MES`, solvents `MPD`/`IPA`/`DMS`/`BME`, polyamines `SPM`/`SPD`, crystallographic heavy-atom probes `NCO`/`IRI`/`RHD`, `HEZ`) are blacklisted. Everything else is injected as a `CCD_*` entity.\n- **Heterodimer handling.** Multi-chain RNA targets use a separate `rnaSequence` entity per distinct chain, with the competition MSA column-sliced per chain.\n- **Long sequences.** Overlapping chunks, per-chunk MSA column slicing, Kabsch alignment on the full overlap, and a narrow ±6-residue linear blend around each handover midpoint (hard switch outside that window — averaging over long overlaps hurt on my val).\n- **Time budgets.** Both phases have wall-clock budgets with adaptive per-target limits so that a single slow target can't starve the rest.\n\n## Findings which worked\n\n- **GPU US-align → CA predictor.** The main lever. Without cheap TM labeling there's no supervised signal to rank templates *structurally*. Without the rerank, slot 2 is \"take rank-2 by sequence\", which is noisy.\n- **Diverse TBM slot filling (bio / model / combined). I believe this was the biggest single source of Bo5 gain.** Strictly better than any single ranking on my held-out targets — the bio rank anchors, the model rank finds structurally similar but sequence-diverse templates, the combined rank catches the cases where both agree.\n- **mmseqs-cluster diversity inside TBM.** Requiring the three TBM slots to come from different `mmseqs_0.300` clusters was bigger than any scoring tweak — it eliminates near-duplicate templates that all fail the same way.\n- **Keeping PTX and RNAPro as independent single-slot de-novo predictions** beat letting either model take 2 slots on any consistent subset of targets in my evals. It would probably be more efficient with a conditional slot allocator, but I didn't have enough time to tune such a selector to the point where it consistently earned its keep.\n\n## What didn't work (or helped only a little)\n\n- **Coordinate-space blending.** Whether uniform, learned, or anchor-based, averaging Kabsch-aligned templates moved scores a little on close templates and hurt on diverse ones. Pick-one-template was simpler and better.\n- **Explicit secondary-structure features.** I tried folding SS in via common tools like ViennaRNA — both as extra inputs to the Cross-Attention selector and as constraints when stitching chunk-wise de-novo predictions. Neither meaningfully moved the score. Entirely possible this is a skill issue on my end: I'm not confident I was using SS in the right way.\n- **Better chunk-to-chunk blending.** Tried various stitching schemes — linear vs cubic-spline blends, wider overlap windows, even an L-BFGS fit over the overlap region — but the underlying de-novo chunks were probably just not accurate enough for this to show up in the end score.\n- **Metadata-based routing of PTX vs. RNAPro slots.** Code computes `ptx_heavy` / `rnapro_heavy` / `balanced` labels from length, protein/DNA presence, and keywords, but the final allocation is flat (1 PTX + 1 RNAPro for everyone). The routing labels are reported for statistics only. Attempts to actually route to 2-of-one model lost consistency on the val set.\n\n## Attachments\n\n  - https://www.kaggle.com/datasets/gapchenko/rna3d2-pairwise-tm-labels — 1.42M ground-truth TM-scores for pairs of RNA\n  structures from the competition training set\n  - https://www.kaggle.com/code/gapchenko/gpu-accelerated-usalign/notebook —CUDA port of TMalign_main from\n  US-align. Includes performance and accuracy comparison against CPU original version.\n  - train_tm_cross_attn.py — training script for the Cross-Attention TM predictor. Reference\n  implementation showing the full pipeline end-to-end (data loading, feature construction, model, training loop, eval by\n   Spearman).\n\n## Acknowledgements\n\n- Organizers for running Part 2 so soon after Part 1, and for the clean data release.\n- `@theoviel` and `@jaejohn` for the baseline TBM and RNAPro notebooks that everyone built on.\n- Protenix (ByteDance) and RNAPro (NVIDIA Digital Bio) for the open checkpoints.\n- Claude (Anthropic) was my pair-programmer for most of this, especially for the CUDA port.\n",
      "votes": 9
    },
    {
      "id": 3463073,
      "postDate": "2026-05-25T19:46:57.977Z",
      "content": "<p>Very interesting approach for finding new templates. It's what I have been working on in this competition. My approach is representation-based alignment using a representation derived from an RNA language model. However, I don't see much improvement in this method. I am curious about how much improvement there is compared to sequence alignment in your approach.</p>",
      "rawMarkdown": "Very interesting approach for finding new templates. It's what I have been working on in this competition. My approach is representation-based alignment using a representation derived from an RNA language model. However, I don't see much improvement in this method. I am curious about how much improvement there is compared to sequence alignment in your approach.",
      "votes": 1,
      "replies": [
        {
          "id": 3463077,
          "postDate": "2026-05-25T19:57:31.943Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3446070,
      "postDate": "2026-04-20T23:33:14.850Z",
      "content": "<p><a href=\"https://www.kaggle.com/gapchenko\" target=\"_blank\">@gapchenko</a> This is an impressive contribution, thanks for posting!  I'm curious to see some of the details -- would you be amenable to also posting the inference notebook?</p>",
      "rawMarkdown": "@gapchenko This is an impressive contribution, thanks for posting!  I'm curious to see some of the details -- would you be amenable to also posting the inference notebook?",
      "votes": 1,
      "replies": [
        {
          "id": 3446490,
          "postDate": "2026-04-21T17:30:35.390Z",
          "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> Thanks for the kind words! \nI've attached the inference notebook to the writeup (lightly cleaned for readability, but semantically the same as the scored submission).</p>",
          "rawMarkdown": "@rhijudas Thanks for the kind words! \nI've attached the inference notebook to the writeup (lightly cleaned for readability, but semantically the same as the scored submission).",
          "replies": [
            {
              "id": 3446586,
              "postDate": "2026-04-21T20:00:20.910Z",
              "content": "<p>I wasn't sure if the notebook was semantically the same -- thanks for verifying. Congratulations on the solution!</p>",
              "rawMarkdown": "I wasn't sure if the notebook was semantically the same -- thanks for verifying. Congratulations on the solution!"
            },
            {
              "id": 3446644,
              "postDate": "2026-04-21T22:23:17.260Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3446649,
              "postDate": "2026-04-21T22:50:08.993Z",
              "content": "<p>Sorry for the mix-up!\n  <a href=\"https://www.kaggle.com/code/gapchenko/gpu-accelerated-usalign\" target=\"_blank\">https://www.kaggle.com/code/gapchenko/gpu-accelerated-usalign</a> is just the CUDA port of USAlign — only used for labeling template pairs.\n   <a href=\"https://www.kaggle.com/code/gapchenko/rna3d2-cross-attn-reranker\" target=\"_blank\">https://www.kaggle.com/code/gapchenko/rna3d2-cross-attn-reranker</a> is the actual inference notebook I attached today🙂 \nLet me know if any questions arise!</p>",
              "rawMarkdown": "Sorry for the mix-up!\n  https://www.kaggle.com/code/gapchenko/gpu-accelerated-usalign is just the CUDA port of USAlign — only used for labeling template pairs.\n   https://www.kaggle.com/code/gapchenko/rna3d2-cross-attn-reranker is the actual inference notebook I attached today🙂 \nLet me know if any questions arise!"
            }
          ]
        }
      ]
    },
    {
      "id": 3468417,
      "postDate": "2026-06-08T23:57:02.850Z",
      "content": "<p>I am curious to know if you perform an ablation study to isolate the impact of cross-attention reranker </p>",
      "rawMarkdown": "I am curious to know if you perform an ablation study to isolate the impact of cross-attention reranker "
    }
  ],
  "comments": [
    {
      "id": 3463073,
      "author_name": "Guojun1",
      "author_url": "",
      "post_date": "2026-05-25T19:46:57.977000",
      "content": "<p>Very interesting approach for finding new templates. It's what I have been working on in this competition. My approach is representation-based alignment using a representation derived from an RNA language model. However, I don't see much improvement in this method. I am curious about how much improvement there is compared to sequence alignment in your approach.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3463077,
          "author_name": "",
          "author_url": "",
          "post_date": "2026-05-25T19:57:31.943000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3446070,
      "author_name": "Rhiju Das",
      "author_url": "",
      "post_date": "2026-04-20T23:33:14.850000",
      "content": "<p><a href=\"https://www.kaggle.com/gapchenko\" target=\"_blank\">@gapchenko</a> This is an impressive contribution, thanks for posting!  I'm curious to see some of the details -- would you be amenable to also posting the inference notebook?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3446490,
          "author_name": "aleksandr3312",
          "author_url": "",
          "post_date": "2026-04-21T17:30:35.390000",
          "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> Thanks for the kind words! \nI've attached the inference notebook to the writeup (lightly cleaned for readability, but semantically the same as the scored submission).</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3446586,
              "author_name": "Rhiju Das",
              "author_url": "",
              "post_date": "2026-04-21T20:00:20.910000",
              "content": "<p>I wasn't sure if the notebook was semantically the same -- thanks for verifying. Congratulations on the solution!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3446644,
              "author_name": "",
              "author_url": "",
              "post_date": "2026-04-21T22:23:17.260000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3446649,
              "author_name": "aleksandr3312",
              "author_url": "",
              "post_date": "2026-04-21T22:50:08.993000",
              "content": "<p>Sorry for the mix-up!\n  <a href=\"https://www.kaggle.com/code/gapchenko/gpu-accelerated-usalign\" target=\"_blank\">https://www.kaggle.com/code/gapchenko/gpu-accelerated-usalign</a> is just the CUDA port of USAlign — only used for labeling template pairs.\n   <a href=\"https://www.kaggle.com/code/gapchenko/rna3d2-cross-attn-reranker\" target=\"_blank\">https://www.kaggle.com/code/gapchenko/rna3d2-cross-attn-reranker</a> is the actual inference notebook I attached today🙂 \nLet me know if any questions arise!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3468417,
      "author_name": "Guojun1",
      "author_url": "",
      "post_date": "2026-06-08T23:57:02.850000",
      "content": "<p>I am curious to know if you perform an ablation study to isolate the impact of cross-attention reranker </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3444714": "# 4th Place Solution — Cross-Attention Template Reranker + Protenix + RNAPro\n\n## Summary\n\nThis was my first bioinformatics competition, and more generally my first serious Kaggle competition overall. I was curious to see how far I could go without domain expertise, and with AI tools as my main leverage. Big thanks to the organizers for the huge amount of work they put into Part 2, to the community for the excellent public datasets and kernels, and congratulations to everyone on the prize list!\n\nThe pipeline is **Protenix + RNAPro + template search reranked by a Cross-Attention TM predictor**. The core idea everything else hangs off is a **CUDA port of US-align** that made it cheap to label TM-scores for a very large number of template–template pairs (~1.42M). That labeled set is what the Cross-Attention model is trained on, which turns \"take top-N by sequence identity\" into a ranking that reorders candidates by **predicted structural TM**.\n\n## Slot allocation (5 predictions per target)\n\n| Slot | Source |\n|------|--------|\n| 1 | TBM — **Bio top-1** (`Bio.Align.PairwiseAligner` in global mode, competition-tuned scores) |\n| 2 | TBM — **Model top-1** (Cross-Attn predicted TM, mmseqs-cluster-diverse from slot 1) |\n| 3 | TBM — **Combined top-1** (`sim * model_score`, cluster-diverse from slots 1-2) |\n| 4 | **Protenix v1** (co-fold with protein/DNA/ligands if they fit) |\n| 5 | **RNAPro** |\n\n\nFor targets longer than 512 nt the Cross-Attn model isn't run (it was trained up to 512) and slots 2-3 fall back to Bio ranks 2 and 3 with cluster diversity.\n\n## GPU US-align (the labeling engine)\n\nhttps://www.kaggle.com/code/gapchenko/gpu-accelerated-usalign/notebook\n\nI re-implemented `TMalign_main` from US-align in CUDA — faithful to the CPU source. All five phases (initial alignment, SS alignment, local fragment superposition, SS + distance matrix, fragment gapless threading) run on GPU; `DP_iter_gpu`, `gpu_TMscore8_search`, and `get_initial5_gpu_fast` are the heavy kernels. A binary stdin batch mode takes coordinate pairs directly via `struct.pack` — no PDB I/O. A separate PDB-batch mode uses a CPU thread pool with one thread per CUDA stream, so the GPU stays saturated while files are parsed.\n\nOn P100-class hardware this gets to hundreds of milliseconds per RNA pair, vs. seconds on the CPU binary — with the biggest speedup on longer molecules (>2k nt). That turned a \"days on CPU\" labeling problem into a \"hours on one GPU\" one, and is the prerequisite for the Cross-Attention predictor below.\n\n**What got labeled**: every pair of available templates in the DB with length ratio in [0.75, 1.25] (~1.42M pairs). These pairs together with their C1' ground-truth coordinates are the supervised signal — TM-score is what I want the network to predict.\n\n## Cross-Attention TM predictor\n\nThe model takes **two nucleotide sequences** (query and template) as independent inputs. Each passes through a shared encoder (embed -> 5 self-attention layers); the two encoded sequences then interact through one bidirectional cross-attention block. The query is pooled, fused with a small context vector, and mapped to a single logit for `P(TM >= 0.5)`.\n\n```\n        query seq                  template seq\n      (N <= 512 nt)              (M <= 512 nt)\n           |                           |\n           |   shared encoder applied  |\n           |   to each independently   |\n           v                           v\n    embed(vocab=5, d=128) + learned pos(max_len=512, d=128)\n           |                           |\n           v                           v\n    5 x [ self-attention (h=4, d=128) + FFN ]\n           |                           |\n           v                           v\n       q_enc (N, 128)           t_enc (M, 128)\n           \\                         /\n            \\                       /\n             v                     v\n      +---------------------------------+\n      |   bidirectional cross-attention |\n      |     q <- q + MHA(Q=q, K=t, V=t) |   # q attends to t\n      |     t <- t + MHA(Q=t, K=q, V=q) |   # t attends to q\n      +-----------------+---------------+\n                        |   (t is then discarded)\n                        v\n              FFN + LayerNorm on q\n                        |\n                        v\n            masked-mean pool over N\n                        |\n                        v\n                pooled in R^128\n                        |\n                        |     context vector (10d):\n                        |       4 common feats:\n                        |         log(q_len)/6, log(t_len)/6, q_len/(t_len+1), Levenshtein similarity\n                        |       6 auxiliary legacy feats (filled with zeros):\n                        |         \n                        |                       | \n                        |                       |\n                        |                Linear(10 -> 128)\n                        |                       |\n                        +------- concat <-------+\n                                 (256d)\n                                   |\n                                   v\n                        Linear(256 -> 128) -> GELU\n                                   |\n                                   v\n                           Linear(128 -> 1)\n                                   |\n                                   v\n                                sigmoid\n                                   |\n                                   v\n                             P(TM >= 0.5)\n```\n\nd = 128, 4 heads, 5 self-attention layers, 1 bidirectional cross-attention layer. Embedding is plain learned (vocab of 5 for `AUGC` + pad), positions are learned up to 512.\n\n**Training.**\n- Labels: 1.42M template-template pairs from GPU US-align, binarized at `TM >= 0.5`.\n- Split: `mmseqs_0.300` clusters — 15% of clusters held out for validation; train pairs must cross clusters, val pairs must involve at least one val-cluster target.\n- Filters: length ratio in [0.75, 1.25], TM in [0.12, 0.92] (drop trivially similar and trivially different), negatives downsampled to 1/3 of positives.\n- Loss: `BCEWithLogitsLoss(pos_weight)` to compensate for residual class imbalance.\n- Optimizer: AdamW (lr = 3e-4, wd = 1e-4), cosine schedule, early stop on val Spearman.\n- This config (d=128, heads=4, layers=5) was the winner of a small architecture sweep over depth (1-8), width (24-192), and heads (2-8).\n\n**Inference.** For each test target take top-500 templates by sequence similarity, score each one with\n      the CA predictor in batches of 64, sigmoid the logits — that's the `model_score`. Slot 2 picks `argma\n      x(model_score)` in a fresh `mmseqs_0.300` cluster; slot 3 picks `argmax(sim × model_score)` in yet ano\n      ther fresh cluster. The three TBM slots thus cover three different reasons to like a template: high sequence identity, high predicted TM, and both agreeing at once.\n\n## Protenix & RNAPro integration\n\nBoth run in-memory — model loaded once, configs updated per target, no subprocess restart. One caveat up front: **according to my comparisons, the marginal benefit of explicit MSA usage turned out to be minimal for both Protenix and RNAPro on this eval setup** — so a lot of the MSA wiring below is really about *not breaking* the models rather than squeezing more out of MSA.\n\nOn top of the public inference recipes:\n\n- **Co-fold.** Non-RNA chains go in if they fit the token budget; if they miss by ≤100 tokens the largest protein ≥500 aa is trimmed symmetrically on both ends (and its MSA is column-trimmed to match).\n- **Ligand filtering.** Buffer and crystallization artifacts (`GOL`, `EDO`, polyethylene glycols `PEG`/`PG4`/`P6G`/`1PE`/`PE4`/`EPE`, `SO4`, `PO4`, salts and buffers `CIT`/`ACT`/`FMT`/`ACY`/`TRS`/`MES`, solvents `MPD`/`IPA`/`DMS`/`BME`, polyamines `SPM`/`SPD`, crystallographic heavy-atom probes `NCO`/`IRI`/`RHD`, `HEZ`) are blacklisted. Everything else is injected as a `CCD_*` entity.\n- **Heterodimer handling.** Multi-chain RNA targets use a separate `rnaSequence` entity per distinct chain, with the competition MSA column-sliced per chain.\n- **Long sequences.** Overlapping chunks, per-chunk MSA column slicing, Kabsch alignment on the full overlap, and a narrow ±6-residue linear blend around each handover midpoint (hard switch outside that window — averaging over long overlaps hurt on my val).\n- **Time budgets.** Both phases have wall-clock budgets with adaptive per-target limits so that a single slow target can't starve the rest.\n\n## Findings which worked\n\n- **GPU US-align → CA predictor.** The main lever. Without cheap TM labeling there's no supervised signal to rank templates *structurally*. Without the rerank, slot 2 is \"take rank-2 by sequence\", which is noisy.\n- **Diverse TBM slot filling (bio / model / combined). I believe this was the biggest single source of Bo5 gain.** Strictly better than any single ranking on my held-out targets — the bio rank anchors, the model rank finds structurally similar but sequence-diverse templates, the combined rank catches the cases where both agree.\n- **mmseqs-cluster diversity inside TBM.** Requiring the three TBM slots to come from different `mmseqs_0.300` clusters was bigger than any scoring tweak — it eliminates near-duplicate templates that all fail the same way.\n- **Keeping PTX and RNAPro as independent single-slot de-novo predictions** beat letting either model take 2 slots on any consistent subset of targets in my evals. It would probably be more efficient with a conditional slot allocator, but I didn't have enough time to tune such a selector to the point where it consistently earned its keep.\n\n## What didn't work (or helped only a little)\n\n- **Coordinate-space blending.** Whether uniform, learned, or anchor-based, averaging Kabsch-aligned templates moved scores a little on close templates and hurt on diverse ones. Pick-one-template was simpler and better.\n- **Explicit secondary-structure features.** I tried folding SS in via common tools like ViennaRNA — both as extra inputs to the Cross-Attention selector and as constraints when stitching chunk-wise de-novo predictions. Neither meaningfully moved the score. Entirely possible this is a skill issue on my end: I'm not confident I was using SS in the right way.\n- **Better chunk-to-chunk blending.** Tried various stitching schemes — linear vs cubic-spline blends, wider overlap windows, even an L-BFGS fit over the overlap region — but the underlying de-novo chunks were probably just not accurate enough for this to show up in the end score.\n- **Metadata-based routing of PTX vs. RNAPro slots.** Code computes `ptx_heavy` / `rnapro_heavy` / `balanced` labels from length, protein/DNA presence, and keywords, but the final allocation is flat (1 PTX + 1 RNAPro for everyone). The routing labels are reported for statistics only. Attempts to actually route to 2-of-one model lost consistency on the val set.\n\n## Attachments\n\n  - https://www.kaggle.com/datasets/gapchenko/rna3d2-pairwise-tm-labels — 1.42M ground-truth TM-scores for pairs of RNA\n  structures from the competition training set\n  - https://www.kaggle.com/code/gapchenko/gpu-accelerated-usalign/notebook —CUDA port of TMalign_main from\n  US-align. Includes performance and accuracy comparison against CPU original version.\n  - train_tm_cross_attn.py — training script for the Cross-Attention TM predictor. Reference\n  implementation showing the full pipeline end-to-end (data loading, feature construction, model, training loop, eval by\n   Spearman).\n\n## Acknowledgements\n\n- Organizers for running Part 2 so soon after Part 1, and for the clean data release.\n- `@theoviel` and `@jaejohn` for the baseline TBM and RNAPro notebooks that everyone built on.\n- Protenix (ByteDance) and RNAPro (NVIDIA Digital Bio) for the open checkpoints.\n- Claude (Anthropic) was my pair-programmer for most of this, especially for the CUDA port.\n",
    "3463073": "Very interesting approach for finding new templates. It's what I have been working on in this competition. My approach is representation-based alignment using a representation derived from an RNA language model. However, I don't see much improvement in this method. I am curious about how much improvement there is compared to sequence alignment in your approach.",
    "3446070": "@gapchenko This is an impressive contribution, thanks for posting!  I'm curious to see some of the details -- would you be amenable to also posting the inference notebook?",
    "3468417": "I am curious to know if you perform an ablation study to isolate the impact of cross-attention reranker "
  }
}