{
  "id": 691133,
  "title": "2nd Place (Public 1st Place) Solution",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/691133",
  "author_name": "AyPy",
  "post_date": "2026-04-14T00:52:33.013000",
  "votes": 21,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First, I would like to extend my deepest gratitude to the competition organizing team and Kaggle staff for hosting this fascinating competition. Having previously worked in experimental sciences including organic chemistry, I am well aware of the challenges in obtaining molecular 3D structural data (ground-truth structures). Considering this, I'm deeply impressed that they were able to hold this Part II competition just less than a year after the <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding\" target=\"_blank\">Part I competition</a>.\nDuring these approximately two months of trial-and-error work, I gained both meaningful experience and valuable learning.</p>\n<h2>Summary of My Solution</h2>\n<p>I built \"<strong>BPP-Protenix</strong>\", a model that integrates Base Pair Probability (BPP) features into an AlphaFold3-style architecture. The model design was inspired by the architecture of <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/668412\" target=\"_blank\">RNAPro</a> shared by <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>. My final pipeline is an ensemble of TBM, RNAPro, and BPP-Protenix (with plain Protenix as fallback for targets not suitable for BPP calculation). Although I dropped one place in the private LB, after completing the development of BPP-Protenix, I held 1st place on the public LB for the final two weeks of the competition. I guess almost no other participants modified the neural network architecture in this competition, so this appears to be a unique solution.</p>\n<p><strong>Submission slot allocation is follows</strong>:</p>\n<table>\n<thead>\n<tr>\n<th>slot 1</th>\n<th>slot 2</th>\n<th>slot 3</th>\n<th>slot 4</th>\n<th>slot5</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>TBM</td>\n<td>RNAPro</td>\n<td>RNAPro</td>\n<td><strong>BPP-Protenix</strong><br>(with plain Protenix fallback)</td>\n<td><strong>BPP-Protenix</strong><br>(with plain Protenix fallback)</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><p><strong>TBM</strong>: I used <a href=\"https://www.kaggle.com/code/kami1976/stanford-rna-3d-folding-part-2a18\" target=\"_blank\">the notebook</a> published by <a href=\"https://www.kaggle.com/kami1976\" target=\"_blank\">@kami1976</a> almost exactly as is. Since it appeared to offer good local scores and seed stability, I decided to borrow this one.</p></li>\n<li><p><strong>RNAPro</strong>: Using the template created by the above TBM as input data, I applied the <a href=\"https://www.kaggle.com/code/jaejohn/rnapro-inference-with-tbm\" target=\"_blank\">the pipeline</a> published by <a href=\"https://www.kaggle.com/jaejohn\" target=\"_blank\">@jaejohn</a>.</p></li>\n<li><p><strong>BPP-Protenix</strong>: This is the model that required the most time to develop for this competition. I'll explain it below.</p></li>\n</ul>\n<h2>Core Idea: BPP as Structural Prior</h2>\n<h3>Inspiration: Ribonanza Competition (2023)</h3>\n<p>The idea came from the <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding\" target=\"_blank\">Stanford Ribonanza RNA Folding</a> competition in 2023. <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> (one of the hosts of the Ribonanza competition) had published about <a href=\"https://academic.oup.com/bib/article/24/1/bbac581/6986359\" target=\"_blank\">RNAdegformer</a>, a transformer predicting mRNA degradation at nucleotide level. In that paper, <strong>BPP (Base Pair Probability) matrix was shown to be a very important feature</strong> for predicting RNA stability. Inspired by this, almost all top Ribonanza participants added BPP embeddings as bias terms in their Transformers.</p>\n<h3>From Stability Prediction to 3D Structure Prediction</h3>\n<p>The task of RNAdegformer and Ribonanza models was stability prediction, but RNA stability depends a lot on 3D structure. Furthermore, the base pair formation probability information obtained through secondary structure prediction provides proximity constraints on the 3D structure, thereby functioning as an effective inductive bias. Given that, providing BPP information to a 3D structure predictor seemed like a natural extension.</p>\n<h3>BPP Injection into Pairformer</h3>\n<p>The BPP matrix <code>[N, N]</code> naturally corresponds to the pair representation <code>z [N, N, c]</code> in the Pairformer trunk of AlphaFold3/Protenix architecture — both are pairwise matrices indexed by residue positions. By passing BPP through a simple linear layer (1 → c) and adding the resulting embedding to <code>z_init</code>, we can inject RNA structural knowledge directly into the model.</p>\n<pre><code>BPP [N, N]\n  → unsqueeze → [N, N, 1]\n  → LinearNoBias(1 → 128) → [N, N, 128]\n  → z_init += bpp_embedding\n</code></pre>\n<h2>BPP Calculator Selection: EternaFold vs ViennaRNA</h2>\n<p>I evaluated BPP quality against experimental 3D coordinates from the training data, checking if residue pairs with high BPP are actually close in 3D space. <a href=\"https://github.com/ViennaRNA/ViennaRNA\" target=\"_blank\">ViennaRNA</a> and <a href=\"https://github.com/WaymentSteeleLab/EternaFold\" target=\"_blank\">EternaFold</a>, two of the primary RNA secondary structure computation tools, were tested. These secondary structure prediction tools assume RNA monomers as the target of prediction. Additionally, calculating BPP for long sequences during a Kaggle session is difficult in inference. Trherefore, it was considered that the training should also be performed using shorter sequences. The following filter was applied to the approximately 5,000 sequences of Kaggle's official training data:</p>\n<p><strong>Filter conditions</strong></p>\n<ul>\n<li>RNA monomer (no multimer)</li>\n<li>Sequence length ≤ 1,000</li>\n</ul>\n<p>The BPP quality was evaluated using the <strong>1,532</strong> sequences remaining after applying this filter. Here, a C1'-C1' distance &lt; 12 Å was defined as \"contact.\" However, this judge was only conducted for C1' combinations separated by 4 or more residues (because such residues exist spatially close regardless of base pair formation).</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>ViennaRNA</th>\n<th>EternaFold</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Contact rate (BPP ≥ 0.7)</td>\n<td>81.4%</td>\n<td><strong>96.2%</strong></td>\n</tr>\n<tr>\n<td>Contact rate (BPP ≥ 0.9)</td>\n<td>86.6%</td>\n<td><strong>97.4%</strong></td>\n</tr>\n<tr>\n<td>Random baseline*</td>\n<td>3.6%</td>\n<td>3.6%</td>\n</tr>\n</tbody>\n</table>\n<p>*Random baseline: contact rate when residue pairs are selected randomly (no BPP information).</p>\n<p>EternaFold BPP was clearly more accurate. Also, it is easily deployable in Kaggle notebooks via the <a href=\"https://github.com/DasLab/arnie\" target=\"_blank\">arnie</a> library. So I chose EternaFold for feature engineering.</p>\n<hr>\n<h2>BPP-Protenix training</h2>\n<h3>Setup</h3>\n<ul>\n<li><strong>Pretrained checkpoint</strong>: <code>protenix_base_20250630_v1.0.0</code></li>\n<li><strong>Train data</strong>: 619 PDBs (RNA-only)<ul>\n<li>From the 1,532 RNA monomers (≤1000 nt) used for BPP quality evaluation mentioned above, PDBs with protein/DNA partners were excluded.</li></ul></li>\n<li><strong>Validation data</strong>: 33 PDBs<ul>\n<li>Combined Kaggle's official validation set with my own collected PDBs released between 2025-12-03 and 2026-02-18. After applying the same filter, 33 records remained.</li></ul></li>\n</ul>\n<h3>Hyperparameters</h3>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>diffusion_batch_size</td>\n<td>32</td>\n</tr>\n<tr>\n<td>train_crop_size</td>\n<td>550</td>\n</tr>\n<tr>\n<td>lr</td>\n<td>1e-4</td>\n</tr>\n<tr>\n<td>EMA decay</td>\n<td>0.995</td>\n</tr>\n<tr>\n<td>num_steps</td>\n<td>4960</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>BPP-Protenix Inference</h2>\n<h3>BPP-Available Target Filter &amp; Fallback to Plain Protenix</h3>\n<p>The following filters were applied to determine if the target was appropriate for BPP calculation by EternaFold:</p>\n<ul>\n<li>RNA monomer (no multimer)</li>\n<li>Total entity length ≤ 850 tokens</li>\n<li>No protein/DNA partners</li>\n</ul>\n<p>Targets not passing the filter were predicted by <strong>plain Protenix</strong> as fallback. Targets exceeding 850 tokens were cropped to 850 residues, with remaining coordinates zero-padded.</p>\n<h3>BPP-Protenix Standalone Performance</h3>\n<p>Before evaluating the ensemble, I assessed BPP-Protenix alone (all 5 slots filled with Protenix predictions). Inspired by RNAPro's TemplateEmbedder (Protenix v0.5), I tested three integration methods (A,B,C).</p>\n<ul>\n<li><strong>Method A</strong>: just a linear layer from 1 to 128 dimensions, then embedding is added to <code>z_init</code>.</li>\n<li><strong>Method B</strong>: In addition to the direct linear layer, I also converted BPP into a 20-dimensional one-hot vector by binning, embedded it with another linear layer, and added both features to z_init.</li>\n<li><strong>Method C</strong>: In addition to method B, uses a BppEmbedder. It is the similar design pattern as TemplateEmbedder in RNAPro. Applies a small 2-block Pairformer before the main 48-block Pairformer in each recycling cycle.</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Architecture</th>\n<th>Public LB (standalone*)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>A (Linear)</strong></td>\n<td>LinearNoBias(1→128), add BPP-emb to z_init</td>\n<td><strong>0.340</strong></td>\n</tr>\n<tr>\n<td>B (+ Binning)</td>\n<td>Method A + one_hot(20 bins) → Linear(20→128),  add two embs to z_init</td>\n<td>0.317</td>\n</tr>\n<tr>\n<td>C (+ BppEmbedder)</td>\n<td>Method B + BppEmbedder in recycling loop</td>\n<td>0.320</td>\n</tr>\n<tr>\n<td>Plain Protenix (control)</td>\n<td>-</td>\n<td>0.309</td>\n</tr>\n</tbody>\n</table>\n<p>*Standalone: all 5 slots filled with BPP-Protenix + plain Protenix fallback</p>\n<p>The standalone BPP-Protenix (Method A) scored <strong>0.34</strong> on public LB. This is comparable to TBM-only public notebooks (~0.35), demonstrating that <strong>NN-based predictions can match template-based approaches</strong>.</p>\n<p>Interesting finding: allowing protein/DNA partner targets in the BPP filter caused almost no change in public LB score. Possibly few such targets existed in the test set. Worth investigating the performance for RNA/protein assembly with late submissions.</p>\n<hr>\n<h2>Pseudoknot Analysis</h2>\n<p><strong>Concern</strong>: EternaFold does not predict pseudoknots. Could BPP-Protenix is not suitable for potential pseudoknot targets?</p>\n<p><strong>Post-hoc validation</strong>:</p>\n<ul>\n<li>Detected pseudoknots from ground-truth 3D coordinates using biotite (base pair detection + crossing pair check)</li>\n<li>Validation set (11 RNA-only targets): pseudoknots found in <strong>7 out of 11</strong> targets</li>\n<li>BPP-Protenix outperformed plain Protenix by <strong>~0.06 TM-score</strong> on average across all 11 targets</li>\n<li>Contrary to expectations, even within the pseudoknot group, BPP-Protenix showed higher TM-scores</li>\n</ul>\n<p><strong>Interpretation</strong>: BPP provides a \"pairing tendency hint\" as continuous values, not a complete secondary structure prediction. Even if pseudoknot base pairs are missed, non-pseudoknot BPP information is still valuable. And maybe, the 48-block Pairformer can absorb BPP inaccuracies.</p>\n<p>Based on the local validation result, I have decided not to take any special measures specifically targeting pseudoknots.</p>\n<hr>\n<h2>Ensemble Pipeline</h2>\n<h3>TBM (Slot 1) and RNAPro (Slot 2–3)</h3>\n<ul>\n<li>Used a public notebook shared early in the competition as-is</li>\n</ul>\n<h3>BPP-Protenix (Slot 4–5)</h3>\n<ul>\n<li>Targets passing the BPP-available filter → BPP-Protenix (seed=101, N_sample=5, top 2 by confidence score)</li>\n<li>Targets failing the filter → Plain Protenix fallback (same config, no BPP, Longer targets were cropped to 850 tokens and zero-padded)</li>\n</ul>\n<h3>Public/Private LB Results</h3>\n<table>\n<thead>\n<tr>\n<th>Configuration</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>TBM 2slot / RNAPro 3slot (baseline)</td>\n<td>~0.43</td>\n<td>—</td>\n</tr>\n<tr>\n<td>TBM 1slot / RNAPro 2slot / Protenix 2slot</td>\n<td>0.461</td>\n<td>0.475</td>\n</tr>\n<tr>\n<td><strong>TBM 1slot / RNAPro 2slot / BPP-Protenix 2slot</strong></td>\n<td><strong>0.504 (1st place)</strong></td>\n<td><strong>0.492 (2nd place)</strong></td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>BPP-Protenix yielded <strong>+0.04</strong> over plain Protenix and <strong>+0.07</strong> over the TBM+RNAPro baseline in the public LB score</li>\n<li>BPP-Protenix demonstrated strong accuracy on the private test set as well</li>\n</ul>\n<h2>Acknowledgments</h2>\n<p>Thank you to the competition hosts (Rhiju Das and team) for organizing this fascinating competition. Thanks also to the Kaggle community for sharing public notebooks (TBM, RNAPro) that formed the foundation of my ensemble.</p>",
  "messages": [
    {
      "id": 3441423,
      "postDate": "2026-04-14T00:52:33.013Z",
      "content": "<p>First, I would like to extend my deepest gratitude to the competition organizing team and Kaggle staff for hosting this fascinating competition. Having previously worked in experimental sciences including organic chemistry, I am well aware of the challenges in obtaining molecular 3D structural data (ground-truth structures). Considering this, I'm deeply impressed that they were able to hold this Part II competition just less than a year after the <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding\" target=\"_blank\">Part I competition</a>.\nDuring these approximately two months of trial-and-error work, I gained both meaningful experience and valuable learning.</p>\n<h2>Summary of My Solution</h2>\n<p>I built \"<strong>BPP-Protenix</strong>\", a model that integrates Base Pair Probability (BPP) features into an AlphaFold3-style architecture. The model design was inspired by the architecture of <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/668412\" target=\"_blank\">RNAPro</a> shared by <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>. My final pipeline is an ensemble of TBM, RNAPro, and BPP-Protenix (with plain Protenix as fallback for targets not suitable for BPP calculation). Although I dropped one place in the private LB, after completing the development of BPP-Protenix, I held 1st place on the public LB for the final two weeks of the competition. I guess almost no other participants modified the neural network architecture in this competition, so this appears to be a unique solution.</p>\n<p><strong>Submission slot allocation is follows</strong>:</p>\n<table>\n<thead>\n<tr>\n<th>slot 1</th>\n<th>slot 2</th>\n<th>slot 3</th>\n<th>slot 4</th>\n<th>slot5</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>TBM</td>\n<td>RNAPro</td>\n<td>RNAPro</td>\n<td><strong>BPP-Protenix</strong><br>(with plain Protenix fallback)</td>\n<td><strong>BPP-Protenix</strong><br>(with plain Protenix fallback)</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><p><strong>TBM</strong>: I used <a href=\"https://www.kaggle.com/code/kami1976/stanford-rna-3d-folding-part-2a18\" target=\"_blank\">the notebook</a> published by <a href=\"https://www.kaggle.com/kami1976\" target=\"_blank\">@kami1976</a> almost exactly as is. Since it appeared to offer good local scores and seed stability, I decided to borrow this one.</p></li>\n<li><p><strong>RNAPro</strong>: Using the template created by the above TBM as input data, I applied the <a href=\"https://www.kaggle.com/code/jaejohn/rnapro-inference-with-tbm\" target=\"_blank\">the pipeline</a> published by <a href=\"https://www.kaggle.com/jaejohn\" target=\"_blank\">@jaejohn</a>.</p></li>\n<li><p><strong>BPP-Protenix</strong>: This is the model that required the most time to develop for this competition. I'll explain it below.</p></li>\n</ul>\n<h2>Core Idea: BPP as Structural Prior</h2>\n<h3>Inspiration: Ribonanza Competition (2023)</h3>\n<p>The idea came from the <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding\" target=\"_blank\">Stanford Ribonanza RNA Folding</a> competition in 2023. <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> (one of the hosts of the Ribonanza competition) had published about <a href=\"https://academic.oup.com/bib/article/24/1/bbac581/6986359\" target=\"_blank\">RNAdegformer</a>, a transformer predicting mRNA degradation at nucleotide level. In that paper, <strong>BPP (Base Pair Probability) matrix was shown to be a very important feature</strong> for predicting RNA stability. Inspired by this, almost all top Ribonanza participants added BPP embeddings as bias terms in their Transformers.</p>\n<h3>From Stability Prediction to 3D Structure Prediction</h3>\n<p>The task of RNAdegformer and Ribonanza models was stability prediction, but RNA stability depends a lot on 3D structure. Furthermore, the base pair formation probability information obtained through secondary structure prediction provides proximity constraints on the 3D structure, thereby functioning as an effective inductive bias. Given that, providing BPP information to a 3D structure predictor seemed like a natural extension.</p>\n<h3>BPP Injection into Pairformer</h3>\n<p>The BPP matrix <code>[N, N]</code> naturally corresponds to the pair representation <code>z [N, N, c]</code> in the Pairformer trunk of AlphaFold3/Protenix architecture — both are pairwise matrices indexed by residue positions. By passing BPP through a simple linear layer (1 → c) and adding the resulting embedding to <code>z_init</code>, we can inject RNA structural knowledge directly into the model.</p>\n<pre><code>BPP [N, N]\n  → unsqueeze → [N, N, 1]\n  → LinearNoBias(1 → 128) → [N, N, 128]\n  → z_init += bpp_embedding\n</code></pre>\n<h2>BPP Calculator Selection: EternaFold vs ViennaRNA</h2>\n<p>I evaluated BPP quality against experimental 3D coordinates from the training data, checking if residue pairs with high BPP are actually close in 3D space. <a href=\"https://github.com/ViennaRNA/ViennaRNA\" target=\"_blank\">ViennaRNA</a> and <a href=\"https://github.com/WaymentSteeleLab/EternaFold\" target=\"_blank\">EternaFold</a>, two of the primary RNA secondary structure computation tools, were tested. These secondary structure prediction tools assume RNA monomers as the target of prediction. Additionally, calculating BPP for long sequences during a Kaggle session is difficult in inference. Trherefore, it was considered that the training should also be performed using shorter sequences. The following filter was applied to the approximately 5,000 sequences of Kaggle's official training data:</p>\n<p><strong>Filter conditions</strong></p>\n<ul>\n<li>RNA monomer (no multimer)</li>\n<li>Sequence length ≤ 1,000</li>\n</ul>\n<p>The BPP quality was evaluated using the <strong>1,532</strong> sequences remaining after applying this filter. Here, a C1'-C1' distance &lt; 12 Å was defined as \"contact.\" However, this judge was only conducted for C1' combinations separated by 4 or more residues (because such residues exist spatially close regardless of base pair formation).</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>ViennaRNA</th>\n<th>EternaFold</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Contact rate (BPP ≥ 0.7)</td>\n<td>81.4%</td>\n<td><strong>96.2%</strong></td>\n</tr>\n<tr>\n<td>Contact rate (BPP ≥ 0.9)</td>\n<td>86.6%</td>\n<td><strong>97.4%</strong></td>\n</tr>\n<tr>\n<td>Random baseline*</td>\n<td>3.6%</td>\n<td>3.6%</td>\n</tr>\n</tbody>\n</table>\n<p>*Random baseline: contact rate when residue pairs are selected randomly (no BPP information).</p>\n<p>EternaFold BPP was clearly more accurate. Also, it is easily deployable in Kaggle notebooks via the <a href=\"https://github.com/DasLab/arnie\" target=\"_blank\">arnie</a> library. So I chose EternaFold for feature engineering.</p>\n<hr>\n<h2>BPP-Protenix training</h2>\n<h3>Setup</h3>\n<ul>\n<li><strong>Pretrained checkpoint</strong>: <code>protenix_base_20250630_v1.0.0</code></li>\n<li><strong>Train data</strong>: 619 PDBs (RNA-only)<ul>\n<li>From the 1,532 RNA monomers (≤1000 nt) used for BPP quality evaluation mentioned above, PDBs with protein/DNA partners were excluded.</li></ul></li>\n<li><strong>Validation data</strong>: 33 PDBs<ul>\n<li>Combined Kaggle's official validation set with my own collected PDBs released between 2025-12-03 and 2026-02-18. After applying the same filter, 33 records remained.</li></ul></li>\n</ul>\n<h3>Hyperparameters</h3>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>diffusion_batch_size</td>\n<td>32</td>\n</tr>\n<tr>\n<td>train_crop_size</td>\n<td>550</td>\n</tr>\n<tr>\n<td>lr</td>\n<td>1e-4</td>\n</tr>\n<tr>\n<td>EMA decay</td>\n<td>0.995</td>\n</tr>\n<tr>\n<td>num_steps</td>\n<td>4960</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>BPP-Protenix Inference</h2>\n<h3>BPP-Available Target Filter &amp; Fallback to Plain Protenix</h3>\n<p>The following filters were applied to determine if the target was appropriate for BPP calculation by EternaFold:</p>\n<ul>\n<li>RNA monomer (no multimer)</li>\n<li>Total entity length ≤ 850 tokens</li>\n<li>No protein/DNA partners</li>\n</ul>\n<p>Targets not passing the filter were predicted by <strong>plain Protenix</strong> as fallback. Targets exceeding 850 tokens were cropped to 850 residues, with remaining coordinates zero-padded.</p>\n<h3>BPP-Protenix Standalone Performance</h3>\n<p>Before evaluating the ensemble, I assessed BPP-Protenix alone (all 5 slots filled with Protenix predictions). Inspired by RNAPro's TemplateEmbedder (Protenix v0.5), I tested three integration methods (A,B,C).</p>\n<ul>\n<li><strong>Method A</strong>: just a linear layer from 1 to 128 dimensions, then embedding is added to <code>z_init</code>.</li>\n<li><strong>Method B</strong>: In addition to the direct linear layer, I also converted BPP into a 20-dimensional one-hot vector by binning, embedded it with another linear layer, and added both features to z_init.</li>\n<li><strong>Method C</strong>: In addition to method B, uses a BppEmbedder. It is the similar design pattern as TemplateEmbedder in RNAPro. Applies a small 2-block Pairformer before the main 48-block Pairformer in each recycling cycle.</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Architecture</th>\n<th>Public LB (standalone*)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>A (Linear)</strong></td>\n<td>LinearNoBias(1→128), add BPP-emb to z_init</td>\n<td><strong>0.340</strong></td>\n</tr>\n<tr>\n<td>B (+ Binning)</td>\n<td>Method A + one_hot(20 bins) → Linear(20→128),  add two embs to z_init</td>\n<td>0.317</td>\n</tr>\n<tr>\n<td>C (+ BppEmbedder)</td>\n<td>Method B + BppEmbedder in recycling loop</td>\n<td>0.320</td>\n</tr>\n<tr>\n<td>Plain Protenix (control)</td>\n<td>-</td>\n<td>0.309</td>\n</tr>\n</tbody>\n</table>\n<p>*Standalone: all 5 slots filled with BPP-Protenix + plain Protenix fallback</p>\n<p>The standalone BPP-Protenix (Method A) scored <strong>0.34</strong> on public LB. This is comparable to TBM-only public notebooks (~0.35), demonstrating that <strong>NN-based predictions can match template-based approaches</strong>.</p>\n<p>Interesting finding: allowing protein/DNA partner targets in the BPP filter caused almost no change in public LB score. Possibly few such targets existed in the test set. Worth investigating the performance for RNA/protein assembly with late submissions.</p>\n<hr>\n<h2>Pseudoknot Analysis</h2>\n<p><strong>Concern</strong>: EternaFold does not predict pseudoknots. Could BPP-Protenix is not suitable for potential pseudoknot targets?</p>\n<p><strong>Post-hoc validation</strong>:</p>\n<ul>\n<li>Detected pseudoknots from ground-truth 3D coordinates using biotite (base pair detection + crossing pair check)</li>\n<li>Validation set (11 RNA-only targets): pseudoknots found in <strong>7 out of 11</strong> targets</li>\n<li>BPP-Protenix outperformed plain Protenix by <strong>~0.06 TM-score</strong> on average across all 11 targets</li>\n<li>Contrary to expectations, even within the pseudoknot group, BPP-Protenix showed higher TM-scores</li>\n</ul>\n<p><strong>Interpretation</strong>: BPP provides a \"pairing tendency hint\" as continuous values, not a complete secondary structure prediction. Even if pseudoknot base pairs are missed, non-pseudoknot BPP information is still valuable. And maybe, the 48-block Pairformer can absorb BPP inaccuracies.</p>\n<p>Based on the local validation result, I have decided not to take any special measures specifically targeting pseudoknots.</p>\n<hr>\n<h2>Ensemble Pipeline</h2>\n<h3>TBM (Slot 1) and RNAPro (Slot 2–3)</h3>\n<ul>\n<li>Used a public notebook shared early in the competition as-is</li>\n</ul>\n<h3>BPP-Protenix (Slot 4–5)</h3>\n<ul>\n<li>Targets passing the BPP-available filter → BPP-Protenix (seed=101, N_sample=5, top 2 by confidence score)</li>\n<li>Targets failing the filter → Plain Protenix fallback (same config, no BPP, Longer targets were cropped to 850 tokens and zero-padded)</li>\n</ul>\n<h3>Public/Private LB Results</h3>\n<table>\n<thead>\n<tr>\n<th>Configuration</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>TBM 2slot / RNAPro 3slot (baseline)</td>\n<td>~0.43</td>\n<td>—</td>\n</tr>\n<tr>\n<td>TBM 1slot / RNAPro 2slot / Protenix 2slot</td>\n<td>0.461</td>\n<td>0.475</td>\n</tr>\n<tr>\n<td><strong>TBM 1slot / RNAPro 2slot / BPP-Protenix 2slot</strong></td>\n<td><strong>0.504 (1st place)</strong></td>\n<td><strong>0.492 (2nd place)</strong></td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>BPP-Protenix yielded <strong>+0.04</strong> over plain Protenix and <strong>+0.07</strong> over the TBM+RNAPro baseline in the public LB score</li>\n<li>BPP-Protenix demonstrated strong accuracy on the private test set as well</li>\n</ul>\n<h2>Acknowledgments</h2>\n<p>Thank you to the competition hosts (Rhiju Das and team) for organizing this fascinating competition. Thanks also to the Kaggle community for sharing public notebooks (TBM, RNAPro) that formed the foundation of my ensemble.</p>",
      "rawMarkdown": "First, I would like to extend my deepest gratitude to the competition organizing team and Kaggle staff for hosting this fascinating competition. Having previously worked in experimental sciences including organic chemistry, I am well aware of the challenges in obtaining molecular 3D structural data (ground-truth structures). Considering this, I'm deeply impressed that they were able to hold this Part II competition just less than a year after the [Part I competition](https://www.kaggle.com/competitions/stanford-rna-3d-folding).\nDuring these approximately two months of trial-and-error work, I gained both meaningful experience and valuable learning.\n\n## Summary of My Solution\n\nI built \"**BPP-Protenix**\", a model that integrates Base Pair Probability (BPP) features into an AlphaFold3-style architecture. The model design was inspired by the architecture of [RNAPro](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/668412) shared by @theoviel. My final pipeline is an ensemble of TBM, RNAPro, and BPP-Protenix (with plain Protenix as fallback for targets not suitable for BPP calculation). Although I dropped one place in the private LB, after completing the development of BPP-Protenix, I held 1st place on the public LB for the final two weeks of the competition. I guess almost no other participants modified the neural network architecture in this competition, so this appears to be a unique solution.\n\n**Submission slot allocation is follows**:\n\n| slot 1 | slot 2 | slot 3 |                    slot 4                     |                     slot5                     |\n| :----: | :----: | :----: | :-------------------------------------------: | :-------------------------------------------: |\n|  TBM   | RNAPro | RNAPro | **BPP-Protenix**<br>(with plain Protenix fallback) | **BPP-Protenix**<br>(with plain Protenix fallback) |\n\n- **TBM**: I used [the notebook](https://www.kaggle.com/code/kami1976/stanford-rna-3d-folding-part-2a18) published by @kami1976 almost exactly as is. Since it appeared to offer good local scores and seed stability, I decided to borrow this one.\n\n- **RNAPro**: Using the template created by the above TBM as input data, I applied the [the pipeline](https://www.kaggle.com/code/jaejohn/rnapro-inference-with-tbm) published by @jaejohn.\n\n- **BPP-Protenix**: This is the model that required the most time to develop for this competition. I'll explain it below.\n\n\n## Core Idea: BPP as Structural Prior\n\n### Inspiration: Ribonanza Competition (2023)\n\nThe idea came from the [Stanford Ribonanza RNA Folding](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding) competition in 2023. @shujun717 (one of the hosts of the Ribonanza competition) had published about [RNAdegformer](https://academic.oup.com/bib/article/24/1/bbac581/6986359), a transformer predicting mRNA degradation at nucleotide level. In that paper, **BPP (Base Pair Probability) matrix was shown to be a very important feature** for predicting RNA stability. Inspired by this, almost all top Ribonanza participants added BPP embeddings as bias terms in their Transformers.\n\n### From Stability Prediction to 3D Structure Prediction\n\nThe task of RNAdegformer and Ribonanza models was stability prediction, but RNA stability depends a lot on 3D structure. Furthermore, the base pair formation probability information obtained through secondary structure prediction provides proximity constraints on the 3D structure, thereby functioning as an effective inductive bias. Given that, providing BPP information to a 3D structure predictor seemed like a natural extension.\n\n### BPP Injection into Pairformer\n\nThe BPP matrix `[N, N]` naturally corresponds to the pair representation `z [N, N, c]` in the Pairformer trunk of AlphaFold3/Protenix architecture — both are pairwise matrices indexed by residue positions. By passing BPP through a simple linear layer (1 → c) and adding the resulting embedding to `z_init`, we can inject RNA structural knowledge directly into the model.\n\n```\nBPP [N, N]\n  → unsqueeze → [N, N, 1]\n  → LinearNoBias(1 → 128) → [N, N, 128]\n  → z_init += bpp_embedding\n```\n\n## BPP Calculator Selection: EternaFold vs ViennaRNA\n\nI evaluated BPP quality against experimental 3D coordinates from the training data, checking if residue pairs with high BPP are actually close in 3D space. [ViennaRNA](https://github.com/ViennaRNA/ViennaRNA) and [EternaFold](https://github.com/WaymentSteeleLab/EternaFold), two of the primary RNA secondary structure computation tools, were tested. These secondary structure prediction tools assume RNA monomers as the target of prediction. Additionally, calculating BPP for long sequences during a Kaggle session is difficult in inference. Trherefore, it was considered that the training should also be performed using shorter sequences. The following filter was applied to the approximately 5,000 sequences of Kaggle's official training data:\n\n**Filter conditions**\n- RNA monomer (no multimer)\n- Sequence length ≤ 1,000\n\nThe BPP quality was evaluated using the **1,532** sequences remaining after applying this filter. Here, a C1'-C1' distance < 12 Å was defined as \"contact.\" However, this judge was only conducted for C1' combinations separated by 4 or more residues (because such residues exist spatially close regardless of base pair formation).\n\n|                          | ViennaRNA | EternaFold |\n| ------------------------ | --------- | ---------- |\n| Contact rate (BPP ≥ 0.7) | 81.4%     | **96.2%**  |\n| Contact rate (BPP ≥ 0.9) | 86.6%     | **97.4%**  |\n| Random baseline*         | 3.6%      | 3.6%       |\n\n*Random baseline: contact rate when residue pairs are selected randomly (no BPP information).\n\nEternaFold BPP was clearly more accurate. Also, it is easily deployable in Kaggle notebooks via the [arnie](https://github.com/DasLab/arnie) library. So I chose EternaFold for feature engineering.\n\n---\n\n## BPP-Protenix training\n\n### Setup\n\n- **Pretrained checkpoint**: `protenix_base_20250630_v1.0.0`\n- **Train data**: 619 PDBs (RNA-only)\n  - From the 1,532 RNA monomers (≤1000 nt) used for BPP quality evaluation mentioned above, PDBs with protein/DNA partners were excluded.\n- **Validation data**: 33 PDBs\n  - Combined Kaggle's official validation set with my own collected PDBs released between 2025-12-03 and 2026-02-18. After applying the same filter, 33 records remained.\n\n### Hyperparameters\n\n| Parameter            | Value                           |\n| -------------------- | ------------------------------- |\n| diffusion_batch_size | 32                              |\n| train_crop_size      | 550                             |\n| lr                   | 1e-4                            |\n| EMA decay            | 0.995                           |\n| num_steps            | 4960                            |\n\n---\n\n## BPP-Protenix Inference\n\n### BPP-Available Target Filter & Fallback to Plain Protenix\n\nThe following filters were applied to determine if the target was appropriate for BPP calculation by EternaFold:\n- RNA monomer (no multimer)\n- Total entity length ≤ 850 tokens\n- No protein/DNA partners\n\nTargets not passing the filter were predicted by **plain Protenix** as fallback. Targets exceeding 850 tokens were cropped to 850 residues, with remaining coordinates zero-padded.\n\n### BPP-Protenix Standalone Performance\n\nBefore evaluating the ensemble, I assessed BPP-Protenix alone (all 5 slots filled with Protenix predictions). Inspired by RNAPro's TemplateEmbedder (Protenix v0.5), I tested three integration methods (A,B,C).\n\n- **Method A**: just a linear layer from 1 to 128 dimensions, then embedding is added to `z_init`.\n- **Method B**: In addition to the direct linear layer, I also converted BPP into a 20-dimensional one-hot vector by binning, embedded it with another linear layer, and added both features to z_init.\n- **Method C**: In addition to method B, uses a BppEmbedder. It is the similar design pattern as TemplateEmbedder in RNAPro. Applies a small 2-block Pairformer before the main 48-block Pairformer in each recycling cycle.\n\n| Method                   | Architecture                                                 | Public LB (standalone*) |\n| ------------------------ | ------------------------------------------------------------ | ----------------------- |\n| **A (Linear)**           | LinearNoBias(1→128), add BPP-emb to z_init                 | **0.340**               |\n| B (+ Binning)            | Method A + one_hot(20 bins) → Linear(20→128),  add two embs to z_init | 0.317                   |\n| C (+ BppEmbedder)        | Method B + BppEmbedder in recycling loop                     | 0.320                   |\n| Plain Protenix (control) | -                                                            | 0.309                   |\n\n*Standalone: all 5 slots filled with BPP-Protenix + plain Protenix fallback\n\nThe standalone BPP-Protenix (Method A) scored **0.34** on public LB. This is comparable to TBM-only public notebooks (~0.35), demonstrating that **NN-based predictions can match template-based approaches**.\n\nInteresting finding: allowing protein/DNA partner targets in the BPP filter caused almost no change in public LB score. Possibly few such targets existed in the test set. Worth investigating the performance for RNA/protein assembly with late submissions.\n\n---\n\n## Pseudoknot Analysis\n\n**Concern**: EternaFold does not predict pseudoknots. Could BPP-Protenix is not suitable for potential pseudoknot targets?\n\n**Post-hoc validation**:\n- Detected pseudoknots from ground-truth 3D coordinates using biotite (base pair detection + crossing pair check)\n- Validation set (11 RNA-only targets): pseudoknots found in **7 out of 11** targets\n- BPP-Protenix outperformed plain Protenix by **~0.06 TM-score** on average across all 11 targets\n- Contrary to expectations, even within the pseudoknot group, BPP-Protenix showed higher TM-scores\n\n**Interpretation**: BPP provides a \"pairing tendency hint\" as continuous values, not a complete secondary structure prediction. Even if pseudoknot base pairs are missed, non-pseudoknot BPP information is still valuable. And maybe, the 48-block Pairformer can absorb BPP inaccuracies.\n\nBased on the local validation result, I have decided not to take any special measures specifically targeting pseudoknots.\n\n---\n\n## Ensemble Pipeline\n\n### TBM (Slot 1) and RNAPro (Slot 2–3)\n- Used a public notebook shared early in the competition as-is\n\n### BPP-Protenix (Slot 4–5)\n- Targets passing the BPP-available filter → BPP-Protenix (seed=101, N_sample=5, top 2 by confidence score)\n- Targets failing the filter → Plain Protenix fallback (same config, no BPP, Longer targets were cropped to 850 tokens and zero-padded)\n\n### Public/Private LB Results\n\n| Configuration                                     | Public LB             |Private LB            |\n| ------------------------------------------------- | --------------------- |--------------------- |\n| TBM 2slot / RNAPro 3slot (baseline)               | ~0.43                 |—                     |\n| TBM 1slot / RNAPro 2slot / Protenix 2slot         | 0.461                 |0.475                 |\n| **TBM 1slot / RNAPro 2slot / BPP-Protenix 2slot** | **0.504 (1st place)** |**0.492 (2nd place)** |\n\n- BPP-Protenix yielded **+0.04** over plain Protenix and **+0.07** over the TBM+RNAPro baseline in the public LB score\n- BPP-Protenix demonstrated strong accuracy on the private test set as well\n\n\n## Acknowledgments\n\nThank you to the competition hosts (Rhiju Das and team) for organizing this fascinating competition. Thanks also to the Kaggle community for sharing public notebooks (TBM, RNAPro) that formed the foundation of my ensemble.",
      "votes": 21
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3441423": "First, I would like to extend my deepest gratitude to the competition organizing team and Kaggle staff for hosting this fascinating competition. Having previously worked in experimental sciences including organic chemistry, I am well aware of the challenges in obtaining molecular 3D structural data (ground-truth structures). Considering this, I'm deeply impressed that they were able to hold this Part II competition just less than a year after the [Part I competition](https://www.kaggle.com/competitions/stanford-rna-3d-folding).\nDuring these approximately two months of trial-and-error work, I gained both meaningful experience and valuable learning.\n\n## Summary of My Solution\n\nI built \"**BPP-Protenix**\", a model that integrates Base Pair Probability (BPP) features into an AlphaFold3-style architecture. The model design was inspired by the architecture of [RNAPro](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/668412) shared by @theoviel. My final pipeline is an ensemble of TBM, RNAPro, and BPP-Protenix (with plain Protenix as fallback for targets not suitable for BPP calculation). Although I dropped one place in the private LB, after completing the development of BPP-Protenix, I held 1st place on the public LB for the final two weeks of the competition. I guess almost no other participants modified the neural network architecture in this competition, so this appears to be a unique solution.\n\n**Submission slot allocation is follows**:\n\n| slot 1 | slot 2 | slot 3 |                    slot 4                     |                     slot5                     |\n| :----: | :----: | :----: | :-------------------------------------------: | :-------------------------------------------: |\n|  TBM   | RNAPro | RNAPro | **BPP-Protenix**<br>(with plain Protenix fallback) | **BPP-Protenix**<br>(with plain Protenix fallback) |\n\n- **TBM**: I used [the notebook](https://www.kaggle.com/code/kami1976/stanford-rna-3d-folding-part-2a18) published by @kami1976 almost exactly as is. Since it appeared to offer good local scores and seed stability, I decided to borrow this one.\n\n- **RNAPro**: Using the template created by the above TBM as input data, I applied the [the pipeline](https://www.kaggle.com/code/jaejohn/rnapro-inference-with-tbm) published by @jaejohn.\n\n- **BPP-Protenix**: This is the model that required the most time to develop for this competition. I'll explain it below.\n\n\n## Core Idea: BPP as Structural Prior\n\n### Inspiration: Ribonanza Competition (2023)\n\nThe idea came from the [Stanford Ribonanza RNA Folding](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding) competition in 2023. @shujun717 (one of the hosts of the Ribonanza competition) had published about [RNAdegformer](https://academic.oup.com/bib/article/24/1/bbac581/6986359), a transformer predicting mRNA degradation at nucleotide level. In that paper, **BPP (Base Pair Probability) matrix was shown to be a very important feature** for predicting RNA stability. Inspired by this, almost all top Ribonanza participants added BPP embeddings as bias terms in their Transformers.\n\n### From Stability Prediction to 3D Structure Prediction\n\nThe task of RNAdegformer and Ribonanza models was stability prediction, but RNA stability depends a lot on 3D structure. Furthermore, the base pair formation probability information obtained through secondary structure prediction provides proximity constraints on the 3D structure, thereby functioning as an effective inductive bias. Given that, providing BPP information to a 3D structure predictor seemed like a natural extension.\n\n### BPP Injection into Pairformer\n\nThe BPP matrix `[N, N]` naturally corresponds to the pair representation `z [N, N, c]` in the Pairformer trunk of AlphaFold3/Protenix architecture — both are pairwise matrices indexed by residue positions. By passing BPP through a simple linear layer (1 → c) and adding the resulting embedding to `z_init`, we can inject RNA structural knowledge directly into the model.\n\n```\nBPP [N, N]\n  → unsqueeze → [N, N, 1]\n  → LinearNoBias(1 → 128) → [N, N, 128]\n  → z_init += bpp_embedding\n```\n\n## BPP Calculator Selection: EternaFold vs ViennaRNA\n\nI evaluated BPP quality against experimental 3D coordinates from the training data, checking if residue pairs with high BPP are actually close in 3D space. [ViennaRNA](https://github.com/ViennaRNA/ViennaRNA) and [EternaFold](https://github.com/WaymentSteeleLab/EternaFold), two of the primary RNA secondary structure computation tools, were tested. These secondary structure prediction tools assume RNA monomers as the target of prediction. Additionally, calculating BPP for long sequences during a Kaggle session is difficult in inference. Trherefore, it was considered that the training should also be performed using shorter sequences. The following filter was applied to the approximately 5,000 sequences of Kaggle's official training data:\n\n**Filter conditions**\n- RNA monomer (no multimer)\n- Sequence length ≤ 1,000\n\nThe BPP quality was evaluated using the **1,532** sequences remaining after applying this filter. Here, a C1'-C1' distance < 12 Å was defined as \"contact.\" However, this judge was only conducted for C1' combinations separated by 4 or more residues (because such residues exist spatially close regardless of base pair formation).\n\n|                          | ViennaRNA | EternaFold |\n| ------------------------ | --------- | ---------- |\n| Contact rate (BPP ≥ 0.7) | 81.4%     | **96.2%**  |\n| Contact rate (BPP ≥ 0.9) | 86.6%     | **97.4%**  |\n| Random baseline*         | 3.6%      | 3.6%       |\n\n*Random baseline: contact rate when residue pairs are selected randomly (no BPP information).\n\nEternaFold BPP was clearly more accurate. Also, it is easily deployable in Kaggle notebooks via the [arnie](https://github.com/DasLab/arnie) library. So I chose EternaFold for feature engineering.\n\n---\n\n## BPP-Protenix training\n\n### Setup\n\n- **Pretrained checkpoint**: `protenix_base_20250630_v1.0.0`\n- **Train data**: 619 PDBs (RNA-only)\n  - From the 1,532 RNA monomers (≤1000 nt) used for BPP quality evaluation mentioned above, PDBs with protein/DNA partners were excluded.\n- **Validation data**: 33 PDBs\n  - Combined Kaggle's official validation set with my own collected PDBs released between 2025-12-03 and 2026-02-18. After applying the same filter, 33 records remained.\n\n### Hyperparameters\n\n| Parameter            | Value                           |\n| -------------------- | ------------------------------- |\n| diffusion_batch_size | 32                              |\n| train_crop_size      | 550                             |\n| lr                   | 1e-4                            |\n| EMA decay            | 0.995                           |\n| num_steps            | 4960                            |\n\n---\n\n## BPP-Protenix Inference\n\n### BPP-Available Target Filter & Fallback to Plain Protenix\n\nThe following filters were applied to determine if the target was appropriate for BPP calculation by EternaFold:\n- RNA monomer (no multimer)\n- Total entity length ≤ 850 tokens\n- No protein/DNA partners\n\nTargets not passing the filter were predicted by **plain Protenix** as fallback. Targets exceeding 850 tokens were cropped to 850 residues, with remaining coordinates zero-padded.\n\n### BPP-Protenix Standalone Performance\n\nBefore evaluating the ensemble, I assessed BPP-Protenix alone (all 5 slots filled with Protenix predictions). Inspired by RNAPro's TemplateEmbedder (Protenix v0.5), I tested three integration methods (A,B,C).\n\n- **Method A**: just a linear layer from 1 to 128 dimensions, then embedding is added to `z_init`.\n- **Method B**: In addition to the direct linear layer, I also converted BPP into a 20-dimensional one-hot vector by binning, embedded it with another linear layer, and added both features to z_init.\n- **Method C**: In addition to method B, uses a BppEmbedder. It is the similar design pattern as TemplateEmbedder in RNAPro. Applies a small 2-block Pairformer before the main 48-block Pairformer in each recycling cycle.\n\n| Method                   | Architecture                                                 | Public LB (standalone*) |\n| ------------------------ | ------------------------------------------------------------ | ----------------------- |\n| **A (Linear)**           | LinearNoBias(1→128), add BPP-emb to z_init                 | **0.340**               |\n| B (+ Binning)            | Method A + one_hot(20 bins) → Linear(20→128),  add two embs to z_init | 0.317                   |\n| C (+ BppEmbedder)        | Method B + BppEmbedder in recycling loop                     | 0.320                   |\n| Plain Protenix (control) | -                                                            | 0.309                   |\n\n*Standalone: all 5 slots filled with BPP-Protenix + plain Protenix fallback\n\nThe standalone BPP-Protenix (Method A) scored **0.34** on public LB. This is comparable to TBM-only public notebooks (~0.35), demonstrating that **NN-based predictions can match template-based approaches**.\n\nInteresting finding: allowing protein/DNA partner targets in the BPP filter caused almost no change in public LB score. Possibly few such targets existed in the test set. Worth investigating the performance for RNA/protein assembly with late submissions.\n\n---\n\n## Pseudoknot Analysis\n\n**Concern**: EternaFold does not predict pseudoknots. Could BPP-Protenix is not suitable for potential pseudoknot targets?\n\n**Post-hoc validation**:\n- Detected pseudoknots from ground-truth 3D coordinates using biotite (base pair detection + crossing pair check)\n- Validation set (11 RNA-only targets): pseudoknots found in **7 out of 11** targets\n- BPP-Protenix outperformed plain Protenix by **~0.06 TM-score** on average across all 11 targets\n- Contrary to expectations, even within the pseudoknot group, BPP-Protenix showed higher TM-scores\n\n**Interpretation**: BPP provides a \"pairing tendency hint\" as continuous values, not a complete secondary structure prediction. Even if pseudoknot base pairs are missed, non-pseudoknot BPP information is still valuable. And maybe, the 48-block Pairformer can absorb BPP inaccuracies.\n\nBased on the local validation result, I have decided not to take any special measures specifically targeting pseudoknots.\n\n---\n\n## Ensemble Pipeline\n\n### TBM (Slot 1) and RNAPro (Slot 2–3)\n- Used a public notebook shared early in the competition as-is\n\n### BPP-Protenix (Slot 4–5)\n- Targets passing the BPP-available filter → BPP-Protenix (seed=101, N_sample=5, top 2 by confidence score)\n- Targets failing the filter → Plain Protenix fallback (same config, no BPP, Longer targets were cropped to 850 tokens and zero-padded)\n\n### Public/Private LB Results\n\n| Configuration                                     | Public LB             |Private LB            |\n| ------------------------------------------------- | --------------------- |--------------------- |\n| TBM 2slot / RNAPro 3slot (baseline)               | ~0.43                 |—                     |\n| TBM 1slot / RNAPro 2slot / Protenix 2slot         | 0.461                 |0.475                 |\n| **TBM 1slot / RNAPro 2slot / BPP-Protenix 2slot** | **0.504 (1st place)** |**0.492 (2nd place)** |\n\n- BPP-Protenix yielded **+0.04** over plain Protenix and **+0.07** over the TBM+RNAPro baseline in the public LB score\n- BPP-Protenix demonstrated strong accuracy on the private test set as well\n\n\n## Acknowledgments\n\nThank you to the competition hosts (Rhiju Das and team) for organizing this fascinating competition. Thanks also to the Kaggle community for sharing public notebooks (TBM, RNAPro) that formed the foundation of my ensemble."
  }
}