{
  "id": 686820,
  "title": "10th Place Solution",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/686820",
  "author_name": "Kh0a",
  "post_date": "2026-04-01T09:41:18.709000",
  "votes": 13,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First word: thanks the host for organize this competition and kagglers for their contribution. This solution absolutely would not have been possible without the incredible contributions of the Kaggle community, so a huge thank you to everyone who shared their insights and baselines. </p>\n<hr>\n<p>I definitely expected a big shake-up in this competition, but I honestly didn't expect to rank this high! </p>\n<p>The toughest part of this competition for me was figuring out how to properly validate and diversify the 5 required predictions. To finalize my submission, I meticulously tracked both my local validation score (<code>validation_sequences.csv</code>) and the Public LB score. </p>\n<p>Here is a breakdown of the core implementation that led to this result.</p>\n<h3>📌 TL;DR</h3>\n<p>My final pipeline relies on an ensemble of:</p>\n<ul>\n<li><strong>TBM</strong> + <strong>Protenix</strong> + <strong>RNAPro</strong></li>\n<li><strong>Chunking</strong> long sequences specifically for Protenix.</li>\n</ul>\n<p>The 5 submission spots were allocated as follows:</p>\n<ul>\n<li><strong>2 spots for TBM</strong></li>\n<li><strong>1 spot for Protenix</strong></li>\n<li><strong>2 spots for RNAPro</strong></li>\n</ul>\n<p>My two final selected submissions used the exact same logic, but with one key difference for diversity: one had <code>USE_RNA_MSA=false</code>, and the other had <code>USE_RNA_MSA=true</code>.</p>\n<hr>\n<h2>Section 1: Inference Baseline</h2>\n<h3>1. TBM (Template-Based Modeling)</h3>\n<p>The TBM logic builds upon several great public notebooks. Merging the validation and training sets to act as a larger template pool for searching when submitting only. train set only when running local validation</p>\n<p><strong>The Pipeline:</strong></p>\n<ol>\n<li><strong>Search:</strong> Scan up to <code>top_n = 30</code> candidates using a global <code>PairwiseAligner</code>. Rank these by their normalized alignment score: $$\\text{norm_score} = \\frac{\\text{aln.score}}{2 \\times \\min(\\text{len_query}, \\text{len_template})}$$  and percent identity. Filter the results by <code>MIN_SIMILARITY</code> and <code>MIN_PERCENT_IDENTITY</code>.</li>\n<li><strong>Adapt:</strong> Map template coordinates to the query using alignment-aware copying alongside interpolation/extrapolation (<code>adapt_template_to_query</code>). This ensures chain segments and stoichiometry are respected so multi-chain templates map correctly.</li>\n<li><strong>Produce Predictions (TBM = 2):</strong><ul>\n<li><strong>Slot 1:</strong> Adapted template (straight copy).</li>\n<li><strong>Slot 2:</strong> Adapted template + small perturbation Gaussian noise with </li></ul></li>\n</ol>\n<p>$$\\sigma = \\max(0.01, (0.40 - \\text{sim}) \\times 0.06)$$ .If similarity is very low, a geometric perturbation (hinge/jitter) is used instead.</p>\n<ol>\n<li><strong>Refine:</strong> Apply soft geometric smoothing via <code>adaptive_rna_constraints()</code> to keep local geometry realistic. </li>\n</ol>\n<p><strong>Key Parameters:</strong></p>\n<ul>\n<li><code>top_n=30</code></li>\n<li><code>adaptive_rna_constraints</code> passes=2</li>\n<li><code>apply_hinge</code> uses $$\\text{deg} \\approx 22$$</li>\n</ul>\n<h3>2. Protenix</h3>\n<ul>\n<li><strong>Quality Control:</strong> Extract the C1' coordinates robustly. Reject collapsed or degenerate outputs (e.g., near-zero inter-residue distances), and pad or repeat valid samples as needed so every target ends up with the required number of predictions.</li>\n<li><strong>Long-sequence Strategy:</strong> Long RNAs were split into overlapping segments. Each segment was inferred independently, and the full-length structure was rebuilt by aligning overlapping regions and smoothly blending coordinates to ensure a continuous backbone. I used <code>MAX_SEQ_LEN ≈ 512</code> and <code>CHUNK_OVERLAP ≈ 64</code>. </li>\n</ul>\n<p>There were only 2 samples in the validation set that required chunking (<code>9MME</code>, <code>9ZCC</code>), but the chunking approach showed consistent local improvement over non-chunking:</p>\n<table>\n<thead>\n<tr>\n<th>Target</th>\n<th>No Chunking</th>\n<th>With Chunking</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>9MME</strong></td>\n<td>~0.1376</td>\n<td><strong>0.1855</strong></td>\n</tr>\n<tr>\n<td><strong>9ZCC</strong></td>\n<td>~0.2320</td>\n<td><strong>0.2781</strong></td>\n</tr>\n</tbody>\n</table>\n<h3>3. RNAPro</h3>\n<p>I passed the TBM-derived predictions directly to RNAPro as structural templates for refinement. RNAPro was run with MSA + RibonanzaNet2 embeddings, 10 recycling cycles, and 200 diffusion steps.</p>\n<hr>\n<h2>Section 2: How I Came Up With This Solution</h2>\n<p>This was the hardest part. I ran inference on the 28 validation samples repeatedly with different settings. Here is what I discovered:</p>\n<ul>\n<li>Protenix predictions are non-deterministic, but generating multiple Protenix predictions for a single runtime didn't yield much helpful diversity.</li>\n<li>RNAPro performed worse on the Public LB but was performing <em>very well</em> locally. </li>\n<li>Setting RNA_MSA to True greatly enhance local validation score by <code>~0.03</code> but LB scores bellow 0.4 when combining with TBM.</li>\n</ul>\n<p>Here is a log of my tracked runs:</p>\n<table>\n<thead>\n<tr>\n<th>Model / Strategy</th>\n<th>Local Validation</th>\n<th>Public LB</th>\n<th>Private LB</th>\n<th>Final Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2 TBM - 1 Protenix - 2 RNAPro <br> <em>(with RNA_MSA False for Protenix)</em></td>\n<td>-</td>\n<td>[0.399;0.424]</td>\n<td>[0.565;0.586]</td>\n<td>0.474</td>\n</tr>\n<tr>\n<td>2 TBM - 1 Protenix - 2 RNAPro <br> <em>(with RNA_MSA True for Protenix)</em></td>\n<td>-</td>\n<td>[0.400;0.414]</td>\n<td>[0.578;0.588]</td>\n<td></td>\n</tr>\n<tr>\n<td>2 TBM - 2 Protenix - 1 RNAPro <br> <em>(with RNA_MSA True for Protenix)</em></td>\n<td>-</td>\n<td>[0.395;0.397]</td>\n<td>0.578</td>\n<td>0.471</td>\n</tr>\n<tr>\n<td>2 TBM - 2 Protenix - 1 RNAPro <br> <em>(with chunking long sequences for Protenix)</em></td>\n<td>0.45</td>\n<td>[0.399;0.424]</td>\n<td>[0.565;0.586]</td>\n<td></td>\n</tr>\n<tr>\n<td>3 TBM - 2 Protenix <br> <em>(With RNA_MSA on)</em></td>\n<td>0.47</td>\n<td>0.401</td>\n<td>0.540</td>\n<td></td>\n</tr>\n<tr>\n<td>TBM + reserve 2 spot for Protenix</td>\n<td>0.43</td>\n<td>[0.409;0.419]</td>\n<td>[0.560;0.590]</td>\n<td></td>\n</tr>\n<tr>\n<td>TBM + reserve 1 spot for Protenix</td>\n<td>0.44</td>\n<td>[0.407;0.419]</td>\n<td>[0.572;0.586]</td>\n<td></td>\n</tr>\n<tr>\n<td>main TBM + Protenix to fill leftover spot <br> <em>(my public notebook)</em></td>\n<td>0.42</td>\n<td>[0.396;0.429]</td>\n<td>[0.503;0.537]</td>\n<td></td>\n</tr>\n<tr>\n<td>RNAPro (seq &lt; 1000) + fill with TBM</td>\n<td>0.47</td>\n<td>0.357</td>\n<td>0.489</td>\n<td></td>\n</tr>\n<tr>\n<td>pure Protenix</td>\n<td></td>\n<td>0.249</td>\n<td>0.503</td>\n<td></td>\n</tr>\n<tr>\n<td>pure TBM</td>\n<td>0.36</td>\n<td>0.368</td>\n<td>0.458</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<p><strong>The Turning Point:</strong>\nThe first completed Protenix inference <a href=\"https://www.kaggle.com/code/llkh0a/stanford-rna-3d-folding-part-2-protenix-tbm\" target=\"_blank\">notebook</a> (which I published) had a critical bottleneck: it prevented making Protenix predictions entirely if 5 TBM predictions passed the threshold. </p>\n<p>After fixing this to ensure Protenix samples were explicitly generated, my local validation improved by <code>0.02</code>. However, there was no significant Public LB improvement. My hypothesis was that the Public LB was heavily dominated by the TBM approach. And many later public notebooks that included Protenix didn't fix this bottleneck either. </p>\n<hr>\n<h2>Section 3: What Didn't Work (Or What I Gave Up On)</h2>\n<ul>\n<li><strong>Reranking Protenix by Model Confidence:</strong> Creating a large pool of Protenix samples and reranking them based on the model's own confidence scores just cost way too much GPU time for too little return.</li>\n<li><strong>Finetuning Protenix:</strong> This was slightly out of my current comprehension level. I did try to preprocess and cache a small subset of data, but I ultimately ran out of time to actually get a training run working.</li>\n<li><strong>Using Templates for Protenix:</strong> I already set up on my public notebook to use template successfully, but I observed absolutely no difference on the Public LB and my local validation scores.</li>\n</ul>\n<hr>\n<h2>References:</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/amirrezaaleyasin/openrnafold-starter3\" target=\"_blank\">https://www.kaggle.com/code/amirrezaaleyasin/openrnafold-starter3</a></li>\n<li><a href=\"https://www.kaggle.com/datasets/qiweiyin/protenix-v1-adjusted\" target=\"_blank\">https://www.kaggle.com/datasets/qiweiyin/protenix-v1-adjusted</a></li>\n<li><a href=\"https://www.kaggle.com/code/alexxanderlarko/protenix-v1\" target=\"_blank\">https://www.kaggle.com/code/alexxanderlarko/protenix-v1</a></li>\n<li><a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">https://github.com/bytedance/Protenix</a></li>\n<li><a href=\"https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference\" target=\"_blank\">https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference</a></li>\n</ul>\n<hr>\n<p>Thanks again to everyone, and congratulations to all the winners!</p>",
  "messages": [
    {
      "id": 3433147,
      "postDate": "2026-04-01T09:41:18.710Z",
      "content": "<p>First word: thanks the host for organize this competition and kagglers for their contribution. This solution absolutely would not have been possible without the incredible contributions of the Kaggle community, so a huge thank you to everyone who shared their insights and baselines. </p>\n<hr>\n<p>I definitely expected a big shake-up in this competition, but I honestly didn't expect to rank this high! </p>\n<p>The toughest part of this competition for me was figuring out how to properly validate and diversify the 5 required predictions. To finalize my submission, I meticulously tracked both my local validation score (<code>validation_sequences.csv</code>) and the Public LB score. </p>\n<p>Here is a breakdown of the core implementation that led to this result.</p>\n<h3>📌 TL;DR</h3>\n<p>My final pipeline relies on an ensemble of:</p>\n<ul>\n<li><strong>TBM</strong> + <strong>Protenix</strong> + <strong>RNAPro</strong></li>\n<li><strong>Chunking</strong> long sequences specifically for Protenix.</li>\n</ul>\n<p>The 5 submission spots were allocated as follows:</p>\n<ul>\n<li><strong>2 spots for TBM</strong></li>\n<li><strong>1 spot for Protenix</strong></li>\n<li><strong>2 spots for RNAPro</strong></li>\n</ul>\n<p>My two final selected submissions used the exact same logic, but with one key difference for diversity: one had <code>USE_RNA_MSA=false</code>, and the other had <code>USE_RNA_MSA=true</code>.</p>\n<hr>\n<h2>Section 1: Inference Baseline</h2>\n<h3>1. TBM (Template-Based Modeling)</h3>\n<p>The TBM logic builds upon several great public notebooks. Merging the validation and training sets to act as a larger template pool for searching when submitting only. train set only when running local validation</p>\n<p><strong>The Pipeline:</strong></p>\n<ol>\n<li><strong>Search:</strong> Scan up to <code>top_n = 30</code> candidates using a global <code>PairwiseAligner</code>. Rank these by their normalized alignment score: $$\\text{norm_score} = \\frac{\\text{aln.score}}{2 \\times \\min(\\text{len_query}, \\text{len_template})}$$  and percent identity. Filter the results by <code>MIN_SIMILARITY</code> and <code>MIN_PERCENT_IDENTITY</code>.</li>\n<li><strong>Adapt:</strong> Map template coordinates to the query using alignment-aware copying alongside interpolation/extrapolation (<code>adapt_template_to_query</code>). This ensures chain segments and stoichiometry are respected so multi-chain templates map correctly.</li>\n<li><strong>Produce Predictions (TBM = 2):</strong><ul>\n<li><strong>Slot 1:</strong> Adapted template (straight copy).</li>\n<li><strong>Slot 2:</strong> Adapted template + small perturbation Gaussian noise with </li></ul></li>\n</ol>\n<p>$$\\sigma = \\max(0.01, (0.40 - \\text{sim}) \\times 0.06)$$ .If similarity is very low, a geometric perturbation (hinge/jitter) is used instead.</p>\n<ol>\n<li><strong>Refine:</strong> Apply soft geometric smoothing via <code>adaptive_rna_constraints()</code> to keep local geometry realistic. </li>\n</ol>\n<p><strong>Key Parameters:</strong></p>\n<ul>\n<li><code>top_n=30</code></li>\n<li><code>adaptive_rna_constraints</code> passes=2</li>\n<li><code>apply_hinge</code> uses $$\\text{deg} \\approx 22$$</li>\n</ul>\n<h3>2. Protenix</h3>\n<ul>\n<li><strong>Quality Control:</strong> Extract the C1' coordinates robustly. Reject collapsed or degenerate outputs (e.g., near-zero inter-residue distances), and pad or repeat valid samples as needed so every target ends up with the required number of predictions.</li>\n<li><strong>Long-sequence Strategy:</strong> Long RNAs were split into overlapping segments. Each segment was inferred independently, and the full-length structure was rebuilt by aligning overlapping regions and smoothly blending coordinates to ensure a continuous backbone. I used <code>MAX_SEQ_LEN ≈ 512</code> and <code>CHUNK_OVERLAP ≈ 64</code>. </li>\n</ul>\n<p>There were only 2 samples in the validation set that required chunking (<code>9MME</code>, <code>9ZCC</code>), but the chunking approach showed consistent local improvement over non-chunking:</p>\n<table>\n<thead>\n<tr>\n<th>Target</th>\n<th>No Chunking</th>\n<th>With Chunking</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>9MME</strong></td>\n<td>~0.1376</td>\n<td><strong>0.1855</strong></td>\n</tr>\n<tr>\n<td><strong>9ZCC</strong></td>\n<td>~0.2320</td>\n<td><strong>0.2781</strong></td>\n</tr>\n</tbody>\n</table>\n<h3>3. RNAPro</h3>\n<p>I passed the TBM-derived predictions directly to RNAPro as structural templates for refinement. RNAPro was run with MSA + RibonanzaNet2 embeddings, 10 recycling cycles, and 200 diffusion steps.</p>\n<hr>\n<h2>Section 2: How I Came Up With This Solution</h2>\n<p>This was the hardest part. I ran inference on the 28 validation samples repeatedly with different settings. Here is what I discovered:</p>\n<ul>\n<li>Protenix predictions are non-deterministic, but generating multiple Protenix predictions for a single runtime didn't yield much helpful diversity.</li>\n<li>RNAPro performed worse on the Public LB but was performing <em>very well</em> locally. </li>\n<li>Setting RNA_MSA to True greatly enhance local validation score by <code>~0.03</code> but LB scores bellow 0.4 when combining with TBM.</li>\n</ul>\n<p>Here is a log of my tracked runs:</p>\n<table>\n<thead>\n<tr>\n<th>Model / Strategy</th>\n<th>Local Validation</th>\n<th>Public LB</th>\n<th>Private LB</th>\n<th>Final Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2 TBM - 1 Protenix - 2 RNAPro <br> <em>(with RNA_MSA False for Protenix)</em></td>\n<td>-</td>\n<td>[0.399;0.424]</td>\n<td>[0.565;0.586]</td>\n<td>0.474</td>\n</tr>\n<tr>\n<td>2 TBM - 1 Protenix - 2 RNAPro <br> <em>(with RNA_MSA True for Protenix)</em></td>\n<td>-</td>\n<td>[0.400;0.414]</td>\n<td>[0.578;0.588]</td>\n<td></td>\n</tr>\n<tr>\n<td>2 TBM - 2 Protenix - 1 RNAPro <br> <em>(with RNA_MSA True for Protenix)</em></td>\n<td>-</td>\n<td>[0.395;0.397]</td>\n<td>0.578</td>\n<td>0.471</td>\n</tr>\n<tr>\n<td>2 TBM - 2 Protenix - 1 RNAPro <br> <em>(with chunking long sequences for Protenix)</em></td>\n<td>0.45</td>\n<td>[0.399;0.424]</td>\n<td>[0.565;0.586]</td>\n<td></td>\n</tr>\n<tr>\n<td>3 TBM - 2 Protenix <br> <em>(With RNA_MSA on)</em></td>\n<td>0.47</td>\n<td>0.401</td>\n<td>0.540</td>\n<td></td>\n</tr>\n<tr>\n<td>TBM + reserve 2 spot for Protenix</td>\n<td>0.43</td>\n<td>[0.409;0.419]</td>\n<td>[0.560;0.590]</td>\n<td></td>\n</tr>\n<tr>\n<td>TBM + reserve 1 spot for Protenix</td>\n<td>0.44</td>\n<td>[0.407;0.419]</td>\n<td>[0.572;0.586]</td>\n<td></td>\n</tr>\n<tr>\n<td>main TBM + Protenix to fill leftover spot <br> <em>(my public notebook)</em></td>\n<td>0.42</td>\n<td>[0.396;0.429]</td>\n<td>[0.503;0.537]</td>\n<td></td>\n</tr>\n<tr>\n<td>RNAPro (seq &lt; 1000) + fill with TBM</td>\n<td>0.47</td>\n<td>0.357</td>\n<td>0.489</td>\n<td></td>\n</tr>\n<tr>\n<td>pure Protenix</td>\n<td></td>\n<td>0.249</td>\n<td>0.503</td>\n<td></td>\n</tr>\n<tr>\n<td>pure TBM</td>\n<td>0.36</td>\n<td>0.368</td>\n<td>0.458</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<p><strong>The Turning Point:</strong>\nThe first completed Protenix inference <a href=\"https://www.kaggle.com/code/llkh0a/stanford-rna-3d-folding-part-2-protenix-tbm\" target=\"_blank\">notebook</a> (which I published) had a critical bottleneck: it prevented making Protenix predictions entirely if 5 TBM predictions passed the threshold. </p>\n<p>After fixing this to ensure Protenix samples were explicitly generated, my local validation improved by <code>0.02</code>. However, there was no significant Public LB improvement. My hypothesis was that the Public LB was heavily dominated by the TBM approach. And many later public notebooks that included Protenix didn't fix this bottleneck either. </p>\n<hr>\n<h2>Section 3: What Didn't Work (Or What I Gave Up On)</h2>\n<ul>\n<li><strong>Reranking Protenix by Model Confidence:</strong> Creating a large pool of Protenix samples and reranking them based on the model's own confidence scores just cost way too much GPU time for too little return.</li>\n<li><strong>Finetuning Protenix:</strong> This was slightly out of my current comprehension level. I did try to preprocess and cache a small subset of data, but I ultimately ran out of time to actually get a training run working.</li>\n<li><strong>Using Templates for Protenix:</strong> I already set up on my public notebook to use template successfully, but I observed absolutely no difference on the Public LB and my local validation scores.</li>\n</ul>\n<hr>\n<h2>References:</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/amirrezaaleyasin/openrnafold-starter3\" target=\"_blank\">https://www.kaggle.com/code/amirrezaaleyasin/openrnafold-starter3</a></li>\n<li><a href=\"https://www.kaggle.com/datasets/qiweiyin/protenix-v1-adjusted\" target=\"_blank\">https://www.kaggle.com/datasets/qiweiyin/protenix-v1-adjusted</a></li>\n<li><a href=\"https://www.kaggle.com/code/alexxanderlarko/protenix-v1\" target=\"_blank\">https://www.kaggle.com/code/alexxanderlarko/protenix-v1</a></li>\n<li><a href=\"https://github.com/bytedance/Protenix\" target=\"_blank\">https://github.com/bytedance/Protenix</a></li>\n<li><a href=\"https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference\" target=\"_blank\">https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference</a></li>\n</ul>\n<hr>\n<p>Thanks again to everyone, and congratulations to all the winners!</p>",
      "rawMarkdown": "First word: thanks the host for organize this competition and kagglers for their contribution. This solution absolutely would not have been possible without the incredible contributions of the Kaggle community, so a huge thank you to everyone who shared their insights and baselines. \n\n---\nI definitely expected a big shake-up in this competition, but I honestly didn't expect to rank this high! \n\nThe toughest part of this competition for me was figuring out how to properly validate and diversify the 5 required predictions. To finalize my submission, I meticulously tracked both my local validation score (`validation_sequences.csv`) and the Public LB score. \n\nHere is a breakdown of the core implementation that led to this result.\n\n### 📌 TL;DR\nMy final pipeline relies on an ensemble of:\n* **TBM** + **Protenix** + **RNAPro**\n* **Chunking** long sequences specifically for Protenix.\n\nThe 5 submission spots were allocated as follows:\n* **2 spots for TBM**\n* **1 spot for Protenix**\n* **2 spots for RNAPro**\n\nMy two final selected submissions used the exact same logic, but with one key difference for diversity: one had `USE_RNA_MSA=false`, and the other had `USE_RNA_MSA=true`.\n\n---\n\n## Section 1: Inference Baseline\n\n### 1. TBM (Template-Based Modeling)\nThe TBM logic builds upon several great public notebooks. Merging the validation and training sets to act as a larger template pool for searching when submitting only. train set only when running local validation\n\n**The Pipeline:**\n1. **Search:** Scan up to `top_n = 30` candidates using a global `PairwiseAligner`. Rank these by their normalized alignment score: $$\\text{norm\\_score} = \\frac{\\text{aln.score}}{2 \\times \\min(\\text{len\\_query}, \\text{len\\_template})}$$  and percent identity. Filter the results by `MIN_SIMILARITY` and `MIN_PERCENT_IDENTITY`.\n2. **Adapt:** Map template coordinates to the query using alignment-aware copying alongside interpolation/extrapolation (`adapt_template_to_query`). This ensures chain segments and stoichiometry are respected so multi-chain templates map correctly.\n3. **Produce Predictions (TBM = 2):**\n   * **Slot 1:** Adapted template (straight copy).\n   * **Slot 2:** Adapted template + small perturbation Gaussian noise with \n\n$$\\sigma = \\max(0.01, (0.40 - \\text{sim}) \\times 0.06)$$ .If similarity is very low, a geometric perturbation (hinge/jitter) is used instead.\n4. **Refine:** Apply soft geometric smoothing via `adaptive_rna_constraints()` to keep local geometry realistic. \n\n**Key Parameters:**\n* `top_n=30`\n* `adaptive_rna_constraints` passes=2\n* `apply_hinge` uses $$\\text{deg} \\approx 22$$\n\n### 2. Protenix\n* **Quality Control:** Extract the C1' coordinates robustly. Reject collapsed or degenerate outputs (e.g., near-zero inter-residue distances), and pad or repeat valid samples as needed so every target ends up with the required number of predictions.\n* **Long-sequence Strategy:** Long RNAs were split into overlapping segments. Each segment was inferred independently, and the full-length structure was rebuilt by aligning overlapping regions and smoothly blending coordinates to ensure a continuous backbone. I used `MAX_SEQ_LEN ≈ 512` and `CHUNK_OVERLAP ≈ 64`. \n\nThere were only 2 samples in the validation set that required chunking (`9MME`, `9ZCC`), but the chunking approach showed consistent local improvement over non-chunking:\n\n| Target | No Chunking | With Chunking |\n| :--- | :--- | :--- |\n| **9MME** | ~0.1376 | **0.1855** |\n| **9ZCC** | ~0.2320 | **0.2781** |\n\n### 3. RNAPro\nI passed the TBM-derived predictions directly to RNAPro as structural templates for refinement. RNAPro was run with MSA + RibonanzaNet2 embeddings, 10 recycling cycles, and 200 diffusion steps.\n\n---\n\n## Section 2: How I Came Up With This Solution\n\nThis was the hardest part. I ran inference on the 28 validation samples repeatedly with different settings. Here is what I discovered:\n* Protenix predictions are non-deterministic, but generating multiple Protenix predictions for a single runtime didn't yield much helpful diversity.\n* RNAPro performed worse on the Public LB but was performing *very well* locally. \n* Setting RNA_MSA to True greatly enhance local validation score by `~0.03` but LB scores bellow 0.4 when combining with TBM.\n\nHere is a log of my tracked runs:\n\n| Model / Strategy | Local Validation | Public LB | Private LB | Final Private LB |\n| :--- | :--- | :--- | :--- | :--- |\n| 2 TBM - 1 Protenix - 2 RNAPro <br> *(with RNA_MSA False for Protenix)* | - | [0.399;0.424] | [0.565;0.586] | 0.474 |\n| 2 TBM - 1 Protenix - 2 RNAPro <br> *(with RNA_MSA True for Protenix)* | - | [0.400;0.414] | [0.578;0.588] | |\n| 2 TBM - 2 Protenix - 1 RNAPro <br> *(with RNA_MSA True for Protenix)* | - | [0.395;0.397] | 0.578 | 0.471 |\n| 2 TBM - 2 Protenix - 1 RNAPro <br> *(with chunking long sequences for Protenix)*| 0.45 | [0.399;0.424] | [0.565;0.586] | |\n| 3 TBM - 2 Protenix <br> *(With RNA_MSA on)* | 0.47 | 0.401 | 0.540 | |\n| TBM + reserve 2 spot for Protenix | 0.43 | [0.409;0.419] | [0.560;0.590] | |\n| TBM + reserve 1 spot for Protenix | 0.44 | [0.407;0.419] | [0.572;0.586] | |\n| main TBM + Protenix to fill leftover spot <br> *(my public notebook)* | 0.42 | [0.396;0.429] | [0.503;0.537] | |\n| RNAPro (seq < 1000) + fill with TBM | 0.47 | 0.357 | 0.489 | |\n| pure Protenix | | 0.249 | 0.503 | |\n| pure TBM | 0.36 | 0.368 | 0.458 | |\n\n**The Turning Point:**\nThe first completed Protenix inference [notebook](https://www.kaggle.com/code/llkh0a/stanford-rna-3d-folding-part-2-protenix-tbm) (which I published) had a critical bottleneck: it prevented making Protenix predictions entirely if 5 TBM predictions passed the threshold. \n\nAfter fixing this to ensure Protenix samples were explicitly generated, my local validation improved by `0.02`. However, there was no significant Public LB improvement. My hypothesis was that the Public LB was heavily dominated by the TBM approach. And many later public notebooks that included Protenix didn't fix this bottleneck either. \n\n---\n\n## Section 3: What Didn't Work (Or What I Gave Up On)\n\n* **Reranking Protenix by Model Confidence:** Creating a large pool of Protenix samples and reranking them based on the model's own confidence scores just cost way too much GPU time for too little return.\n* **Finetuning Protenix:** This was slightly out of my current comprehension level. I did try to preprocess and cache a small subset of data, but I ultimately ran out of time to actually get a training run working.\n* **Using Templates for Protenix:** I already set up on my public notebook to use template successfully, but I observed absolutely no difference on the Public LB and my local validation scores.\n\n---\n\n## References:\n* https://www.kaggle.com/code/amirrezaaleyasin/openrnafold-starter3\n* https://www.kaggle.com/datasets/qiweiyin/protenix-v1-adjusted\n* https://www.kaggle.com/code/alexxanderlarko/protenix-v1\n* https://github.com/bytedance/Protenix\n* https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference\n\n---\n\nThanks again to everyone, and congratulations to all the winners!\n\n",
      "votes": 13
    },
    {
      "id": 3433694,
      "postDate": "2026-04-02T00:25:01.310Z",
      "content": "<p>It's fascinating that your focus on <code>validation_sequences.csv</code> led to strong performance.  One of the cases in there, <strong>9MME</strong>, has 8 RNA chains and there are existing templates for the entire complex. Did you treat multi-chain cases specially?</p>",
      "rawMarkdown": "It's fascinating that your focus on `validation_sequences.csv` led to strong performance.  One of the cases in there, **9MME**, has 8 RNA chains and there are existing templates for the entire complex. Did you treat multi-chain cases specially?",
      "votes": 2,
      "replies": [
        {
          "id": 3433713,
          "postDate": "2026-04-02T01:57:12.723Z",
          "content": "<p>Thank you for hosting this competition!</p>\n<p>For my final solution, I did not treat multi-chain cases specially. I only split sequences longer than 512 into multiple chunks before feeding them into Protenix. For 9MME (length 4640), chunking improved the score from 0.13 to 0.18 for Protenix prediction, but TBM yielded a far superior score of 0.8493 for that case.</p>",
          "rawMarkdown": "Thank you for hosting this competition!\n\nFor my final solution, I did not treat multi-chain cases specially. I only split sequences longer than 512 into multiple chunks before feeding them into Protenix. For 9MME (length 4640), chunking improved the score from 0.13 to 0.18 for Protenix prediction, but TBM yielded a far superior score of 0.8493 for that case.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3433160,
      "postDate": "2026-04-01T10:13:08.967Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3433694,
      "author_name": "Rhiju Das",
      "author_url": "",
      "post_date": "2026-04-02T00:25:01.310000",
      "content": "<p>It's fascinating that your focus on <code>validation_sequences.csv</code> led to strong performance.  One of the cases in there, <strong>9MME</strong>, has 8 RNA chains and there are existing templates for the entire complex. Did you treat multi-chain cases specially?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3433713,
          "author_name": "Kh0a",
          "author_url": "",
          "post_date": "2026-04-02T01:57:12.723000",
          "content": "<p>Thank you for hosting this competition!</p>\n<p>For my final solution, I did not treat multi-chain cases specially. I only split sequences longer than 512 into multiple chunks before feeding them into Protenix. For 9MME (length 4640), chunking improved the score from 0.13 to 0.18 for Protenix prediction, but TBM yielded a far superior score of 0.8493 for that case.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3433160,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-04-01T10:13:08.967000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3433147": "First word: thanks the host for organize this competition and kagglers for their contribution. This solution absolutely would not have been possible without the incredible contributions of the Kaggle community, so a huge thank you to everyone who shared their insights and baselines. \n\n---\nI definitely expected a big shake-up in this competition, but I honestly didn't expect to rank this high! \n\nThe toughest part of this competition for me was figuring out how to properly validate and diversify the 5 required predictions. To finalize my submission, I meticulously tracked both my local validation score (`validation_sequences.csv`) and the Public LB score. \n\nHere is a breakdown of the core implementation that led to this result.\n\n### 📌 TL;DR\nMy final pipeline relies on an ensemble of:\n* **TBM** + **Protenix** + **RNAPro**\n* **Chunking** long sequences specifically for Protenix.\n\nThe 5 submission spots were allocated as follows:\n* **2 spots for TBM**\n* **1 spot for Protenix**\n* **2 spots for RNAPro**\n\nMy two final selected submissions used the exact same logic, but with one key difference for diversity: one had `USE_RNA_MSA=false`, and the other had `USE_RNA_MSA=true`.\n\n---\n\n## Section 1: Inference Baseline\n\n### 1. TBM (Template-Based Modeling)\nThe TBM logic builds upon several great public notebooks. Merging the validation and training sets to act as a larger template pool for searching when submitting only. train set only when running local validation\n\n**The Pipeline:**\n1. **Search:** Scan up to `top_n = 30` candidates using a global `PairwiseAligner`. Rank these by their normalized alignment score: $$\\text{norm\\_score} = \\frac{\\text{aln.score}}{2 \\times \\min(\\text{len\\_query}, \\text{len\\_template})}$$  and percent identity. Filter the results by `MIN_SIMILARITY` and `MIN_PERCENT_IDENTITY`.\n2. **Adapt:** Map template coordinates to the query using alignment-aware copying alongside interpolation/extrapolation (`adapt_template_to_query`). This ensures chain segments and stoichiometry are respected so multi-chain templates map correctly.\n3. **Produce Predictions (TBM = 2):**\n   * **Slot 1:** Adapted template (straight copy).\n   * **Slot 2:** Adapted template + small perturbation Gaussian noise with \n\n$$\\sigma = \\max(0.01, (0.40 - \\text{sim}) \\times 0.06)$$ .If similarity is very low, a geometric perturbation (hinge/jitter) is used instead.\n4. **Refine:** Apply soft geometric smoothing via `adaptive_rna_constraints()` to keep local geometry realistic. \n\n**Key Parameters:**\n* `top_n=30`\n* `adaptive_rna_constraints` passes=2\n* `apply_hinge` uses $$\\text{deg} \\approx 22$$\n\n### 2. Protenix\n* **Quality Control:** Extract the C1' coordinates robustly. Reject collapsed or degenerate outputs (e.g., near-zero inter-residue distances), and pad or repeat valid samples as needed so every target ends up with the required number of predictions.\n* **Long-sequence Strategy:** Long RNAs were split into overlapping segments. Each segment was inferred independently, and the full-length structure was rebuilt by aligning overlapping regions and smoothly blending coordinates to ensure a continuous backbone. I used `MAX_SEQ_LEN ≈ 512` and `CHUNK_OVERLAP ≈ 64`. \n\nThere were only 2 samples in the validation set that required chunking (`9MME`, `9ZCC`), but the chunking approach showed consistent local improvement over non-chunking:\n\n| Target | No Chunking | With Chunking |\n| :--- | :--- | :--- |\n| **9MME** | ~0.1376 | **0.1855** |\n| **9ZCC** | ~0.2320 | **0.2781** |\n\n### 3. RNAPro\nI passed the TBM-derived predictions directly to RNAPro as structural templates for refinement. RNAPro was run with MSA + RibonanzaNet2 embeddings, 10 recycling cycles, and 200 diffusion steps.\n\n---\n\n## Section 2: How I Came Up With This Solution\n\nThis was the hardest part. I ran inference on the 28 validation samples repeatedly with different settings. Here is what I discovered:\n* Protenix predictions are non-deterministic, but generating multiple Protenix predictions for a single runtime didn't yield much helpful diversity.\n* RNAPro performed worse on the Public LB but was performing *very well* locally. \n* Setting RNA_MSA to True greatly enhance local validation score by `~0.03` but LB scores bellow 0.4 when combining with TBM.\n\nHere is a log of my tracked runs:\n\n| Model / Strategy | Local Validation | Public LB | Private LB | Final Private LB |\n| :--- | :--- | :--- | :--- | :--- |\n| 2 TBM - 1 Protenix - 2 RNAPro <br> *(with RNA_MSA False for Protenix)* | - | [0.399;0.424] | [0.565;0.586] | 0.474 |\n| 2 TBM - 1 Protenix - 2 RNAPro <br> *(with RNA_MSA True for Protenix)* | - | [0.400;0.414] | [0.578;0.588] | |\n| 2 TBM - 2 Protenix - 1 RNAPro <br> *(with RNA_MSA True for Protenix)* | - | [0.395;0.397] | 0.578 | 0.471 |\n| 2 TBM - 2 Protenix - 1 RNAPro <br> *(with chunking long sequences for Protenix)*| 0.45 | [0.399;0.424] | [0.565;0.586] | |\n| 3 TBM - 2 Protenix <br> *(With RNA_MSA on)* | 0.47 | 0.401 | 0.540 | |\n| TBM + reserve 2 spot for Protenix | 0.43 | [0.409;0.419] | [0.560;0.590] | |\n| TBM + reserve 1 spot for Protenix | 0.44 | [0.407;0.419] | [0.572;0.586] | |\n| main TBM + Protenix to fill leftover spot <br> *(my public notebook)* | 0.42 | [0.396;0.429] | [0.503;0.537] | |\n| RNAPro (seq < 1000) + fill with TBM | 0.47 | 0.357 | 0.489 | |\n| pure Protenix | | 0.249 | 0.503 | |\n| pure TBM | 0.36 | 0.368 | 0.458 | |\n\n**The Turning Point:**\nThe first completed Protenix inference [notebook](https://www.kaggle.com/code/llkh0a/stanford-rna-3d-folding-part-2-protenix-tbm) (which I published) had a critical bottleneck: it prevented making Protenix predictions entirely if 5 TBM predictions passed the threshold. \n\nAfter fixing this to ensure Protenix samples were explicitly generated, my local validation improved by `0.02`. However, there was no significant Public LB improvement. My hypothesis was that the Public LB was heavily dominated by the TBM approach. And many later public notebooks that included Protenix didn't fix this bottleneck either. \n\n---\n\n## Section 3: What Didn't Work (Or What I Gave Up On)\n\n* **Reranking Protenix by Model Confidence:** Creating a large pool of Protenix samples and reranking them based on the model's own confidence scores just cost way too much GPU time for too little return.\n* **Finetuning Protenix:** This was slightly out of my current comprehension level. I did try to preprocess and cache a small subset of data, but I ultimately ran out of time to actually get a training run working.\n* **Using Templates for Protenix:** I already set up on my public notebook to use template successfully, but I observed absolutely no difference on the Public LB and my local validation scores.\n\n---\n\n## References:\n* https://www.kaggle.com/code/amirrezaaleyasin/openrnafold-starter3\n* https://www.kaggle.com/datasets/qiweiyin/protenix-v1-adjusted\n* https://www.kaggle.com/code/alexxanderlarko/protenix-v1\n* https://github.com/bytedance/Protenix\n* https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference\n\n---\n\nThanks again to everyone, and congratulations to all the winners!\n\n",
    "3433694": "It's fascinating that your focus on `validation_sequences.csv` led to strong performance.  One of the cases in there, **9MME**, has 8 RNA chains and there are existing templates for the entire complex. Did you treat multi-chain cases specially?",
    "3433160": ""
  }
}