{
  "id": 688036,
  "title": "5th Place Solution — msa and protenix and TBM",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/688036",
  "author_name": "huyang111",
  "post_date": "2026-04-04T12:30:47.415000",
  "votes": 15,
  "comment_count": 1,
  "views": 0,
  "content": "<h1>Acknowledgments</h1>\n<p>As a graduate student in bioinformatics at University of Science and Technology of China, \"I have learned a lot from this competition, and it is a great help to me and my future. I am grateful for the help of my teammates and believe this was a great collaboration. Most of the open-source notebooks in this competition were built based on Qiwei's dataset; he has extraordinary code insight. Flop provided great ideas for our code improvement in the later stages. Other teammates also cleared up our confusion with professional RNA domain knowledge and provided computational power support.\nCongrats to all the winners and participants!Here's a quick summary of what I did.</p>\n<h1>conclusion</h1>\n<h2>Phase 1: Template-Based Modeling</h2>\n<p>We searched the combined training and validation set using Biopython's PairwiseAligner with RNA-tuned gap penalties. Atomic coordinates were transferred to the query using a geometry-aware gap-filling procedure: missing residues are interpolated along RNA-like helical trajectories (C1'–C1' step ~5.9 Å) rather than placed at zero or left as NaN. Diversity transforms — hinge rotations, chain jitter, smooth wiggle — are applied slot-by-slot from the same template source, followed by RNA geometry constraint refinement (bond-length correction + Laplacian smoothing + self-avoidance).\nTo avoid slot redundancy, we apply TM-score-based greedy diversity selection: from all TBM candidates, we greedily pick structures that maximize a quality-diversity trade-off score, ensuring the two TBM slots are structurally distinct rather than near-duplicates of the top hit.\nSlot allocation is length-adaptive:\nShort sequences (≤ 512 nt): max 2 TBM slots, 3 reserved for Protenix\nLong sequences (&gt; 512 nt): TBM-first, up to 5 TBM slots if sufficient templates exist</p>\n<h2>Phase 2 Core Insight: MSA Depth</h2>\n<p>The mainstream open-source approach treats MSA as a binary switch — either full MSA or no MSA — . I argue this framing discards useful structure in the MSA quality signal.\nMy analysis of the 28 test targets revealed a clear depth distribution:\nDepth ≥ 700  : 15 targets (54%) — strong evolutionary signal\nDepth 10–699 :  7 targets (25%) — moderate, noise risk\nDepth ≤ 9    :  6 targets (21%) — MSA nearly uninformative</p>\n<p>I ran ablations across thresholds (100 / 300  / 700 / 1500) and found depth ≥ 700 to be the empirically optimal cutoff: below it, distant homologs inject alignment noise that degrades diffusion quality; above it, evolutionary co-variation provides reliable structural constraints.\nThis leads to our three-mode Protenix inference design:</p>\n<pre><code>Slot 3: msa700 sample 1  — official MSA, depth ≥ 700\n                            high-fidelity evolutionary constraint\nSlot 4: msa700 sample 2  — same MSA policy, independent diffusion\n                            diversity from stochastic diffusion,\n                            NOT from MSA noise\nSlot 5: msa_full         — relaxed threshold, depth ≥ 1\n                            intentional diversity hedge:\n                            noisy-but-different MSA features\n                            only deployed AFTER slots 3&amp;4 are secured\n</code></pre>\n<p>The key distinction from the binary on/off approach: slots 3 and 4 share identical high-quality MSA features; diversity between them comes purely from independent diffusion sampling. Slot 5 is a deliberate hedge that accepts some MSA noise in exchange for a structurally distinct prediction.\nFor long sequences (&gt; 512 nt), only msa700 is used with N_sample=3 — three independent diffusion samples in a single forward pass, sharing the expensive feature computation while diversifying outputs.\nFor targets requiring multi-chunk inference, we replaced the standard linear weight ramp with a quintic smoothstep function $$w(t) = 6t^5 - 15t^4 + 10t^3, \\quad t \\in [0, 1]$$. While linear blending results in first-derivative discontinuities at chunk boundaries, our approach guarantees  c2(curvature) continuity. Combined with Kabsch-based reference alignment in the 128-nt overlap regions, this eliminates geometric 'kinks' and ensures the global global topology remains smooth and physically valid.</p>\n<h1>Phase 3</h1>\n<p>Standard Protenix deployments produce different outputs across runs due to non-deterministic CUDA operations, MSA subsampling, and floating-point ordering effects. We enforced full reproducibility through three mechanisms:\n1.Disabling non-deterministic attention backends — FlashAttention and memory-efficient attention are both switched off, forcing deterministic standard attention\n2.Disabling MC Dropout (mc_dropout_apply_rate=0.0) to remove inference-time stochastic regularization\n3.Forcing torch-native LayerNorm and triangle attention kernels instead of cuDNN/cuEquivariance backends, which have non-deterministic implementations\nThis guarantees that any improvement in our pipeline comes from algorithmic changes, not RNG variation — a critical property for reliable ablation studies in a competition setting.</p>\n<pre><code>test_sequences.csv\n       |\n       v\n+----------------------+\n|       Phase 1        |   PairwiseAligner, top_n=30\n|         TBM          |   geometry-aware gap filling\n|   Diversity Selection|   TM-score greedy selection\n+----------------------+\n       |\n       |  short (≤512 nt): 2 TBM slots\n       |  long  (&gt;512 nt): up to 5 TBM slots (TBM-first)\n       v\n+----------------------+\n|       Phase 2        |   Dual-GPU Protenix\n|      Protenix        |\n|   3 inference modes: |   slot 3: msa700 sample 1\n|   msa700_1           |           (depth≥700, deterministic)\n|   msa700_2           |   slot 4: msa700 sample 2\n|   msa_full           |           (independent diffusion)\n|                      |   slot 5: msa_full\n|                      |           (depth≥1, diversity hedge)\n+----------------------+\n       |\n       | long seqs (&gt;512 nt):\n       | chunk-level greedy bin-packing across 2 GPUs\n       | → Kabsch SVD align + smoothstep stitch\n       v\n  submission.csv\n  (5 slots per target, fully deterministic &amp; reproducible)\n</code></pre>\n<p>Reference</p>\n<p>[1]<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/qiweiyin/protenix-v1-inference-2026</a></p>\n<p>[2]<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/nihilisticneuralnet/0-409-stanford-rna-folding-2-protenix-template</a></p>\n<p>[3]<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/alexxanderlarko/protenix-v1</a></p>",
  "messages": [
    {
      "id": 3435576,
      "postDate": "2026-04-04T12:30:47.417Z",
      "content": "<h1>Acknowledgments</h1>\n<p>As a graduate student in bioinformatics at University of Science and Technology of China, \"I have learned a lot from this competition, and it is a great help to me and my future. I am grateful for the help of my teammates and believe this was a great collaboration. Most of the open-source notebooks in this competition were built based on Qiwei's dataset; he has extraordinary code insight. Flop provided great ideas for our code improvement in the later stages. Other teammates also cleared up our confusion with professional RNA domain knowledge and provided computational power support.\nCongrats to all the winners and participants!Here's a quick summary of what I did.</p>\n<h1>conclusion</h1>\n<h2>Phase 1: Template-Based Modeling</h2>\n<p>We searched the combined training and validation set using Biopython's PairwiseAligner with RNA-tuned gap penalties. Atomic coordinates were transferred to the query using a geometry-aware gap-filling procedure: missing residues are interpolated along RNA-like helical trajectories (C1'–C1' step ~5.9 Å) rather than placed at zero or left as NaN. Diversity transforms — hinge rotations, chain jitter, smooth wiggle — are applied slot-by-slot from the same template source, followed by RNA geometry constraint refinement (bond-length correction + Laplacian smoothing + self-avoidance).\nTo avoid slot redundancy, we apply TM-score-based greedy diversity selection: from all TBM candidates, we greedily pick structures that maximize a quality-diversity trade-off score, ensuring the two TBM slots are structurally distinct rather than near-duplicates of the top hit.\nSlot allocation is length-adaptive:\nShort sequences (≤ 512 nt): max 2 TBM slots, 3 reserved for Protenix\nLong sequences (&gt; 512 nt): TBM-first, up to 5 TBM slots if sufficient templates exist</p>\n<h2>Phase 2 Core Insight: MSA Depth</h2>\n<p>The mainstream open-source approach treats MSA as a binary switch — either full MSA or no MSA — . I argue this framing discards useful structure in the MSA quality signal.\nMy analysis of the 28 test targets revealed a clear depth distribution:\nDepth ≥ 700  : 15 targets (54%) — strong evolutionary signal\nDepth 10–699 :  7 targets (25%) — moderate, noise risk\nDepth ≤ 9    :  6 targets (21%) — MSA nearly uninformative</p>\n<p>I ran ablations across thresholds (100 / 300  / 700 / 1500) and found depth ≥ 700 to be the empirically optimal cutoff: below it, distant homologs inject alignment noise that degrades diffusion quality; above it, evolutionary co-variation provides reliable structural constraints.\nThis leads to our three-mode Protenix inference design:</p>\n<pre><code>Slot 3: msa700 sample 1  — official MSA, depth ≥ 700\n                            high-fidelity evolutionary constraint\nSlot 4: msa700 sample 2  — same MSA policy, independent diffusion\n                            diversity from stochastic diffusion,\n                            NOT from MSA noise\nSlot 5: msa_full         — relaxed threshold, depth ≥ 1\n                            intentional diversity hedge:\n                            noisy-but-different MSA features\n                            only deployed AFTER slots 3&amp;4 are secured\n</code></pre>\n<p>The key distinction from the binary on/off approach: slots 3 and 4 share identical high-quality MSA features; diversity between them comes purely from independent diffusion sampling. Slot 5 is a deliberate hedge that accepts some MSA noise in exchange for a structurally distinct prediction.\nFor long sequences (&gt; 512 nt), only msa700 is used with N_sample=3 — three independent diffusion samples in a single forward pass, sharing the expensive feature computation while diversifying outputs.\nFor targets requiring multi-chunk inference, we replaced the standard linear weight ramp with a quintic smoothstep function $$w(t) = 6t^5 - 15t^4 + 10t^3, \\quad t \\in [0, 1]$$. While linear blending results in first-derivative discontinuities at chunk boundaries, our approach guarantees  c2(curvature) continuity. Combined with Kabsch-based reference alignment in the 128-nt overlap regions, this eliminates geometric 'kinks' and ensures the global global topology remains smooth and physically valid.</p>\n<h1>Phase 3</h1>\n<p>Standard Protenix deployments produce different outputs across runs due to non-deterministic CUDA operations, MSA subsampling, and floating-point ordering effects. We enforced full reproducibility through three mechanisms:\n1.Disabling non-deterministic attention backends — FlashAttention and memory-efficient attention are both switched off, forcing deterministic standard attention\n2.Disabling MC Dropout (mc_dropout_apply_rate=0.0) to remove inference-time stochastic regularization\n3.Forcing torch-native LayerNorm and triangle attention kernels instead of cuDNN/cuEquivariance backends, which have non-deterministic implementations\nThis guarantees that any improvement in our pipeline comes from algorithmic changes, not RNG variation — a critical property for reliable ablation studies in a competition setting.</p>\n<pre><code>test_sequences.csv\n       |\n       v\n+----------------------+\n|       Phase 1        |   PairwiseAligner, top_n=30\n|         TBM          |   geometry-aware gap filling\n|   Diversity Selection|   TM-score greedy selection\n+----------------------+\n       |\n       |  short (≤512 nt): 2 TBM slots\n       |  long  (&gt;512 nt): up to 5 TBM slots (TBM-first)\n       v\n+----------------------+\n|       Phase 2        |   Dual-GPU Protenix\n|      Protenix        |\n|   3 inference modes: |   slot 3: msa700 sample 1\n|   msa700_1           |           (depth≥700, deterministic)\n|   msa700_2           |   slot 4: msa700 sample 2\n|   msa_full           |           (independent diffusion)\n|                      |   slot 5: msa_full\n|                      |           (depth≥1, diversity hedge)\n+----------------------+\n       |\n       | long seqs (&gt;512 nt):\n       | chunk-level greedy bin-packing across 2 GPUs\n       | → Kabsch SVD align + smoothstep stitch\n       v\n  submission.csv\n  (5 slots per target, fully deterministic &amp; reproducible)\n</code></pre>\n<p>Reference</p>\n<p>[1]<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/qiweiyin/protenix-v1-inference-2026</a></p>\n<p>[2]<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/nihilisticneuralnet/0-409-stanford-rna-folding-2-protenix-template</a></p>\n<p>[3]<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/alexxanderlarko/protenix-v1</a></p>",
      "rawMarkdown": "# Acknowledgments\nAs a graduate student in bioinformatics at University of Science and Technology of China, \"I have learned a lot from this competition, and it is a great help to me and my future. I am grateful for the help of my teammates and believe this was a great collaboration. Most of the open-source notebooks in this competition were built based on Qiwei's dataset; he has extraordinary code insight. Flop provided great ideas for our code improvement in the later stages. Other teammates also cleared up our confusion with professional RNA domain knowledge and provided computational power support.\nCongrats to all the winners and participants!Here's a quick summary of what I did.\n\n# conclusion\n## Phase 1: Template-Based Modeling\nWe searched the combined training and validation set using Biopython's PairwiseAligner with RNA-tuned gap penalties. Atomic coordinates were transferred to the query using a geometry-aware gap-filling procedure: missing residues are interpolated along RNA-like helical trajectories (C1'–C1' step ~5.9 Å) rather than placed at zero or left as NaN. Diversity transforms — hinge rotations, chain jitter, smooth wiggle — are applied slot-by-slot from the same template source, followed by RNA geometry constraint refinement (bond-length correction + Laplacian smoothing + self-avoidance).\nTo avoid slot redundancy, we apply TM-score-based greedy diversity selection: from all TBM candidates, we greedily pick structures that maximize a quality-diversity trade-off score, ensuring the two TBM slots are structurally distinct rather than near-duplicates of the top hit.\nSlot allocation is length-adaptive:\nShort sequences (≤ 512 nt): max 2 TBM slots, 3 reserved for Protenix\nLong sequences (> 512 nt): TBM-first, up to 5 TBM slots if sufficient templates exist\n\n## Phase 2 Core Insight: MSA Depth\nThe mainstream open-source approach treats MSA as a binary switch — either full MSA or no MSA — . I argue this framing discards useful structure in the MSA quality signal.\nMy analysis of the 28 test targets revealed a clear depth distribution:\nDepth ≥ 700  : 15 targets (54%) — strong evolutionary signal\nDepth 10–699 :  7 targets (25%) — moderate, noise risk\nDepth ≤ 9    :  6 targets (21%) — MSA nearly uninformative\n\nI ran ablations across thresholds (100 / 300  / 700 / 1500) and found depth ≥ 700 to be the empirically optimal cutoff: below it, distant homologs inject alignment noise that degrades diffusion quality; above it, evolutionary co-variation provides reliable structural constraints.\nThis leads to our three-mode Protenix inference design:\n~~~~\nSlot 3: msa700 sample 1  — official MSA, depth ≥ 700\n                            high-fidelity evolutionary constraint\nSlot 4: msa700 sample 2  — same MSA policy, independent diffusion\n                            diversity from stochastic diffusion,\n                            NOT from MSA noise\nSlot 5: msa_full         — relaxed threshold, depth ≥ 1\n                            intentional diversity hedge:\n                            noisy-but-different MSA features\n                            only deployed AFTER slots 3&4 are secured\n~~~~\nThe key distinction from the binary on/off approach: slots 3 and 4 share identical high-quality MSA features; diversity between them comes purely from independent diffusion sampling. Slot 5 is a deliberate hedge that accepts some MSA noise in exchange for a structurally distinct prediction.\nFor long sequences (> 512 nt), only msa700 is used with N_sample=3 — three independent diffusion samples in a single forward pass, sharing the expensive feature computation while diversifying outputs.\nFor targets requiring multi-chunk inference, we replaced the standard linear weight ramp with a quintic smoothstep function $$w(t) = 6t^5 - 15t^4 + 10t^3, \\quad t \\in [0, 1]$$. While linear blending results in first-derivative discontinuities at chunk boundaries, our approach guarantees  c2(curvature) continuity. Combined with Kabsch-based reference alignment in the 128-nt overlap regions, this eliminates geometric 'kinks' and ensures the global global topology remains smooth and physically valid.\n\n# Phase 3\nStandard Protenix deployments produce different outputs across runs due to non-deterministic CUDA operations, MSA subsampling, and floating-point ordering effects. We enforced full reproducibility through three mechanisms:\n1.Disabling non-deterministic attention backends — FlashAttention and memory-efficient attention are both switched off, forcing deterministic standard attention\n2.Disabling MC Dropout (mc_dropout_apply_rate=0.0) to remove inference-time stochastic regularization\n3.Forcing torch-native LayerNorm and triangle attention kernels instead of cuDNN/cuEquivariance backends, which have non-deterministic implementations\nThis guarantees that any improvement in our pipeline comes from algorithmic changes, not RNG variation — a critical property for reliable ablation studies in a competition setting.\n\n~~~~\ntest_sequences.csv\n       |\n       v\n+----------------------+\n|       Phase 1        |   PairwiseAligner, top_n=30\n|         TBM          |   geometry-aware gap filling\n|   Diversity Selection|   TM-score greedy selection\n+----------------------+\n       |\n       |  short (≤512 nt): 2 TBM slots\n       |  long  (>512 nt): up to 5 TBM slots (TBM-first)\n       v\n+----------------------+\n|       Phase 2        |   Dual-GPU Protenix\n|      Protenix        |\n|   3 inference modes: |   slot 3: msa700 sample 1\n|   msa700_1           |           (depth≥700, deterministic)\n|   msa700_2           |   slot 4: msa700 sample 2\n|   msa_full           |           (independent diffusion)\n|                      |   slot 5: msa_full\n|                      |           (depth≥1, diversity hedge)\n+----------------------+\n       |\n       | long seqs (>512 nt):\n       | chunk-level greedy bin-packing across 2 GPUs\n       | → Kabsch SVD align + smoothstep stitch\n       v\n  submission.csv\n  (5 slots per target, fully deterministic & reproducible)\n~~~~\n\n\n\n\n\nReference\n\n[1][https://www.kaggle.com/code/qiweiyin/protenix-v1-inference-2026](url)\n\n[2][https://www.kaggle.com/code/nihilisticneuralnet/0-409-stanford-rna-folding-2-protenix-template](url)\n\n[3][https://www.kaggle.com/code/alexxanderlarko/protenix-v1](url)",
      "votes": 15
    },
    {
      "id": 3441322,
      "postDate": "2026-04-13T19:57:10.840Z",
      "content": "<p>Congratz on the nice finish !\nThe insights on MSA use regarding sequence lengths are very interesting, I want to see how these insights combine with other top team ideas, and RNAPro.</p>\n<p>Would you mind making the inference code and associated datasets public ? \nThanks !</p>",
      "rawMarkdown": "Congratz on the nice finish !\nThe insights on MSA use regarding sequence lengths are very interesting, I want to see how these insights combine with other top team ideas, and RNAPro.\n\nWould you mind making the inference code and associated datasets public ? \nThanks !"
    }
  ],
  "comments": [
    {
      "id": 3441322,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2026-04-13T19:57:10.840000",
      "content": "<p>Congratz on the nice finish !\nThe insights on MSA use regarding sequence lengths are very interesting, I want to see how these insights combine with other top team ideas, and RNAPro.</p>\n<p>Would you mind making the inference code and associated datasets public ? \nThanks !</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3435576": "# Acknowledgments\nAs a graduate student in bioinformatics at University of Science and Technology of China, \"I have learned a lot from this competition, and it is a great help to me and my future. I am grateful for the help of my teammates and believe this was a great collaboration. Most of the open-source notebooks in this competition were built based on Qiwei's dataset; he has extraordinary code insight. Flop provided great ideas for our code improvement in the later stages. Other teammates also cleared up our confusion with professional RNA domain knowledge and provided computational power support.\nCongrats to all the winners and participants!Here's a quick summary of what I did.\n\n# conclusion\n## Phase 1: Template-Based Modeling\nWe searched the combined training and validation set using Biopython's PairwiseAligner with RNA-tuned gap penalties. Atomic coordinates were transferred to the query using a geometry-aware gap-filling procedure: missing residues are interpolated along RNA-like helical trajectories (C1'–C1' step ~5.9 Å) rather than placed at zero or left as NaN. Diversity transforms — hinge rotations, chain jitter, smooth wiggle — are applied slot-by-slot from the same template source, followed by RNA geometry constraint refinement (bond-length correction + Laplacian smoothing + self-avoidance).\nTo avoid slot redundancy, we apply TM-score-based greedy diversity selection: from all TBM candidates, we greedily pick structures that maximize a quality-diversity trade-off score, ensuring the two TBM slots are structurally distinct rather than near-duplicates of the top hit.\nSlot allocation is length-adaptive:\nShort sequences (≤ 512 nt): max 2 TBM slots, 3 reserved for Protenix\nLong sequences (> 512 nt): TBM-first, up to 5 TBM slots if sufficient templates exist\n\n## Phase 2 Core Insight: MSA Depth\nThe mainstream open-source approach treats MSA as a binary switch — either full MSA or no MSA — . I argue this framing discards useful structure in the MSA quality signal.\nMy analysis of the 28 test targets revealed a clear depth distribution:\nDepth ≥ 700  : 15 targets (54%) — strong evolutionary signal\nDepth 10–699 :  7 targets (25%) — moderate, noise risk\nDepth ≤ 9    :  6 targets (21%) — MSA nearly uninformative\n\nI ran ablations across thresholds (100 / 300  / 700 / 1500) and found depth ≥ 700 to be the empirically optimal cutoff: below it, distant homologs inject alignment noise that degrades diffusion quality; above it, evolutionary co-variation provides reliable structural constraints.\nThis leads to our three-mode Protenix inference design:\n~~~~\nSlot 3: msa700 sample 1  — official MSA, depth ≥ 700\n                            high-fidelity evolutionary constraint\nSlot 4: msa700 sample 2  — same MSA policy, independent diffusion\n                            diversity from stochastic diffusion,\n                            NOT from MSA noise\nSlot 5: msa_full         — relaxed threshold, depth ≥ 1\n                            intentional diversity hedge:\n                            noisy-but-different MSA features\n                            only deployed AFTER slots 3&4 are secured\n~~~~\nThe key distinction from the binary on/off approach: slots 3 and 4 share identical high-quality MSA features; diversity between them comes purely from independent diffusion sampling. Slot 5 is a deliberate hedge that accepts some MSA noise in exchange for a structurally distinct prediction.\nFor long sequences (> 512 nt), only msa700 is used with N_sample=3 — three independent diffusion samples in a single forward pass, sharing the expensive feature computation while diversifying outputs.\nFor targets requiring multi-chunk inference, we replaced the standard linear weight ramp with a quintic smoothstep function $$w(t) = 6t^5 - 15t^4 + 10t^3, \\quad t \\in [0, 1]$$. While linear blending results in first-derivative discontinuities at chunk boundaries, our approach guarantees  c2(curvature) continuity. Combined with Kabsch-based reference alignment in the 128-nt overlap regions, this eliminates geometric 'kinks' and ensures the global global topology remains smooth and physically valid.\n\n# Phase 3\nStandard Protenix deployments produce different outputs across runs due to non-deterministic CUDA operations, MSA subsampling, and floating-point ordering effects. We enforced full reproducibility through three mechanisms:\n1.Disabling non-deterministic attention backends — FlashAttention and memory-efficient attention are both switched off, forcing deterministic standard attention\n2.Disabling MC Dropout (mc_dropout_apply_rate=0.0) to remove inference-time stochastic regularization\n3.Forcing torch-native LayerNorm and triangle attention kernels instead of cuDNN/cuEquivariance backends, which have non-deterministic implementations\nThis guarantees that any improvement in our pipeline comes from algorithmic changes, not RNG variation — a critical property for reliable ablation studies in a competition setting.\n\n~~~~\ntest_sequences.csv\n       |\n       v\n+----------------------+\n|       Phase 1        |   PairwiseAligner, top_n=30\n|         TBM          |   geometry-aware gap filling\n|   Diversity Selection|   TM-score greedy selection\n+----------------------+\n       |\n       |  short (≤512 nt): 2 TBM slots\n       |  long  (>512 nt): up to 5 TBM slots (TBM-first)\n       v\n+----------------------+\n|       Phase 2        |   Dual-GPU Protenix\n|      Protenix        |\n|   3 inference modes: |   slot 3: msa700 sample 1\n|   msa700_1           |           (depth≥700, deterministic)\n|   msa700_2           |   slot 4: msa700 sample 2\n|   msa_full           |           (independent diffusion)\n|                      |   slot 5: msa_full\n|                      |           (depth≥1, diversity hedge)\n+----------------------+\n       |\n       | long seqs (>512 nt):\n       | chunk-level greedy bin-packing across 2 GPUs\n       | → Kabsch SVD align + smoothstep stitch\n       v\n  submission.csv\n  (5 slots per target, fully deterministic & reproducible)\n~~~~\n\n\n\n\n\nReference\n\n[1][https://www.kaggle.com/code/qiweiyin/protenix-v1-inference-2026](url)\n\n[2][https://www.kaggle.com/code/nihilisticneuralnet/0-409-stanford-rna-folding-2-protenix-template](url)\n\n[3][https://www.kaggle.com/code/alexxanderlarko/protenix-v1](url)",
    "3441322": "Congratz on the nice finish !\nThe insights on MSA use regarding sequence lengths are very interesting, I want to see how these insights combine with other top team ideas, and RNAPro.\n\nWould you mind making the inference code and associated datasets public ? \nThanks !"
  }
}