{
  "id": 687766,
  "title": "68th Public, 142nd Private Solution - Protenix and Local Aligner for Gap Filling",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/687766",
  "author_name": "Andrew Lee",
  "post_date": "2026-04-03T16:01:45.790000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<h1><strong>1. Approach</strong></h1>\n<p>Without extensive access to GPUs, I realized that to create a viable solution, I couldn't change or fine tune the underlying models, I had to change their usage and pipeline. The main problem I found with many of the top notebooks was their gap interpolation. Basically, when the template used in TBM had a gap that did not match with the query sequence, linear interpolation was used almost exclusively to fill in the gaps, essentially creating a straight line of coordinates in between known coordinates. To solve this, I came up with two methods: Filling with Protenix and with parts of existing TBM templates using Biopython's pairwise aligner.</p>\n<h1><strong>2. Method</strong></h1>\n<h2><strong>2.1 Protenix Gap Filling</strong></h2>\n<p>For Protenix, I took gaps of longer than 3 nt, then ran inference with Protenix using 20 flanking residues on each side to provide extra context, then ran rescaling to ensure that the Protenix output was sized reasonably to a distance of around 5.95 Å between each residue. Then to align it to the rest of the structure I used 5 runs of Iterative Weighted Kabsch Alignment with an RMSD gate of 8 Å, threshold of 3, and a minimum of 10 inliers. My Protenix implementation in my notebook also used dynamic chunking for all predictions, using a minimum chunk size of 20 nt and a maximum of the sequence length divided by 4.</p>\n<h2><strong>2.2 Biopython Local Pairwise Aligner</strong></h2>\n<p>While reading Biopython's aligner documentation, I came across the local mode for the aligner, which I used to take the query gap as an input and find matching subsequences in the training set. After this, I used the same Iterative Weighted Kabsch Alignment as previously mentioned in order to align the gap to the rest of the sequence.</p>\n<h1><strong>3. Implementation</strong></h1>\n<p>My 2 notebooks for submission featured one with pure Protenix for gap interpolation, and one with that ran Biopython's local aligner first, then had Protenix as a fallback for when quality thresholds weren't met for gap sequences.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F31669346%2F3bfa0309ca6dd79012dfd5de7f78cec9%2Frna_structure_prediction_pipeline.svg?generation=1775256307987486&amp;alt=media\" alt=\"\"></p>\n<h1><strong>4. Other Explorations not Submitted</strong></h1>\n<ul>\n<li>One of the earlier public implementations of deep learning methods with TBM, where I used Boltz + RNAPro + TBM, shared below</li>\n<li>Using the Qwen Reasoning model to filter the RNA sequences into groups based on their description and suggest RNA sequences to try as a modeling template, shared below</li>\n<li>Combining TBM + Protenix + Boltz2</li>\n<li>Combining TBM + Protenix + Drfold2</li>\n<li>Averaging multiple mid to high scoring TBM templates (scores between 0.5 and 0.8)</li>\n</ul>\n<h1><strong>5. Conclusion</strong></h1>\n<p>Being a complete beginner to this field and a college freshman, this competition has been an extremely valuable learning experience for me, teaching me how to think innovatively, turn my ideas into code, and work with large datasets. I enjoyed how this competition rewarded diversity of solutions and innovation, and I strove to create that with my pipeline. Thank you to the competition organizers and Rhiju Das for this amazing competition.</p>",
  "messages": [
    {
      "id": 3435053,
      "postDate": "2026-04-03T16:01:45.790Z",
      "content": "<h1><strong>1. Approach</strong></h1>\n<p>Without extensive access to GPUs, I realized that to create a viable solution, I couldn't change or fine tune the underlying models, I had to change their usage and pipeline. The main problem I found with many of the top notebooks was their gap interpolation. Basically, when the template used in TBM had a gap that did not match with the query sequence, linear interpolation was used almost exclusively to fill in the gaps, essentially creating a straight line of coordinates in between known coordinates. To solve this, I came up with two methods: Filling with Protenix and with parts of existing TBM templates using Biopython's pairwise aligner.</p>\n<h1><strong>2. Method</strong></h1>\n<h2><strong>2.1 Protenix Gap Filling</strong></h2>\n<p>For Protenix, I took gaps of longer than 3 nt, then ran inference with Protenix using 20 flanking residues on each side to provide extra context, then ran rescaling to ensure that the Protenix output was sized reasonably to a distance of around 5.95 Å between each residue. Then to align it to the rest of the structure I used 5 runs of Iterative Weighted Kabsch Alignment with an RMSD gate of 8 Å, threshold of 3, and a minimum of 10 inliers. My Protenix implementation in my notebook also used dynamic chunking for all predictions, using a minimum chunk size of 20 nt and a maximum of the sequence length divided by 4.</p>\n<h2><strong>2.2 Biopython Local Pairwise Aligner</strong></h2>\n<p>While reading Biopython's aligner documentation, I came across the local mode for the aligner, which I used to take the query gap as an input and find matching subsequences in the training set. After this, I used the same Iterative Weighted Kabsch Alignment as previously mentioned in order to align the gap to the rest of the sequence.</p>\n<h1><strong>3. Implementation</strong></h1>\n<p>My 2 notebooks for submission featured one with pure Protenix for gap interpolation, and one with that ran Biopython's local aligner first, then had Protenix as a fallback for when quality thresholds weren't met for gap sequences.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F31669346%2F3bfa0309ca6dd79012dfd5de7f78cec9%2Frna_structure_prediction_pipeline.svg?generation=1775256307987486&amp;alt=media\" alt=\"\"></p>\n<h1><strong>4. Other Explorations not Submitted</strong></h1>\n<ul>\n<li>One of the earlier public implementations of deep learning methods with TBM, where I used Boltz + RNAPro + TBM, shared below</li>\n<li>Using the Qwen Reasoning model to filter the RNA sequences into groups based on their description and suggest RNA sequences to try as a modeling template, shared below</li>\n<li>Combining TBM + Protenix + Boltz2</li>\n<li>Combining TBM + Protenix + Drfold2</li>\n<li>Averaging multiple mid to high scoring TBM templates (scores between 0.5 and 0.8)</li>\n</ul>\n<h1><strong>5. Conclusion</strong></h1>\n<p>Being a complete beginner to this field and a college freshman, this competition has been an extremely valuable learning experience for me, teaching me how to think innovatively, turn my ideas into code, and work with large datasets. I enjoyed how this competition rewarded diversity of solutions and innovation, and I strove to create that with my pipeline. Thank you to the competition organizers and Rhiju Das for this amazing competition.</p>",
      "rawMarkdown": "# **1. Approach**\nWithout extensive access to GPUs, I realized that to create a viable solution, I couldn't change or fine tune the underlying models, I had to change their usage and pipeline. The main problem I found with many of the top notebooks was their gap interpolation. Basically, when the template used in TBM had a gap that did not match with the query sequence, linear interpolation was used almost exclusively to fill in the gaps, essentially creating a straight line of coordinates in between known coordinates. To solve this, I came up with two methods: Filling with Protenix and with parts of existing TBM templates using Biopython's pairwise aligner.\n\n# **2. Method**\n## **2.1 Protenix Gap Filling**\nFor Protenix, I took gaps of longer than 3 nt, then ran inference with Protenix using 20 flanking residues on each side to provide extra context, then ran rescaling to ensure that the Protenix output was sized reasonably to a distance of around 5.95 Å between each residue. Then to align it to the rest of the structure I used 5 runs of Iterative Weighted Kabsch Alignment with an RMSD gate of 8 Å, threshold of 3, and a minimum of 10 inliers. My Protenix implementation in my notebook also used dynamic chunking for all predictions, using a minimum chunk size of 20 nt and a maximum of the sequence length divided by 4.\n\n## **2.2 Biopython Local Pairwise Aligner**\nWhile reading Biopython's aligner documentation, I came across the local mode for the aligner, which I used to take the query gap as an input and find matching subsequences in the training set. After this, I used the same Iterative Weighted Kabsch Alignment as previously mentioned in order to align the gap to the rest of the sequence.\n\n# **3. Implementation**\nMy 2 notebooks for submission featured one with pure Protenix for gap interpolation, and one with that ran Biopython's local aligner first, then had Protenix as a fallback for when quality thresholds weren't met for gap sequences.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F31669346%2F3bfa0309ca6dd79012dfd5de7f78cec9%2Frna_structure_prediction_pipeline.svg?generation=1775256307987486&alt=media)\n\n# **4. Other Explorations not Submitted**\n- One of the earlier public implementations of deep learning methods with TBM, where I used Boltz + RNAPro + TBM, shared below\n- Using the Qwen Reasoning model to filter the RNA sequences into groups based on their description and suggest RNA sequences to try as a modeling template, shared below\n- Combining TBM + Protenix + Boltz2\n- Combining TBM + Protenix + Drfold2\n- Averaging multiple mid to high scoring TBM templates (scores between 0.5 and 0.8)\n\n# **5. Conclusion**\nBeing a complete beginner to this field and a college freshman, this competition has been an extremely valuable learning experience for me, teaching me how to think innovatively, turn my ideas into code, and work with large datasets. I enjoyed how this competition rewarded diversity of solutions and innovation, and I strove to create that with my pipeline. Thank you to the competition organizers and Rhiju Das for this amazing competition.",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3435053": "# **1. Approach**\nWithout extensive access to GPUs, I realized that to create a viable solution, I couldn't change or fine tune the underlying models, I had to change their usage and pipeline. The main problem I found with many of the top notebooks was their gap interpolation. Basically, when the template used in TBM had a gap that did not match with the query sequence, linear interpolation was used almost exclusively to fill in the gaps, essentially creating a straight line of coordinates in between known coordinates. To solve this, I came up with two methods: Filling with Protenix and with parts of existing TBM templates using Biopython's pairwise aligner.\n\n# **2. Method**\n## **2.1 Protenix Gap Filling**\nFor Protenix, I took gaps of longer than 3 nt, then ran inference with Protenix using 20 flanking residues on each side to provide extra context, then ran rescaling to ensure that the Protenix output was sized reasonably to a distance of around 5.95 Å between each residue. Then to align it to the rest of the structure I used 5 runs of Iterative Weighted Kabsch Alignment with an RMSD gate of 8 Å, threshold of 3, and a minimum of 10 inliers. My Protenix implementation in my notebook also used dynamic chunking for all predictions, using a minimum chunk size of 20 nt and a maximum of the sequence length divided by 4.\n\n## **2.2 Biopython Local Pairwise Aligner**\nWhile reading Biopython's aligner documentation, I came across the local mode for the aligner, which I used to take the query gap as an input and find matching subsequences in the training set. After this, I used the same Iterative Weighted Kabsch Alignment as previously mentioned in order to align the gap to the rest of the sequence.\n\n# **3. Implementation**\nMy 2 notebooks for submission featured one with pure Protenix for gap interpolation, and one with that ran Biopython's local aligner first, then had Protenix as a fallback for when quality thresholds weren't met for gap sequences.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F31669346%2F3bfa0309ca6dd79012dfd5de7f78cec9%2Frna_structure_prediction_pipeline.svg?generation=1775256307987486&alt=media)\n\n# **4. Other Explorations not Submitted**\n- One of the earlier public implementations of deep learning methods with TBM, where I used Boltz + RNAPro + TBM, shared below\n- Using the Qwen Reasoning model to filter the RNA sequences into groups based on their description and suggest RNA sequences to try as a modeling template, shared below\n- Combining TBM + Protenix + Boltz2\n- Combining TBM + Protenix + Drfold2\n- Averaging multiple mid to high scoring TBM templates (scores between 0.5 and 0.8)\n\n# **5. Conclusion**\nBeing a complete beginner to this field and a college freshman, this competition has been an extremely valuable learning experience for me, teaching me how to think innovatively, turn my ideas into code, and work with large datasets. I enjoyed how this competition rewarded diversity of solutions and innovation, and I strove to create that with my pipeline. Thank you to the competition organizers and Rhiju Das for this amazing competition."
  }
}