{
  "id": 688937,
  "title": "15th Place Solution - A Beginner Data Scientist's Journey",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/688937",
  "author_name": "PanditaProgrammer",
  "post_date": "2026-04-07T10:48:33.627000",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>Stanford RNA 3D Folding Part 2 - Competition Writeup</h1>\n<p>Hi! Thanks to all the team behind the Stanford RNA 3D Folding Part 2 competition. It was a fantastic challenge, and we had a great time learning throughout the process.</p>\n<p>I am a student at UPM (Polytechnic University of Madrid), and as this was my first serious Kaggle competition, the learning curve was steep but incredibly rewarding. I am representing <a href=\"https://www.whiteboxml.com/\" target=\"_blank\">WhiteBox</a>, a boutique data science and ML consultancy in Spain where I am currently a Data Scientist Intern.</p>\n<h3>Our Philosophy</h3>\n<p>Our team is passionate about bridging the gap between research and real-world applications. This competition was the perfect playground to test our skills in a completely new domain; we approached the problem as data scientists with no prior background in biology, which made the experience even more valuable as it allowed us to test our transferable skills when pivoting to unfamiliar fields.</p>\n<h2>Problem Understanding &amp; Domain Knowledge</h2>\n<p>Before diving into modeling, we spent a significant amount of time understanding the problem and the domain. We read through the competition description, explored the provided datasets, and researched the multiple factors that influence RNA folding, such as sequence similarity, ligand interactions, and the presence other chains. This foundational understanding allowed us to make better decisions and validate our approaches more effectively as we progressed through the competition. </p>\n<h2>Local Validation Strategy</h2>\n<p>Since the competition evaluation is based on TM-score computed via USalign, we used the same metric for out local validation. We though about creating our own train/validation split from the training data, but given that we were not experts in the domain, we were concerned about data leakage and ended up using the default validation set after reading how the split were done. Then we focused on understanding the type of chain we were predicting (monomer, homomer, heteromer, synthetic, the function of that chain, etc.) to try to understand why one method could work and why not.</p>\n<p>One of the keys to our success I would say was the creation of an interactive 3D visualization dashboard that allowed us to visually inspect the predicted vs. ground truth structures. This was crucial for understanding the strengths and weaknesses of our models, especially in cases where the TM-score did not fully capture the quality of the prediction. It helped us identify specific issues such as steric clashes, incorrect chain positioning, and other structural artifacts that we could then address in our modeling approaches.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17092985%2F68d953c2104cf4696868e06849c73049%2Fimage.png?generation=1775558628697326&amp;alt=media\" alt=\"\">\nAs you can see it gives us a side-by-side comparison of the predicted and true structures, allowing us to rotate and zoom in on specific regions to analyze the quality of the prediction in detail. Instead on using only the TM-score (1D metric), this dashboard provided a more intuitive understanding of the 3D structural differences, which was invaluable for debugging and improving our models.</p>\n<h2>The journey: From Bottom to Top 11 to Bottom to Final Top 15</h2>\n<p>Once we have gathered domain knowledge and set up our local validation strategy, we started experimenting with different approaches. We tried a wide range of methods, from template-based modeling (TBM) to finetuning Protenix on the competition data, and even explored other models like DRFold2 and Chai. We also experimented with different ways of incorporating ligand information and handling multimers.</p>\n<p>The key here was to understand the strengths and weaknesses of each approach and how they performed on different types of sequences (monomers, homomers, heteromers, synthetic…). We iteratively refined our models based on the insights we gained from our local validation and the 3D visualization dashboard.</p>\n<p>We spend many hours debugging and improving our TBM approach, which ended up being the backbone of our final pipeline. A different approach we found promising was to weigh the importance of the ligand similarity and the chain count in the template selection along with another TBM more similar to the one used by everyone.</p>\n<p>The second approach we have taken was to use RNAPro, the model from NVIDIA Digital Bio that came from mixing the top solutions from Phase 1. Setting them on default parameters and using the TBM output as template for it worked really well for sequences &lt;= 350nt. For longer sequences it was lacking the precision I would like it to have. At the end we used two RNAPros with different template predictions and parameters to increase the diversity of our ensemble. For example if we use more cycles it could benefit the giant RNAs (9LEC, 9LEL) as it would be able to refine the structure more, but it would also decrease the performance for smaller sequences.</p>\n<p>We have tried also DRFold2, Chai-1 and Boltz-1 but I didn't like them as much as RNAPro and Protenix, so we deprioritized them to focus on Protenix to the short time we had left.</p>\n<p>Protenix, Protenix, and more Protenix. It is a diffusion model that, depending on its random seed, can produce very different predictions. For those who do not know it, it allows you to input a vast amount of data and use different approaches available to those who have taken the time to read the documentation in the era of chatbots.</p>\n<p>I was able to configure them locally without problems but I struggled to upload my solution to Kaggle due to the many things needed for it to function properly. I will share the more valuable insights I got from the predictions on local.</p>\n<p>Protenix Versions: I have tested the 4 models from Protenix (the 202560630, the base one, the v0.5 and the v0.2) and each of them has its own strengths and weaknesses.</p>\n<p>The 2025 one at least seems better overall in tm-score for the test compared to the others, but I would like to highlight that for more simple RNAs the base performance is better, and for more complex ones the 2025 one is better.</p>\n<p>i have played with all the parameters of Protenix and depending on the parameters it could give us completely different results. </p>\n<ul>\n<li>If we are working with other Proteins and DNA it is better to activate the flag, if not don't activate as it decrease the performance for RNA.</li>\n<li>For Giant RNAs (9LEC, 9LEL) it is better to use more cycles and more recycles, but for smaller ones it is better to use less cycles and recycles.</li>\n<li>For viruses don't use Protenix.</li>\n<li>For monomers Protenix works acceptably.</li>\n<li>For homomers and heteromers it is bad at position them correctly in space, it is able to predict the structure of each chain really well. For example the 9KGG is a homomultimer of 2 chains, Protenix is able to predict the structure for each chain but no to position them correctly so instead of a maximun tm-score of 1 we could only get a tm-score maximum of 0.5. We thought of ways to normalize the chains and simplify the positioning but you just need to check other difficult samples like 9G4Q or 9WHV to see that there's no unique answer.</li>\n<li>We know that Protenix does not work well when no template is available, in local we were able to create the fasta files to supply the lack of it but making it work on Kaggle was a nightmare and we didn't have time to do it. </li>\n</ul>\n<p>Knowing the weaknessed we thought about finetunning Protenix to make it work with heteromers and homomers, middleway the results were promising but the variance across random seeds was too high for me to feel confident about it, so we ended up not using it for the final submission. </p>\n<p>Some interesting images (left is the default and right is the fine-tuned one):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17092985%2F57c736a1ba578f9eef37922638937df5%2Fimage2.png?generation=1775558589328801&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17092985%2F05d11146e132491d43c47b1234a10d88%2Fimage3.png?generation=1775558613294141&amp;alt=media\" alt=\"\"></p>\n<p>We have done some tweaks to the Protenix code (changing the loss funtion to a custom derivable approximation of the TM-score, data augmentation…) but we were not able to solve the problem of the high variance across random seeds and with the tight deadline I decided not to pursue it. If you are curious when we fine-tune Protenix as happened before there were no way to deal with the Catastrophic forgetting problem, our model finetune works better for those sequences but forgot how to fold the others.</p>\n<ul>\n<li>The inclusion of ligands can improve the performance in some cases but decrease it in other.</li>\n<li>We have tried using an embedding instead of the sequence and the result imrpoved for some sequences but decreased for others. The improved one were those were we didn't have a good template but with an embedding we were able to capture some of the structural information.</li>\n</ul>\n<p>Those along with other things were discovered thanks to the 3D visualization dashboard. Right now I do not remember much curiosities as I have done a lot of experiments and I have not written down all the insights.</p>\n<h2>Approach</h2>\n<h3>Final Strategy: Best-of-5 Ensemble</h3>\n<p>Our final submission uses 5 prediction slots strategically:</p>\n<p><strong>Slot 1 - Pure TBM with Biopython Alignment</strong></p>\n<ul>\n<li>Align test sequences against all training sequences using Biopython pairwise alignment</li>\n<li>Minimum 50% identity threshold</li>\n<li>Linear interpolation for missing coordinates</li>\n<li>Adaptive constraints: bond distance correction (~5.95 A for adjacent residues), angular constraints (~10.2 A for i+2), Laplacian smoothing, and steric clash prevention</li>\n</ul>\n<p><strong>Slots 2-3 - TBM + RNAPro (Protenix + RibonanzaNet Embeddings)</strong></p>\n<ul>\n<li>Weighted template selection: 85% sequence similarity + 15% ligand similarity</li>\n<li>Also considers number of unique sequences and total chain count</li>\n<li>TBM coordinates serve as templates for the RNAPro model (NVIDIA Digital Bio, Phase 1 winner approach)</li>\n<li>Works best for sequences &lt;= 350 nucleotides</li>\n<li>For longer sequences: window sliding and stitching strategy</li>\n<li>Slot 3 uses slightly different parameters and uses Slot 1 output as its template</li>\n</ul>\n<p><strong>Slots 4-5 - Protenix 2025 from Scratch</strong></p>\n<ul>\n<li>ByteDance's trainable PyTorch reproduction of AlphaFold 3</li>\n<li>JSON-based input with correct stoichiometry information</li>\n<li>No ligand input (to avoid OOM errors on large complexes)</li>\n<li>Window sliding and stitching for sequences &gt; 350 nt</li>\n<li>Slot 5 is a duplicate of Slot 4 (ran out of time)</li>\n<li>Slot 4 should be for the base Protenix but OOT issues.</li>\n</ul>\n<p>If you think carefully instead of doing 5 solutions in our case was more like 3.5-4 solutions as the last two were the same coordinates.</p>\n<p>Our main focus was to make sure each predictions actually makes sense. a major part of the evaluation was done by me visually inspecting the predictions and comparing them to the ground truth using our 3D visualization dashboard. It is important to develop the intuition of data scientist and to not overfit and try to get as much information as you can about the data.</p>\n<h2>What We Would Do Differently</h2>\n<ul>\n<li><strong>Protenix finetuning</strong> - I would have loved to fine-tune Protenix but as I am not familiar with the domain I struggle to do it.</li>\n<li><strong>More sophisticated ligand weighting</strong> in template selection, that was a basic one we check were working right</li>\n<li><strong>Secondary structure integration</strong> - we deprioritized this but it could have improved predictions by restrictions</li>\n<li><strong>Larger ensemble diversity</strong> - combining DRFold2, Chai, and Protenix could have improved coverage</li>\n</ul>\n<h2>Acknowledgments</h2>\n<ul>\n<li>Stanford and Kaggle for hosting the competition</li>\n<li>NVIDIA Digital Bio team for the RNAPro model from Phase 1</li>\n<li>ByteDance for the open-source Protenix (AlphaFold 3 reproduction)</li>\n<li>The Kaggle community for forum discussions and shared notebooks</li>\n<li>My teammates at WhiteBox for their collaboration and support throughout the competition</li>\n</ul>\n<h2>Conclusion</h2>\n<p>Overall, this competition was an incredible learning experience that pushed us to apply our data science skills in a completely new domain. We had a lot of fun experimenting with different approaches, analyzing the results, and iteratively improving our models. We are proud of our final submission and grateful for the opportunity to participate in such a challenging and rewarding competition. We look forward to seeing how the field of RNA 3D folding continues to evolve and hope that our contributions can help advance the state of the art in this exciting area of research.</p>",
  "messages": [
    {
      "id": 3437215,
      "postDate": "2026-04-07T10:48:33.627Z",
      "content": "<h1>Stanford RNA 3D Folding Part 2 - Competition Writeup</h1>\n<p>Hi! Thanks to all the team behind the Stanford RNA 3D Folding Part 2 competition. It was a fantastic challenge, and we had a great time learning throughout the process.</p>\n<p>I am a student at UPM (Polytechnic University of Madrid), and as this was my first serious Kaggle competition, the learning curve was steep but incredibly rewarding. I am representing <a href=\"https://www.whiteboxml.com/\" target=\"_blank\">WhiteBox</a>, a boutique data science and ML consultancy in Spain where I am currently a Data Scientist Intern.</p>\n<h3>Our Philosophy</h3>\n<p>Our team is passionate about bridging the gap between research and real-world applications. This competition was the perfect playground to test our skills in a completely new domain; we approached the problem as data scientists with no prior background in biology, which made the experience even more valuable as it allowed us to test our transferable skills when pivoting to unfamiliar fields.</p>\n<h2>Problem Understanding &amp; Domain Knowledge</h2>\n<p>Before diving into modeling, we spent a significant amount of time understanding the problem and the domain. We read through the competition description, explored the provided datasets, and researched the multiple factors that influence RNA folding, such as sequence similarity, ligand interactions, and the presence other chains. This foundational understanding allowed us to make better decisions and validate our approaches more effectively as we progressed through the competition. </p>\n<h2>Local Validation Strategy</h2>\n<p>Since the competition evaluation is based on TM-score computed via USalign, we used the same metric for out local validation. We though about creating our own train/validation split from the training data, but given that we were not experts in the domain, we were concerned about data leakage and ended up using the default validation set after reading how the split were done. Then we focused on understanding the type of chain we were predicting (monomer, homomer, heteromer, synthetic, the function of that chain, etc.) to try to understand why one method could work and why not.</p>\n<p>One of the keys to our success I would say was the creation of an interactive 3D visualization dashboard that allowed us to visually inspect the predicted vs. ground truth structures. This was crucial for understanding the strengths and weaknesses of our models, especially in cases where the TM-score did not fully capture the quality of the prediction. It helped us identify specific issues such as steric clashes, incorrect chain positioning, and other structural artifacts that we could then address in our modeling approaches.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17092985%2F68d953c2104cf4696868e06849c73049%2Fimage.png?generation=1775558628697326&amp;alt=media\" alt=\"\">\nAs you can see it gives us a side-by-side comparison of the predicted and true structures, allowing us to rotate and zoom in on specific regions to analyze the quality of the prediction in detail. Instead on using only the TM-score (1D metric), this dashboard provided a more intuitive understanding of the 3D structural differences, which was invaluable for debugging and improving our models.</p>\n<h2>The journey: From Bottom to Top 11 to Bottom to Final Top 15</h2>\n<p>Once we have gathered domain knowledge and set up our local validation strategy, we started experimenting with different approaches. We tried a wide range of methods, from template-based modeling (TBM) to finetuning Protenix on the competition data, and even explored other models like DRFold2 and Chai. We also experimented with different ways of incorporating ligand information and handling multimers.</p>\n<p>The key here was to understand the strengths and weaknesses of each approach and how they performed on different types of sequences (monomers, homomers, heteromers, synthetic…). We iteratively refined our models based on the insights we gained from our local validation and the 3D visualization dashboard.</p>\n<p>We spend many hours debugging and improving our TBM approach, which ended up being the backbone of our final pipeline. A different approach we found promising was to weigh the importance of the ligand similarity and the chain count in the template selection along with another TBM more similar to the one used by everyone.</p>\n<p>The second approach we have taken was to use RNAPro, the model from NVIDIA Digital Bio that came from mixing the top solutions from Phase 1. Setting them on default parameters and using the TBM output as template for it worked really well for sequences &lt;= 350nt. For longer sequences it was lacking the precision I would like it to have. At the end we used two RNAPros with different template predictions and parameters to increase the diversity of our ensemble. For example if we use more cycles it could benefit the giant RNAs (9LEC, 9LEL) as it would be able to refine the structure more, but it would also decrease the performance for smaller sequences.</p>\n<p>We have tried also DRFold2, Chai-1 and Boltz-1 but I didn't like them as much as RNAPro and Protenix, so we deprioritized them to focus on Protenix to the short time we had left.</p>\n<p>Protenix, Protenix, and more Protenix. It is a diffusion model that, depending on its random seed, can produce very different predictions. For those who do not know it, it allows you to input a vast amount of data and use different approaches available to those who have taken the time to read the documentation in the era of chatbots.</p>\n<p>I was able to configure them locally without problems but I struggled to upload my solution to Kaggle due to the many things needed for it to function properly. I will share the more valuable insights I got from the predictions on local.</p>\n<p>Protenix Versions: I have tested the 4 models from Protenix (the 202560630, the base one, the v0.5 and the v0.2) and each of them has its own strengths and weaknesses.</p>\n<p>The 2025 one at least seems better overall in tm-score for the test compared to the others, but I would like to highlight that for more simple RNAs the base performance is better, and for more complex ones the 2025 one is better.</p>\n<p>i have played with all the parameters of Protenix and depending on the parameters it could give us completely different results. </p>\n<ul>\n<li>If we are working with other Proteins and DNA it is better to activate the flag, if not don't activate as it decrease the performance for RNA.</li>\n<li>For Giant RNAs (9LEC, 9LEL) it is better to use more cycles and more recycles, but for smaller ones it is better to use less cycles and recycles.</li>\n<li>For viruses don't use Protenix.</li>\n<li>For monomers Protenix works acceptably.</li>\n<li>For homomers and heteromers it is bad at position them correctly in space, it is able to predict the structure of each chain really well. For example the 9KGG is a homomultimer of 2 chains, Protenix is able to predict the structure for each chain but no to position them correctly so instead of a maximun tm-score of 1 we could only get a tm-score maximum of 0.5. We thought of ways to normalize the chains and simplify the positioning but you just need to check other difficult samples like 9G4Q or 9WHV to see that there's no unique answer.</li>\n<li>We know that Protenix does not work well when no template is available, in local we were able to create the fasta files to supply the lack of it but making it work on Kaggle was a nightmare and we didn't have time to do it. </li>\n</ul>\n<p>Knowing the weaknessed we thought about finetunning Protenix to make it work with heteromers and homomers, middleway the results were promising but the variance across random seeds was too high for me to feel confident about it, so we ended up not using it for the final submission. </p>\n<p>Some interesting images (left is the default and right is the fine-tuned one):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17092985%2F57c736a1ba578f9eef37922638937df5%2Fimage2.png?generation=1775558589328801&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17092985%2F05d11146e132491d43c47b1234a10d88%2Fimage3.png?generation=1775558613294141&amp;alt=media\" alt=\"\"></p>\n<p>We have done some tweaks to the Protenix code (changing the loss funtion to a custom derivable approximation of the TM-score, data augmentation…) but we were not able to solve the problem of the high variance across random seeds and with the tight deadline I decided not to pursue it. If you are curious when we fine-tune Protenix as happened before there were no way to deal with the Catastrophic forgetting problem, our model finetune works better for those sequences but forgot how to fold the others.</p>\n<ul>\n<li>The inclusion of ligands can improve the performance in some cases but decrease it in other.</li>\n<li>We have tried using an embedding instead of the sequence and the result imrpoved for some sequences but decreased for others. The improved one were those were we didn't have a good template but with an embedding we were able to capture some of the structural information.</li>\n</ul>\n<p>Those along with other things were discovered thanks to the 3D visualization dashboard. Right now I do not remember much curiosities as I have done a lot of experiments and I have not written down all the insights.</p>\n<h2>Approach</h2>\n<h3>Final Strategy: Best-of-5 Ensemble</h3>\n<p>Our final submission uses 5 prediction slots strategically:</p>\n<p><strong>Slot 1 - Pure TBM with Biopython Alignment</strong></p>\n<ul>\n<li>Align test sequences against all training sequences using Biopython pairwise alignment</li>\n<li>Minimum 50% identity threshold</li>\n<li>Linear interpolation for missing coordinates</li>\n<li>Adaptive constraints: bond distance correction (~5.95 A for adjacent residues), angular constraints (~10.2 A for i+2), Laplacian smoothing, and steric clash prevention</li>\n</ul>\n<p><strong>Slots 2-3 - TBM + RNAPro (Protenix + RibonanzaNet Embeddings)</strong></p>\n<ul>\n<li>Weighted template selection: 85% sequence similarity + 15% ligand similarity</li>\n<li>Also considers number of unique sequences and total chain count</li>\n<li>TBM coordinates serve as templates for the RNAPro model (NVIDIA Digital Bio, Phase 1 winner approach)</li>\n<li>Works best for sequences &lt;= 350 nucleotides</li>\n<li>For longer sequences: window sliding and stitching strategy</li>\n<li>Slot 3 uses slightly different parameters and uses Slot 1 output as its template</li>\n</ul>\n<p><strong>Slots 4-5 - Protenix 2025 from Scratch</strong></p>\n<ul>\n<li>ByteDance's trainable PyTorch reproduction of AlphaFold 3</li>\n<li>JSON-based input with correct stoichiometry information</li>\n<li>No ligand input (to avoid OOM errors on large complexes)</li>\n<li>Window sliding and stitching for sequences &gt; 350 nt</li>\n<li>Slot 5 is a duplicate of Slot 4 (ran out of time)</li>\n<li>Slot 4 should be for the base Protenix but OOT issues.</li>\n</ul>\n<p>If you think carefully instead of doing 5 solutions in our case was more like 3.5-4 solutions as the last two were the same coordinates.</p>\n<p>Our main focus was to make sure each predictions actually makes sense. a major part of the evaluation was done by me visually inspecting the predictions and comparing them to the ground truth using our 3D visualization dashboard. It is important to develop the intuition of data scientist and to not overfit and try to get as much information as you can about the data.</p>\n<h2>What We Would Do Differently</h2>\n<ul>\n<li><strong>Protenix finetuning</strong> - I would have loved to fine-tune Protenix but as I am not familiar with the domain I struggle to do it.</li>\n<li><strong>More sophisticated ligand weighting</strong> in template selection, that was a basic one we check were working right</li>\n<li><strong>Secondary structure integration</strong> - we deprioritized this but it could have improved predictions by restrictions</li>\n<li><strong>Larger ensemble diversity</strong> - combining DRFold2, Chai, and Protenix could have improved coverage</li>\n</ul>\n<h2>Acknowledgments</h2>\n<ul>\n<li>Stanford and Kaggle for hosting the competition</li>\n<li>NVIDIA Digital Bio team for the RNAPro model from Phase 1</li>\n<li>ByteDance for the open-source Protenix (AlphaFold 3 reproduction)</li>\n<li>The Kaggle community for forum discussions and shared notebooks</li>\n<li>My teammates at WhiteBox for their collaboration and support throughout the competition</li>\n</ul>\n<h2>Conclusion</h2>\n<p>Overall, this competition was an incredible learning experience that pushed us to apply our data science skills in a completely new domain. We had a lot of fun experimenting with different approaches, analyzing the results, and iteratively improving our models. We are proud of our final submission and grateful for the opportunity to participate in such a challenging and rewarding competition. We look forward to seeing how the field of RNA 3D folding continues to evolve and hope that our contributions can help advance the state of the art in this exciting area of research.</p>",
      "rawMarkdown": "# Stanford RNA 3D Folding Part 2 - Competition Writeup\n\nHi! Thanks to all the team behind the Stanford RNA 3D Folding Part 2 competition. It was a fantastic challenge, and we had a great time learning throughout the process.\n\nI am a student at UPM (Polytechnic University of Madrid), and as this was my first serious Kaggle competition, the learning curve was steep but incredibly rewarding. I am representing [WhiteBox](https://www.whiteboxml.com/), a boutique data science and ML consultancy in Spain where I am currently a Data Scientist Intern.\n\n### Our Philosophy\nOur team is passionate about bridging the gap between research and real-world applications. This competition was the perfect playground to test our skills in a completely new domain; we approached the problem as data scientists with no prior background in biology, which made the experience even more valuable as it allowed us to test our transferable skills when pivoting to unfamiliar fields.\n\n\n## Problem Understanding & Domain Knowledge\nBefore diving into modeling, we spent a significant amount of time understanding the problem and the domain. We read through the competition description, explored the provided datasets, and researched the multiple factors that influence RNA folding, such as sequence similarity, ligand interactions, and the presence other chains. This foundational understanding allowed us to make better decisions and validate our approaches more effectively as we progressed through the competition. \n\n\n## Local Validation Strategy\nSince the competition evaluation is based on TM-score computed via USalign, we used the same metric for out local validation. We though about creating our own train/validation split from the training data, but given that we were not experts in the domain, we were concerned about data leakage and ended up using the default validation set after reading how the split were done. Then we focused on understanding the type of chain we were predicting (monomer, homomer, heteromer, synthetic, the function of that chain, etc.) to try to understand why one method could work and why not.\n\nOne of the keys to our success I would say was the creation of an interactive 3D visualization dashboard that allowed us to visually inspect the predicted vs. ground truth structures. This was crucial for understanding the strengths and weaknesses of our models, especially in cases where the TM-score did not fully capture the quality of the prediction. It helped us identify specific issues such as steric clashes, incorrect chain positioning, and other structural artifacts that we could then address in our modeling approaches.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17092985%2F68d953c2104cf4696868e06849c73049%2Fimage.png?generation=1775558628697326&alt=media)\nAs you can see it gives us a side-by-side comparison of the predicted and true structures, allowing us to rotate and zoom in on specific regions to analyze the quality of the prediction in detail. Instead on using only the TM-score (1D metric), this dashboard provided a more intuitive understanding of the 3D structural differences, which was invaluable for debugging and improving our models.\n\n## The journey: From Bottom to Top 11 to Bottom to Final Top 15\n\nOnce we have gathered domain knowledge and set up our local validation strategy, we started experimenting with different approaches. We tried a wide range of methods, from template-based modeling (TBM) to finetuning Protenix on the competition data, and even explored other models like DRFold2 and Chai. We also experimented with different ways of incorporating ligand information and handling multimers.\n\nThe key here was to understand the strengths and weaknesses of each approach and how they performed on different types of sequences (monomers, homomers, heteromers, synthetic...). We iteratively refined our models based on the insights we gained from our local validation and the 3D visualization dashboard.\n\nWe spend many hours debugging and improving our TBM approach, which ended up being the backbone of our final pipeline. A different approach we found promising was to weigh the importance of the ligand similarity and the chain count in the template selection along with another TBM more similar to the one used by everyone.\n\nThe second approach we have taken was to use RNAPro, the model from NVIDIA Digital Bio that came from mixing the top solutions from Phase 1. Setting them on default parameters and using the TBM output as template for it worked really well for sequences <= 350nt. For longer sequences it was lacking the precision I would like it to have. At the end we used two RNAPros with different template predictions and parameters to increase the diversity of our ensemble. For example if we use more cycles it could benefit the giant RNAs (9LEC, 9LEL) as it would be able to refine the structure more, but it would also decrease the performance for smaller sequences.\n\nWe have tried also DRFold2, Chai-1 and Boltz-1 but I didn't like them as much as RNAPro and Protenix, so we deprioritized them to focus on Protenix to the short time we had left.\n\nProtenix, Protenix, and more Protenix. It is a diffusion model that, depending on its random seed, can produce very different predictions. For those who do not know it, it allows you to input a vast amount of data and use different approaches available to those who have taken the time to read the documentation in the era of chatbots.\n\nI was able to configure them locally without problems but I struggled to upload my solution to Kaggle due to the many things needed for it to function properly. I will share the more valuable insights I got from the predictions on local.\n\nProtenix Versions: I have tested the 4 models from Protenix (the 202560630, the base one, the v0.5 and the v0.2) and each of them has its own strengths and weaknesses.\n\nThe 2025 one at least seems better overall in tm-score for the test compared to the others, but I would like to highlight that for more simple RNAs the base performance is better, and for more complex ones the 2025 one is better.\n\ni have played with all the parameters of Protenix and depending on the parameters it could give us completely different results. \n\n- If we are working with other Proteins and DNA it is better to activate the flag, if not don't activate as it decrease the performance for RNA.\n- For Giant RNAs (9LEC, 9LEL) it is better to use more cycles and more recycles, but for smaller ones it is better to use less cycles and recycles.\n- For viruses don't use Protenix.\n- For monomers Protenix works acceptably.\n- For homomers and heteromers it is bad at position them correctly in space, it is able to predict the structure of each chain really well. For example the 9KGG is a homomultimer of 2 chains, Protenix is able to predict the structure for each chain but no to position them correctly so instead of a maximun tm-score of 1 we could only get a tm-score maximum of 0.5. We thought of ways to normalize the chains and simplify the positioning but you just need to check other difficult samples like 9G4Q or 9WHV to see that there's no unique answer.\n- We know that Protenix does not work well when no template is available, in local we were able to create the fasta files to supply the lack of it but making it work on Kaggle was a nightmare and we didn't have time to do it. \n\nKnowing the weaknessed we thought about finetunning Protenix to make it work with heteromers and homomers, middleway the results were promising but the variance across random seeds was too high for me to feel confident about it, so we ended up not using it for the final submission. \n\nSome interesting images (left is the default and right is the fine-tuned one):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17092985%2F57c736a1ba578f9eef37922638937df5%2Fimage2.png?generation=1775558589328801&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17092985%2F05d11146e132491d43c47b1234a10d88%2Fimage3.png?generation=1775558613294141&alt=media)\n\nWe have done some tweaks to the Protenix code (changing the loss funtion to a custom derivable approximation of the TM-score, data augmentation...) but we were not able to solve the problem of the high variance across random seeds and with the tight deadline I decided not to pursue it. If you are curious when we fine-tune Protenix as happened before there were no way to deal with the Catastrophic forgetting problem, our model finetune works better for those sequences but forgot how to fold the others.\n\n- The inclusion of ligands can improve the performance in some cases but decrease it in other.\n- We have tried using an embedding instead of the sequence and the result imrpoved for some sequences but decreased for others. The improved one were those were we didn't have a good template but with an embedding we were able to capture some of the structural information.\n\nThose along with other things were discovered thanks to the 3D visualization dashboard. Right now I do not remember much curiosities as I have done a lot of experiments and I have not written down all the insights.\n\n## Approach\n\n### Final Strategy: Best-of-5 Ensemble\n\nOur final submission uses 5 prediction slots strategically:\n\n**Slot 1 - Pure TBM with Biopython Alignment**\n- Align test sequences against all training sequences using Biopython pairwise alignment\n- Minimum 50% identity threshold\n- Linear interpolation for missing coordinates\n- Adaptive constraints: bond distance correction (~5.95 A for adjacent residues), angular constraints (~10.2 A for i+2), Laplacian smoothing, and steric clash prevention\n\n**Slots 2-3 - TBM + RNAPro (Protenix + RibonanzaNet Embeddings)**\n- Weighted template selection: 85% sequence similarity + 15% ligand similarity\n- Also considers number of unique sequences and total chain count\n- TBM coordinates serve as templates for the RNAPro model (NVIDIA Digital Bio, Phase 1 winner approach)\n- Works best for sequences <= 350 nucleotides\n- For longer sequences: window sliding and stitching strategy\n- Slot 3 uses slightly different parameters and uses Slot 1 output as its template\n\n**Slots 4-5 - Protenix 2025 from Scratch**\n- ByteDance's trainable PyTorch reproduction of AlphaFold 3\n- JSON-based input with correct stoichiometry information\n- No ligand input (to avoid OOM errors on large complexes)\n- Window sliding and stitching for sequences > 350 nt\n- Slot 5 is a duplicate of Slot 4 (ran out of time)\n- Slot 4 should be for the base Protenix but OOT issues.\n\nIf you think carefully instead of doing 5 solutions in our case was more like 3.5-4 solutions as the last two were the same coordinates.\n\nOur main focus was to make sure each predictions actually makes sense. a major part of the evaluation was done by me visually inspecting the predictions and comparing them to the ground truth using our 3D visualization dashboard. It is important to develop the intuition of data scientist and to not overfit and try to get as much information as you can about the data.\n\n\n## What We Would Do Differently\n\n- **Protenix finetuning** - I would have loved to fine-tune Protenix but as I am not familiar with the domain I struggle to do it.\n- **More sophisticated ligand weighting** in template selection, that was a basic one we check were working right\n- **Secondary structure integration** - we deprioritized this but it could have improved predictions by restrictions\n- **Larger ensemble diversity** - combining DRFold2, Chai, and Protenix could have improved coverage\n\n\n## Acknowledgments\n\n- Stanford and Kaggle for hosting the competition\n- NVIDIA Digital Bio team for the RNAPro model from Phase 1\n- ByteDance for the open-source Protenix (AlphaFold 3 reproduction)\n- The Kaggle community for forum discussions and shared notebooks\n- My teammates at WhiteBox for their collaboration and support throughout the competition\n\n## Conclusion\nOverall, this competition was an incredible learning experience that pushed us to apply our data science skills in a completely new domain. We had a lot of fun experimenting with different approaches, analyzing the results, and iteratively improving our models. We are proud of our final submission and grateful for the opportunity to participate in such a challenging and rewarding competition. We look forward to seeing how the field of RNA 3D folding continues to evolve and hope that our contributions can help advance the state of the art in this exciting area of research.\n",
      "votes": 8
    },
    {
      "id": 3443399,
      "postDate": "2026-04-16T13:44:06.947Z",
      "content": "<p>it is  helps me to learn data science</p>",
      "rawMarkdown": "it is  helps me to learn data science"
    },
    {
      "id": 3437291,
      "postDate": "2026-04-07T13:09:21.027Z",
      "content": "<p>Our final submission is just the tip of the iceberg. We have extensive logs and code from our experiments, including the dashboard, Protenix fine-tuning results, and other strategies, that didn't make the final cutoff but offer key insights.\nWe are happy to share these to support the CASP17 model synthesis. Feel free to reach out if any of these specific insights could be useful.</p>",
      "rawMarkdown": "Our final submission is just the tip of the iceberg. We have extensive logs and code from our experiments, including the dashboard, Protenix fine-tuning results, and other strategies, that didn't make the final cutoff but offer key insights.\nWe are happy to share these to support the CASP17 model synthesis. Feel free to reach out if any of these specific insights could be useful."
    }
  ],
  "comments": [
    {
      "id": 3443399,
      "author_name": "Suresh Sujan",
      "author_url": "",
      "post_date": "2026-04-16T13:44:06.947000",
      "content": "<p>it is  helps me to learn data science</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3437291,
      "author_name": "PanditaProgrammer",
      "author_url": "",
      "post_date": "2026-04-07T13:09:21.027000",
      "content": "<p>Our final submission is just the tip of the iceberg. We have extensive logs and code from our experiments, including the dashboard, Protenix fine-tuning results, and other strategies, that didn't make the final cutoff but offer key insights.\nWe are happy to share these to support the CASP17 model synthesis. Feel free to reach out if any of these specific insights could be useful.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3437215": "# Stanford RNA 3D Folding Part 2 - Competition Writeup\n\nHi! Thanks to all the team behind the Stanford RNA 3D Folding Part 2 competition. It was a fantastic challenge, and we had a great time learning throughout the process.\n\nI am a student at UPM (Polytechnic University of Madrid), and as this was my first serious Kaggle competition, the learning curve was steep but incredibly rewarding. I am representing [WhiteBox](https://www.whiteboxml.com/), a boutique data science and ML consultancy in Spain where I am currently a Data Scientist Intern.\n\n### Our Philosophy\nOur team is passionate about bridging the gap between research and real-world applications. This competition was the perfect playground to test our skills in a completely new domain; we approached the problem as data scientists with no prior background in biology, which made the experience even more valuable as it allowed us to test our transferable skills when pivoting to unfamiliar fields.\n\n\n## Problem Understanding & Domain Knowledge\nBefore diving into modeling, we spent a significant amount of time understanding the problem and the domain. We read through the competition description, explored the provided datasets, and researched the multiple factors that influence RNA folding, such as sequence similarity, ligand interactions, and the presence other chains. This foundational understanding allowed us to make better decisions and validate our approaches more effectively as we progressed through the competition. \n\n\n## Local Validation Strategy\nSince the competition evaluation is based on TM-score computed via USalign, we used the same metric for out local validation. We though about creating our own train/validation split from the training data, but given that we were not experts in the domain, we were concerned about data leakage and ended up using the default validation set after reading how the split were done. Then we focused on understanding the type of chain we were predicting (monomer, homomer, heteromer, synthetic, the function of that chain, etc.) to try to understand why one method could work and why not.\n\nOne of the keys to our success I would say was the creation of an interactive 3D visualization dashboard that allowed us to visually inspect the predicted vs. ground truth structures. This was crucial for understanding the strengths and weaknesses of our models, especially in cases where the TM-score did not fully capture the quality of the prediction. It helped us identify specific issues such as steric clashes, incorrect chain positioning, and other structural artifacts that we could then address in our modeling approaches.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17092985%2F68d953c2104cf4696868e06849c73049%2Fimage.png?generation=1775558628697326&alt=media)\nAs you can see it gives us a side-by-side comparison of the predicted and true structures, allowing us to rotate and zoom in on specific regions to analyze the quality of the prediction in detail. Instead on using only the TM-score (1D metric), this dashboard provided a more intuitive understanding of the 3D structural differences, which was invaluable for debugging and improving our models.\n\n## The journey: From Bottom to Top 11 to Bottom to Final Top 15\n\nOnce we have gathered domain knowledge and set up our local validation strategy, we started experimenting with different approaches. We tried a wide range of methods, from template-based modeling (TBM) to finetuning Protenix on the competition data, and even explored other models like DRFold2 and Chai. We also experimented with different ways of incorporating ligand information and handling multimers.\n\nThe key here was to understand the strengths and weaknesses of each approach and how they performed on different types of sequences (monomers, homomers, heteromers, synthetic...). We iteratively refined our models based on the insights we gained from our local validation and the 3D visualization dashboard.\n\nWe spend many hours debugging and improving our TBM approach, which ended up being the backbone of our final pipeline. A different approach we found promising was to weigh the importance of the ligand similarity and the chain count in the template selection along with another TBM more similar to the one used by everyone.\n\nThe second approach we have taken was to use RNAPro, the model from NVIDIA Digital Bio that came from mixing the top solutions from Phase 1. Setting them on default parameters and using the TBM output as template for it worked really well for sequences <= 350nt. For longer sequences it was lacking the precision I would like it to have. At the end we used two RNAPros with different template predictions and parameters to increase the diversity of our ensemble. For example if we use more cycles it could benefit the giant RNAs (9LEC, 9LEL) as it would be able to refine the structure more, but it would also decrease the performance for smaller sequences.\n\nWe have tried also DRFold2, Chai-1 and Boltz-1 but I didn't like them as much as RNAPro and Protenix, so we deprioritized them to focus on Protenix to the short time we had left.\n\nProtenix, Protenix, and more Protenix. It is a diffusion model that, depending on its random seed, can produce very different predictions. For those who do not know it, it allows you to input a vast amount of data and use different approaches available to those who have taken the time to read the documentation in the era of chatbots.\n\nI was able to configure them locally without problems but I struggled to upload my solution to Kaggle due to the many things needed for it to function properly. I will share the more valuable insights I got from the predictions on local.\n\nProtenix Versions: I have tested the 4 models from Protenix (the 202560630, the base one, the v0.5 and the v0.2) and each of them has its own strengths and weaknesses.\n\nThe 2025 one at least seems better overall in tm-score for the test compared to the others, but I would like to highlight that for more simple RNAs the base performance is better, and for more complex ones the 2025 one is better.\n\ni have played with all the parameters of Protenix and depending on the parameters it could give us completely different results. \n\n- If we are working with other Proteins and DNA it is better to activate the flag, if not don't activate as it decrease the performance for RNA.\n- For Giant RNAs (9LEC, 9LEL) it is better to use more cycles and more recycles, but for smaller ones it is better to use less cycles and recycles.\n- For viruses don't use Protenix.\n- For monomers Protenix works acceptably.\n- For homomers and heteromers it is bad at position them correctly in space, it is able to predict the structure of each chain really well. For example the 9KGG is a homomultimer of 2 chains, Protenix is able to predict the structure for each chain but no to position them correctly so instead of a maximun tm-score of 1 we could only get a tm-score maximum of 0.5. We thought of ways to normalize the chains and simplify the positioning but you just need to check other difficult samples like 9G4Q or 9WHV to see that there's no unique answer.\n- We know that Protenix does not work well when no template is available, in local we were able to create the fasta files to supply the lack of it but making it work on Kaggle was a nightmare and we didn't have time to do it. \n\nKnowing the weaknessed we thought about finetunning Protenix to make it work with heteromers and homomers, middleway the results were promising but the variance across random seeds was too high for me to feel confident about it, so we ended up not using it for the final submission. \n\nSome interesting images (left is the default and right is the fine-tuned one):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17092985%2F57c736a1ba578f9eef37922638937df5%2Fimage2.png?generation=1775558589328801&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F17092985%2F05d11146e132491d43c47b1234a10d88%2Fimage3.png?generation=1775558613294141&alt=media)\n\nWe have done some tweaks to the Protenix code (changing the loss funtion to a custom derivable approximation of the TM-score, data augmentation...) but we were not able to solve the problem of the high variance across random seeds and with the tight deadline I decided not to pursue it. If you are curious when we fine-tune Protenix as happened before there were no way to deal with the Catastrophic forgetting problem, our model finetune works better for those sequences but forgot how to fold the others.\n\n- The inclusion of ligands can improve the performance in some cases but decrease it in other.\n- We have tried using an embedding instead of the sequence and the result imrpoved for some sequences but decreased for others. The improved one were those were we didn't have a good template but with an embedding we were able to capture some of the structural information.\n\nThose along with other things were discovered thanks to the 3D visualization dashboard. Right now I do not remember much curiosities as I have done a lot of experiments and I have not written down all the insights.\n\n## Approach\n\n### Final Strategy: Best-of-5 Ensemble\n\nOur final submission uses 5 prediction slots strategically:\n\n**Slot 1 - Pure TBM with Biopython Alignment**\n- Align test sequences against all training sequences using Biopython pairwise alignment\n- Minimum 50% identity threshold\n- Linear interpolation for missing coordinates\n- Adaptive constraints: bond distance correction (~5.95 A for adjacent residues), angular constraints (~10.2 A for i+2), Laplacian smoothing, and steric clash prevention\n\n**Slots 2-3 - TBM + RNAPro (Protenix + RibonanzaNet Embeddings)**\n- Weighted template selection: 85% sequence similarity + 15% ligand similarity\n- Also considers number of unique sequences and total chain count\n- TBM coordinates serve as templates for the RNAPro model (NVIDIA Digital Bio, Phase 1 winner approach)\n- Works best for sequences <= 350 nucleotides\n- For longer sequences: window sliding and stitching strategy\n- Slot 3 uses slightly different parameters and uses Slot 1 output as its template\n\n**Slots 4-5 - Protenix 2025 from Scratch**\n- ByteDance's trainable PyTorch reproduction of AlphaFold 3\n- JSON-based input with correct stoichiometry information\n- No ligand input (to avoid OOM errors on large complexes)\n- Window sliding and stitching for sequences > 350 nt\n- Slot 5 is a duplicate of Slot 4 (ran out of time)\n- Slot 4 should be for the base Protenix but OOT issues.\n\nIf you think carefully instead of doing 5 solutions in our case was more like 3.5-4 solutions as the last two were the same coordinates.\n\nOur main focus was to make sure each predictions actually makes sense. a major part of the evaluation was done by me visually inspecting the predictions and comparing them to the ground truth using our 3D visualization dashboard. It is important to develop the intuition of data scientist and to not overfit and try to get as much information as you can about the data.\n\n\n## What We Would Do Differently\n\n- **Protenix finetuning** - I would have loved to fine-tune Protenix but as I am not familiar with the domain I struggle to do it.\n- **More sophisticated ligand weighting** in template selection, that was a basic one we check were working right\n- **Secondary structure integration** - we deprioritized this but it could have improved predictions by restrictions\n- **Larger ensemble diversity** - combining DRFold2, Chai, and Protenix could have improved coverage\n\n\n## Acknowledgments\n\n- Stanford and Kaggle for hosting the competition\n- NVIDIA Digital Bio team for the RNAPro model from Phase 1\n- ByteDance for the open-source Protenix (AlphaFold 3 reproduction)\n- The Kaggle community for forum discussions and shared notebooks\n- My teammates at WhiteBox for their collaboration and support throughout the competition\n\n## Conclusion\nOverall, this competition was an incredible learning experience that pushed us to apply our data science skills in a completely new domain. We had a lot of fun experimenting with different approaches, analyzing the results, and iteratively improving our models. We are proud of our final submission and grateful for the opportunity to participate in such a challenging and rewarding competition. We look forward to seeing how the field of RNA 3D folding continues to evolve and hope that our contributions can help advance the state of the art in this exciting area of research.\n",
    "3443399": "it is  helps me to learn data science",
    "3437291": "Our final submission is just the tip of the iceberg. We have extensive logs and code from our experiments, including the dashboard, Protenix fine-tuning results, and other strategies, that didn't make the final cutoff but offer key insights.\nWe are happy to share these to support the CASP17 model synthesis. Feel free to reach out if any of these specific insights could be useful."
  }
}