{
  "id": 687191,
  "title": "Some Thoughts on the Rerun and Ranks Discrepancies ",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/687191",
  "author_name": "Theo Viel",
  "post_date": "2026-04-02T16:03:24.789000",
  "votes": 11,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Few comments as an experienced Kaggler. I was partly involved with the competition organization but I do not have all the details.\nI do not have access to the private sequences (nor the public ones actually), but I have a more global understanding thanks to the work I did on RNAPro (which is <a href=\"https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference\" target=\"_blank\">public</a> !).</p>\n<p>These are my thoughts and there could be some inaccuracies. Nonetheless, I hope these can help understanding what happened regarding your scores.</p>\n<h2>The rerun</h2>\n<p>There was a rerun, similar to the reruns that happen in forecasting competitions (<a href=\"https://www.kaggle.com/competitions/predict-energy-behavior-of-prosumers/leaderboard\" target=\"_blank\">example</a>, <a href=\"https://www.kaggle.com/competitions/dfl-bundesliga-data-shootout\" target=\"_blank\">other example</a>).</p>\n<p>Which means the private leaderboard gets updated, and that both public and temporary-private leaderboard were populated with data <strong>not necessarily</strong> representative of what will be in the final-private dataset.</p>\n<p>These datasets are populated to provide Kagglers an environment to debug their submissions against. They usually are from the same data source, and for tasks where time is important, hold <strong>weaker guarantees</strong> than the real test data. Public LB can often be used as a way to compare yourself with other teams, but the best performance indicator should (most of the time) be your local validation.\nFor forecasting competitions, it is important to update the hidden test dataset with unseen data, to ensure there is no leakage in the test evaluation. Kaggle competitions should sustain the best standard regarding model evaluation.</p>\n<p>It is the same with RNA competitions: new targets get released as they get available.\nLast year the dataset was curated after the end of the competition. This year it was curated during the competition to make the experience better for Kagglers (having to wait 3 months for the results is not ideal), and because the hypothesis that these sequence will not leak into the training sets during the competition timeline is strong enough.</p>\n<p><strong>Because private LB data is collected after the public LB data, a shift has to be expected.</strong></p>\n<p>Usually the public LB is discarded (to save costs or for simplicity maybe), but here it is still available which I think is a better choice, although it can be confusing.\nI'm 90% sure it contains the same sequences as before - what should explain public LB variations is that</p>\n<ul>\n<li>Public LB only contains the scores of the 2 (or 1) models selected as final submission</li>\n<li>Randomness, sequence order, and other not-obvious random things such as hardware variation</li>\n<li>The MSA directory was updated, and maybe some other files as well (?)</li>\n</ul>\n<p>If there is an issue with the Kaggle rerun (which is not impossible), then let's hope it does not effect private LB standings which are what matters.</p>\n<p>Note that there are new answers from the hosts in <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686938\" target=\"_blank\">this thread</a>. </p>\n<h2>Performance drops</h2>\n<p>In a lot of competitions, you will notice performance shifts between public and private (<a href=\"https://www.kaggle.com/competitions/birdclef-2024\" target=\"_blank\">example</a>, <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection\" target=\"_blank\">example</a>). This happens a lot because:</p>\n<ul>\n<li>Private data can be harder, whether it is intended or not</li>\n<li>Public kernels are often overfitted to public data</li>\n</ul>\n<p>As a competitor, it is one of your goals to ensure robustness against public / private variations. Note that hosts are encouraged to not share information about the private data by the Kaggle team, hence the oftentimes observed lack of transparency. It is better not to share than to share information that will ruin the competition setup.</p>\n<p>The way I see it as a veteran Kaggler is that it's my goal to gather all the information I can to solve the private LB puzzle.\nSometimes the puzzle is trivial and private LB performance is stable. Sometimes the hosts do something unpredictable. Most of the time it's something in-between, and the solution can be figured using a mix of EDA, public score analysis, and carefully reading all information available.</p>\n<p>Not figuring out this puzzle can be frustrating - and I can ensure you that noone has a 100% success rate at that.</p>\n<h2>Solving this year's puzzle</h2>\n<p>From the competition page :</p>\n<blockquote>\n  <p>Now, you’ll face even more complex targets, including ones with no structural templates […]</p>\n</blockquote>\n<p>This is a big hint towards what your model should be robust against. You should build a solution that's good when there are no strong templates. Which means template-based approaches will very likely not be enough for the task. </p>\n<p>However, template based approaches worked well on the public dataset. In fact, templates found by top public kernels were so strong that plugging <a href=\"https://www.kaggle.com/code/jaejohn/rnapro-inference-with-tbm\" target=\"_blank\">RNAPro</a> on top of them did not bring any improvement. This is an indicator that public LB sequences might not meet the \"no strong template\" criterion, or that public kernels performance is masking improvements. And both are hints that a shake-up is likely to happen when newer sequences are collected.</p>\n<p>This is further amplified by the fact that public kernels are crowded with TBM solutions tweaked on the 3rd public LB digit, and that the public test set is quite small. A solution ranked 100th actually performs very similarly to a solution ranked 1000th. This is especially true for the public LB, but holds for the private LB since solutions adapted from public kernels will perform similarly. From then it is a coin flip that decides whether you get a silver medal or end up 500th.</p>\n<p>Note that many teams did a brilliant job at addressing this challenge. For instance 6 of the top 8 teams stayed in the gold zone, and last year's 1st and 2nd place also grabed a gold medal this year. </p>\n<h2>TLDR</h2>\n<ul>\n<li>There was a rerun, only scores of submissions that were selected are relevant.</li>\n<li>Public LB contains the same sequences as before, but only shows scores of your 2 selected submissions.</li>\n<li>Private LB contains newer sequences, on which widely used public template-based approaches turned out to be weaker. I do not think this is intended, but there were hints this was going to happen (the general expected difficulty was stated in the competition description page, and public notebook were overfitted). It is one of your (toughest !) challenges as a competitor to develop solutions robust against public/private shift.</li>\n</ul>",
  "messages": [
    {
      "id": 3434110,
      "postDate": "2026-04-02T16:03:24.790Z",
      "content": "<p>Few comments as an experienced Kaggler. I was partly involved with the competition organization but I do not have all the details.\nI do not have access to the private sequences (nor the public ones actually), but I have a more global understanding thanks to the work I did on RNAPro (which is <a href=\"https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference\" target=\"_blank\">public</a> !).</p>\n<p>These are my thoughts and there could be some inaccuracies. Nonetheless, I hope these can help understanding what happened regarding your scores.</p>\n<h2>The rerun</h2>\n<p>There was a rerun, similar to the reruns that happen in forecasting competitions (<a href=\"https://www.kaggle.com/competitions/predict-energy-behavior-of-prosumers/leaderboard\" target=\"_blank\">example</a>, <a href=\"https://www.kaggle.com/competitions/dfl-bundesliga-data-shootout\" target=\"_blank\">other example</a>).</p>\n<p>Which means the private leaderboard gets updated, and that both public and temporary-private leaderboard were populated with data <strong>not necessarily</strong> representative of what will be in the final-private dataset.</p>\n<p>These datasets are populated to provide Kagglers an environment to debug their submissions against. They usually are from the same data source, and for tasks where time is important, hold <strong>weaker guarantees</strong> than the real test data. Public LB can often be used as a way to compare yourself with other teams, but the best performance indicator should (most of the time) be your local validation.\nFor forecasting competitions, it is important to update the hidden test dataset with unseen data, to ensure there is no leakage in the test evaluation. Kaggle competitions should sustain the best standard regarding model evaluation.</p>\n<p>It is the same with RNA competitions: new targets get released as they get available.\nLast year the dataset was curated after the end of the competition. This year it was curated during the competition to make the experience better for Kagglers (having to wait 3 months for the results is not ideal), and because the hypothesis that these sequence will not leak into the training sets during the competition timeline is strong enough.</p>\n<p><strong>Because private LB data is collected after the public LB data, a shift has to be expected.</strong></p>\n<p>Usually the public LB is discarded (to save costs or for simplicity maybe), but here it is still available which I think is a better choice, although it can be confusing.\nI'm 90% sure it contains the same sequences as before - what should explain public LB variations is that</p>\n<ul>\n<li>Public LB only contains the scores of the 2 (or 1) models selected as final submission</li>\n<li>Randomness, sequence order, and other not-obvious random things such as hardware variation</li>\n<li>The MSA directory was updated, and maybe some other files as well (?)</li>\n</ul>\n<p>If there is an issue with the Kaggle rerun (which is not impossible), then let's hope it does not effect private LB standings which are what matters.</p>\n<p>Note that there are new answers from the hosts in <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686938\" target=\"_blank\">this thread</a>. </p>\n<h2>Performance drops</h2>\n<p>In a lot of competitions, you will notice performance shifts between public and private (<a href=\"https://www.kaggle.com/competitions/birdclef-2024\" target=\"_blank\">example</a>, <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection\" target=\"_blank\">example</a>). This happens a lot because:</p>\n<ul>\n<li>Private data can be harder, whether it is intended or not</li>\n<li>Public kernels are often overfitted to public data</li>\n</ul>\n<p>As a competitor, it is one of your goals to ensure robustness against public / private variations. Note that hosts are encouraged to not share information about the private data by the Kaggle team, hence the oftentimes observed lack of transparency. It is better not to share than to share information that will ruin the competition setup.</p>\n<p>The way I see it as a veteran Kaggler is that it's my goal to gather all the information I can to solve the private LB puzzle.\nSometimes the puzzle is trivial and private LB performance is stable. Sometimes the hosts do something unpredictable. Most of the time it's something in-between, and the solution can be figured using a mix of EDA, public score analysis, and carefully reading all information available.</p>\n<p>Not figuring out this puzzle can be frustrating - and I can ensure you that noone has a 100% success rate at that.</p>\n<h2>Solving this year's puzzle</h2>\n<p>From the competition page :</p>\n<blockquote>\n  <p>Now, you’ll face even more complex targets, including ones with no structural templates […]</p>\n</blockquote>\n<p>This is a big hint towards what your model should be robust against. You should build a solution that's good when there are no strong templates. Which means template-based approaches will very likely not be enough for the task. </p>\n<p>However, template based approaches worked well on the public dataset. In fact, templates found by top public kernels were so strong that plugging <a href=\"https://www.kaggle.com/code/jaejohn/rnapro-inference-with-tbm\" target=\"_blank\">RNAPro</a> on top of them did not bring any improvement. This is an indicator that public LB sequences might not meet the \"no strong template\" criterion, or that public kernels performance is masking improvements. And both are hints that a shake-up is likely to happen when newer sequences are collected.</p>\n<p>This is further amplified by the fact that public kernels are crowded with TBM solutions tweaked on the 3rd public LB digit, and that the public test set is quite small. A solution ranked 100th actually performs very similarly to a solution ranked 1000th. This is especially true for the public LB, but holds for the private LB since solutions adapted from public kernels will perform similarly. From then it is a coin flip that decides whether you get a silver medal or end up 500th.</p>\n<p>Note that many teams did a brilliant job at addressing this challenge. For instance 6 of the top 8 teams stayed in the gold zone, and last year's 1st and 2nd place also grabed a gold medal this year. </p>\n<h2>TLDR</h2>\n<ul>\n<li>There was a rerun, only scores of submissions that were selected are relevant.</li>\n<li>Public LB contains the same sequences as before, but only shows scores of your 2 selected submissions.</li>\n<li>Private LB contains newer sequences, on which widely used public template-based approaches turned out to be weaker. I do not think this is intended, but there were hints this was going to happen (the general expected difficulty was stated in the competition description page, and public notebook were overfitted). It is one of your (toughest !) challenges as a competitor to develop solutions robust against public/private shift.</li>\n</ul>",
      "rawMarkdown": "Few comments as an experienced Kaggler. I was partly involved with the competition organization but I do not have all the details.\nI do not have access to the private sequences (nor the public ones actually), but I have a more global understanding thanks to the work I did on RNAPro (which is [public](https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference) !).\n\nThese are my thoughts and there could be some inaccuracies. Nonetheless, I hope these can help understanding what happened regarding your scores.\n\n## The rerun\n\nThere was a rerun, similar to the reruns that happen in forecasting competitions ([example](https://www.kaggle.com/competitions/predict-energy-behavior-of-prosumers/leaderboard), [other example](https://www.kaggle.com/competitions/dfl-bundesliga-data-shootout)).\n\nWhich means the private leaderboard gets updated, and that both public and temporary-private leaderboard were populated with data **not necessarily** representative of what will be in the final-private dataset.\n\nThese datasets are populated to provide Kagglers an environment to debug their submissions against. They usually are from the same data source, and for tasks where time is important, hold **weaker guarantees** than the real test data. Public LB can often be used as a way to compare yourself with other teams, but the best performance indicator should (most of the time) be your local validation.\nFor forecasting competitions, it is important to update the hidden test dataset with unseen data, to ensure there is no leakage in the test evaluation. Kaggle competitions should sustain the best standard regarding model evaluation.\n\nIt is the same with RNA competitions: new targets get released as they get available.\nLast year the dataset was curated after the end of the competition. This year it was curated during the competition to make the experience better for Kagglers (having to wait 3 months for the results is not ideal), and because the hypothesis that these sequence will not leak into the training sets during the competition timeline is strong enough.\n\n**Because private LB data is collected after the public LB data, a shift has to be expected.**\n\nUsually the public LB is discarded (to save costs or for simplicity maybe), but here it is still available which I think is a better choice, although it can be confusing.\nI'm 90% sure it contains the same sequences as before - what should explain public LB variations is that\n- Public LB only contains the scores of the 2 (or 1) models selected as final submission\n- Randomness, sequence order, and other not-obvious random things such as hardware variation\n- The MSA directory was updated, and maybe some other files as well (?)\n\nIf there is an issue with the Kaggle rerun (which is not impossible), then let's hope it does not effect private LB standings which are what matters.\n\nNote that there are new answers from the hosts in [this thread](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686938). \n\n## Performance drops\n\nIn a lot of competitions, you will notice performance shifts between public and private ([example](https://www.kaggle.com/competitions/birdclef-2024), [example](https://www.kaggle.com/competitions/rsna-breast-cancer-detection)). This happens a lot because:\n- Private data can be harder, whether it is intended or not\n- Public kernels are often overfitted to public data\n\nAs a competitor, it is one of your goals to ensure robustness against public / private variations. Note that hosts are encouraged to not share information about the private data by the Kaggle team, hence the oftentimes observed lack of transparency. It is better not to share than to share information that will ruin the competition setup.\n\nThe way I see it as a veteran Kaggler is that it's my goal to gather all the information I can to solve the private LB puzzle.\nSometimes the puzzle is trivial and private LB performance is stable. Sometimes the hosts do something unpredictable. Most of the time it's something in-between, and the solution can be figured using a mix of EDA, public score analysis, and carefully reading all information available.\n\nNot figuring out this puzzle can be frustrating - and I can ensure you that noone has a 100% success rate at that.\n\n## Solving this year's puzzle\n\nFrom the competition page :\n\n> Now, you’ll face even more complex targets, including ones with no structural templates [...]\n\nThis is a big hint towards what your model should be robust against. You should build a solution that's good when there are no strong templates. Which means template-based approaches will very likely not be enough for the task. \n\nHowever, template based approaches worked well on the public dataset. In fact, templates found by top public kernels were so strong that plugging [RNAPro](https://www.kaggle.com/code/jaejohn/rnapro-inference-with-tbm) on top of them did not bring any improvement. This is an indicator that public LB sequences might not meet the \"no strong template\" criterion, or that public kernels performance is masking improvements. And both are hints that a shake-up is likely to happen when newer sequences are collected.\n\nThis is further amplified by the fact that public kernels are crowded with TBM solutions tweaked on the 3rd public LB digit, and that the public test set is quite small. A solution ranked 100th actually performs very similarly to a solution ranked 1000th. This is especially true for the public LB, but holds for the private LB since solutions adapted from public kernels will perform similarly. From then it is a coin flip that decides whether you get a silver medal or end up 500th.\n\nNote that many teams did a brilliant job at addressing this challenge. For instance 6 of the top 8 teams stayed in the gold zone, and last year's 1st and 2nd place also grabed a gold medal this year. \n\n## TLDR\n\n- There was a rerun, only scores of submissions that were selected are relevant.\n- Public LB contains the same sequences as before, but only shows scores of your 2 selected submissions.\n- Private LB contains newer sequences, on which widely used public template-based approaches turned out to be weaker. I do not think this is intended, but there were hints this was going to happen (the general expected difficulty was stated in the competition description page, and public notebook were overfitted). It is one of your (toughest !) challenges as a competitor to develop solutions robust against public/private shift.",
      "votes": 11
    },
    {
      "id": 3434646,
      "postDate": "2026-04-03T04:49:15.877Z",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> Thanks for the post. </p>\n<p>And everyone else: thanks for flagging these complexities. </p>\n<p>This Stanford RNA 3D Folding Part 2 competition was meant to be quite similar to Part 1, and our checks before release did not reveal these complexities. </p>\n<p>We competition hosts and the Kaggle developers and competition hosts were caught off guard, and we could have done better on responding earlier that we'd need time and your help to track down.</p>\n<p>For a couple of Theo's points above, the Kaggle devs have agreed to let me describe some additional results that we hope can illustrate what happened.</p>\n<p><strong>Public LB variations</strong></p>\n<p>It was a surprise to many participants that their Public LB rankings shifted. </p>\n<p>The rerun of notebooks on those Public LB targets was unavoidable with current Kaggle infrastructure, but we didn't expect such a disruption in ranking. </p>\n<p>Here's why –&nbsp;the scores of notebooks actually changed very little: </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5659684%2F269d9dbbbbe339d2a85c972ed7455fef%2Fkaggle_part2_pub_origpub_scatterplot.png?generation=1775187253196596&amp;alt=media\" alt=\"\"></p>\n<p>This plot contains all the original and final Public LB scores for the two selected notebooks for each team. The reproducibility seemed very good, e.g., the mean absolute deviation between scores in the two runs was low (~0.007).  The Pearson correlation was high (r=0.997). This was similar or actually slightly better than the reproducibility in notebook scores that we saw in Part 1….</p>\n<p>However, there was a crowding of the scores, boxed in red, whose magnitude we didn't notice due to the overlap of points. As <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> and others have pointed out, random fluctuations produced small shifts in the score that then led to extremely large changes in the <em>rankings</em> within this crowd. </p>\n<p>The hosts should have realized this ahead of time. But at least I (now host of 4 Kaggle competitions) had not seen such level of crowding before, even in Part 1 of this competition. </p>\n<p>We are still very interested in how some teams were able to avoid this shakeup. The solution writeups so far have been illuminating -- some top teams set up their own internal validation sets and explored strategies distinct from the most widely shared public notebooks. We look forward to more writeups, more discussion, and more learning. </p>\n<p>And if there are future RNA competitions, we'll need to find a way to prevent reruns of the Public LB.</p>\n<p><strong>Performance drops</strong></p>\n<p>As mentioned elsewhere, it was not known to us ahead of time whether the final private leaderboard would be easier or harder than the Public leaderboard. In the end, it turned out the <code>best_template_oracle</code> score was lower in Private than Public.  </p>\n<p>But an important thing to point out is that many notebooks got <em>higher</em> scores on the Private leaderboard than Public:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5659684%2F091499c8797533e19b7fbcc7d352944c%2Fkaggle_part2_score_pub_priv_scatterplot.png?generation=1775188453951595&amp;alt=media\" alt=\"\"> </p>\n<p>In terms of raw scores, the Pearson correlation between Private and Public leaderboard scores was actually pretty good (r=0.955). However, we overlooked the same crowding phenomenon (red box) and did not anticipate the subsequent shakeup of rankings. </p>\n<p>It will be interesting to again hear from the teams that managed to retain good scores in both Public and Private leaderboards. Also if Kagglers are familiar with other competitions that had this same crowding, please do post.</p>\n<p><strong>Placeholder Private leaderboard</strong></p>\n<p>Although this was not mentioned in <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>'s post, others, including <a href=\"https://www.kaggle.com/hongan\" target=\"_blank\">@hongan</a> <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> in this thread, have noted that it was very confusing to have released the scores for an internal 'placeholder' Private LB test set that was updated at the end with a larger test set.  </p>\n<p>The same thing actually happened in Part 1, but the placeholder Private LB set happened to give lower scores than the final Private LB, so not too many people noted the change. If there are future RNA competitions, hosts and devs now know that we should communicate better about any placeholder Private LB sets -- or avoid them altogether.</p>\n<p>Despite all of this, we hope everyone has learned a lot from their experience and from each other in this competition. The scientific impact of this competition's methods for RNA structure prediction in medicine and biology are likely to be quite high. </p>\n<p>And while we thought this was going to be a straightforward sequel to Part 1 of the Stanford RNA Folding Challenge, it turned out to be more complex -- it makes us hosts appreciate even more the ingenuity of the teams who 'solved the puzzle'.</p>",
      "rawMarkdown": "@theoviel Thanks for the post. \n\nAnd everyone else: thanks for flagging these complexities. \n\nThis Stanford RNA 3D Folding Part 2 competition was meant to be quite similar to Part 1, and our checks before release did not reveal these complexities. \n\nWe competition hosts and the Kaggle developers and competition hosts were caught off guard, and we could have done better on responding earlier that we'd need time and your help to track down.\n\nFor a couple of Theo's points above, the Kaggle devs have agreed to let me describe some additional results that we hope can illustrate what happened.\n\n**Public LB variations**\n\nIt was a surprise to many participants that their Public LB rankings shifted. \n\nThe rerun of notebooks on those Public LB targets was unavoidable with current Kaggle infrastructure, but we didn't expect such a disruption in ranking. \n\nHere's why – the scores of notebooks actually changed very little: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5659684%2F269d9dbbbbe339d2a85c972ed7455fef%2Fkaggle_part2_pub_origpub_scatterplot.png?generation=1775187253196596&alt=media)\n\nThis plot contains all the original and final Public LB scores for the two selected notebooks for each team. The reproducibility seemed very good, e.g., the mean absolute deviation between scores in the two runs was low (~0.007).  The Pearson correlation was high (r=0.997). This was similar or actually slightly better than the reproducibility in notebook scores that we saw in Part 1....\n\nHowever, there was a crowding of the scores, boxed in red, whose magnitude we didn't notice due to the overlap of points. As @theoviel and others have pointed out, random fluctuations produced small shifts in the score that then led to extremely large changes in the *rankings* within this crowd. \n\nThe hosts should have realized this ahead of time. But at least I (now host of 4 Kaggle competitions) had not seen such level of crowding before, even in Part 1 of this competition. \n\nWe are still very interested in how some teams were able to avoid this shakeup. The solution writeups so far have been illuminating -- some top teams set up their own internal validation sets and explored strategies distinct from the most widely shared public notebooks. We look forward to more writeups, more discussion, and more learning. \n\nAnd if there are future RNA competitions, we'll need to find a way to prevent reruns of the Public LB.\n\n**Performance drops**\n\nAs mentioned elsewhere, it was not known to us ahead of time whether the final private leaderboard would be easier or harder than the Public leaderboard. In the end, it turned out the `best_template_oracle` score was lower in Private than Public.  \n\nBut an important thing to point out is that many notebooks got *higher* scores on the Private leaderboard than Public:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5659684%2F091499c8797533e19b7fbcc7d352944c%2Fkaggle_part2_score_pub_priv_scatterplot.png?generation=1775188453951595&alt=media) \n\nIn terms of raw scores, the Pearson correlation between Private and Public leaderboard scores was actually pretty good (r=0.955). However, we overlooked the same crowding phenomenon (red box) and did not anticipate the subsequent shakeup of rankings. \n\nIt will be interesting to again hear from the teams that managed to retain good scores in both Public and Private leaderboards. Also if Kagglers are familiar with other competitions that had this same crowding, please do post.\n\n**Placeholder Private leaderboard**\n\nAlthough this was not mentioned in @theoviel's post, others, including @hongan @sheriytm in this thread, have noted that it was very confusing to have released the scores for an internal 'placeholder' Private LB test set that was updated at the end with a larger test set.  \n\nThe same thing actually happened in Part 1, but the placeholder Private LB set happened to give lower scores than the final Private LB, so not too many people noted the change. If there are future RNA competitions, hosts and devs now know that we should communicate better about any placeholder Private LB sets -- or avoid them altogether.\n\nDespite all of this, we hope everyone has learned a lot from their experience and from each other in this competition. The scientific impact of this competition's methods for RNA structure prediction in medicine and biology are likely to be quite high. \n\nAnd while we thought this was going to be a straightforward sequel to Part 1 of the Stanford RNA Folding Challenge, it turned out to be more complex -- it makes us hosts appreciate even more the ingenuity of the teams who 'solved the puzzle'.\n",
      "votes": 6,
      "replies": [
        {
          "id": 3434709,
          "postDate": "2026-04-03T06:27:13.853Z",
          "content": "<p>Thanks for chiming in! love the clarity here. I'm going to work on my solution writeup tomorrow!</p>",
          "rawMarkdown": "Thanks for chiming in! love the clarity here. I'm going to work on my solution writeup tomorrow!"
        },
        {
          "id": 3435188,
          "postDate": "2026-04-03T19:30:44.883Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3435189,
          "postDate": "2026-04-03T19:31:27.677Z",
          "content": "<p>I'm devastated! <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F59648%2F7b3daef3c0b2d332f688a313ad47890d%2FScreenshot%202026-04-03%20202723.png?generation=1775244680251238&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "I'm devastated! ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F59648%2F7b3daef3c0b2d332f688a313ad47890d%2FScreenshot%202026-04-03%20202723.png?generation=1775244680251238&alt=media)"
        }
      ]
    },
    {
      "id": 3434119,
      "postDate": "2026-04-02T16:18:48.913Z",
      "content": "<p>The confusion comes from that </p>\n<ul>\n<li>public LB was cleared and when it was then revealed, the scores are not same (host explains that because the order of the sequences has changed, and that is why the public LB score has changed).</li>\n<li>there are two private scores. I understand that there was an update on private test set. But still usually in rerun you only see the final private score, not two different scores with huge differences.</li>\n</ul>\n<p>It is not common to see these practices in competitions. So naturally questions were asked! But now I think it is explained by the hosts in a more transparent way that at least cleared all my doubt. I just hope they can explain what happened in the first place.</p>",
      "rawMarkdown": "The confusion comes from that \n\n- public LB was cleared and when it was then revealed, the scores are not same (host explains that because the order of the sequences has changed, and that is why the public LB score has changed).\n- there are two private scores. I understand that there was an update on private test set. But still usually in rerun you only see the final private score, not two different scores with huge differences.\n\nIt is not common to see these practices in competitions. So naturally questions were asked! But now I think it is explained by the hosts in a more transparent way that at least cleared all my doubt. I just hope they can explain what happened in the first place.",
      "votes": 3
    },
    {
      "id": 3434358,
      "postDate": "2026-04-02T21:13:36.913Z",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>, do you really think by them ignoring answering our genuine be it uncomfortable questions, they are doing us as a community and Kaggle's reputation a favor?</p>",
      "rawMarkdown": "@theoviel, do you really think by them ignoring answering our genuine be it uncomfortable questions, they are doing us as a community and Kaggle's reputation a favor?",
      "votes": -1
    }
  ],
  "comments": [
    {
      "id": 3434646,
      "author_name": "Rhiju Das",
      "author_url": "",
      "post_date": "2026-04-03T04:49:15.877000",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> Thanks for the post. </p>\n<p>And everyone else: thanks for flagging these complexities. </p>\n<p>This Stanford RNA 3D Folding Part 2 competition was meant to be quite similar to Part 1, and our checks before release did not reveal these complexities. </p>\n<p>We competition hosts and the Kaggle developers and competition hosts were caught off guard, and we could have done better on responding earlier that we'd need time and your help to track down.</p>\n<p>For a couple of Theo's points above, the Kaggle devs have agreed to let me describe some additional results that we hope can illustrate what happened.</p>\n<p><strong>Public LB variations</strong></p>\n<p>It was a surprise to many participants that their Public LB rankings shifted. </p>\n<p>The rerun of notebooks on those Public LB targets was unavoidable with current Kaggle infrastructure, but we didn't expect such a disruption in ranking. </p>\n<p>Here's why –&nbsp;the scores of notebooks actually changed very little: </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5659684%2F269d9dbbbbe339d2a85c972ed7455fef%2Fkaggle_part2_pub_origpub_scatterplot.png?generation=1775187253196596&amp;alt=media\" alt=\"\"></p>\n<p>This plot contains all the original and final Public LB scores for the two selected notebooks for each team. The reproducibility seemed very good, e.g., the mean absolute deviation between scores in the two runs was low (~0.007).  The Pearson correlation was high (r=0.997). This was similar or actually slightly better than the reproducibility in notebook scores that we saw in Part 1….</p>\n<p>However, there was a crowding of the scores, boxed in red, whose magnitude we didn't notice due to the overlap of points. As <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> and others have pointed out, random fluctuations produced small shifts in the score that then led to extremely large changes in the <em>rankings</em> within this crowd. </p>\n<p>The hosts should have realized this ahead of time. But at least I (now host of 4 Kaggle competitions) had not seen such level of crowding before, even in Part 1 of this competition. </p>\n<p>We are still very interested in how some teams were able to avoid this shakeup. The solution writeups so far have been illuminating -- some top teams set up their own internal validation sets and explored strategies distinct from the most widely shared public notebooks. We look forward to more writeups, more discussion, and more learning. </p>\n<p>And if there are future RNA competitions, we'll need to find a way to prevent reruns of the Public LB.</p>\n<p><strong>Performance drops</strong></p>\n<p>As mentioned elsewhere, it was not known to us ahead of time whether the final private leaderboard would be easier or harder than the Public leaderboard. In the end, it turned out the <code>best_template_oracle</code> score was lower in Private than Public.  </p>\n<p>But an important thing to point out is that many notebooks got <em>higher</em> scores on the Private leaderboard than Public:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5659684%2F091499c8797533e19b7fbcc7d352944c%2Fkaggle_part2_score_pub_priv_scatterplot.png?generation=1775188453951595&amp;alt=media\" alt=\"\"> </p>\n<p>In terms of raw scores, the Pearson correlation between Private and Public leaderboard scores was actually pretty good (r=0.955). However, we overlooked the same crowding phenomenon (red box) and did not anticipate the subsequent shakeup of rankings. </p>\n<p>It will be interesting to again hear from the teams that managed to retain good scores in both Public and Private leaderboards. Also if Kagglers are familiar with other competitions that had this same crowding, please do post.</p>\n<p><strong>Placeholder Private leaderboard</strong></p>\n<p>Although this was not mentioned in <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>'s post, others, including <a href=\"https://www.kaggle.com/hongan\" target=\"_blank\">@hongan</a> <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> in this thread, have noted that it was very confusing to have released the scores for an internal 'placeholder' Private LB test set that was updated at the end with a larger test set.  </p>\n<p>The same thing actually happened in Part 1, but the placeholder Private LB set happened to give lower scores than the final Private LB, so not too many people noted the change. If there are future RNA competitions, hosts and devs now know that we should communicate better about any placeholder Private LB sets -- or avoid them altogether.</p>\n<p>Despite all of this, we hope everyone has learned a lot from their experience and from each other in this competition. The scientific impact of this competition's methods for RNA structure prediction in medicine and biology are likely to be quite high. </p>\n<p>And while we thought this was going to be a straightforward sequel to Part 1 of the Stanford RNA Folding Challenge, it turned out to be more complex -- it makes us hosts appreciate even more the ingenuity of the teams who 'solved the puzzle'.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 3434709,
          "author_name": "hongan",
          "author_url": "",
          "post_date": "2026-04-03T06:27:13.853000",
          "content": "<p>Thanks for chiming in! love the clarity here. I'm going to work on my solution writeup tomorrow!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3435188,
          "author_name": "",
          "author_url": "",
          "post_date": "2026-04-03T19:30:44.883000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3435189,
          "author_name": "epigene",
          "author_url": "",
          "post_date": "2026-04-03T19:31:27.677000",
          "content": "<p>I'm devastated! <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F59648%2F7b3daef3c0b2d332f688a313ad47890d%2FScreenshot%202026-04-03%20202723.png?generation=1775244680251238&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3434119,
      "author_name": "hongan",
      "author_url": "",
      "post_date": "2026-04-02T16:18:48.913000",
      "content": "<p>The confusion comes from that </p>\n<ul>\n<li>public LB was cleared and when it was then revealed, the scores are not same (host explains that because the order of the sequences has changed, and that is why the public LB score has changed).</li>\n<li>there are two private scores. I understand that there was an update on private test set. But still usually in rerun you only see the final private score, not two different scores with huge differences.</li>\n</ul>\n<p>It is not common to see these practices in competitions. So naturally questions were asked! But now I think it is explained by the hosts in a more transparent way that at least cleared all my doubt. I just hope they can explain what happened in the first place.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3434358,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2026-04-02T21:13:36.913000",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a>, do you really think by them ignoring answering our genuine be it uncomfortable questions, they are doing us as a community and Kaggle's reputation a favor?</p>",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3434110": "Few comments as an experienced Kaggler. I was partly involved with the competition organization but I do not have all the details.\nI do not have access to the private sequences (nor the public ones actually), but I have a more global understanding thanks to the work I did on RNAPro (which is [public](https://www.kaggle.com/code/theoviel/stanford-rna-3d-folding-pt2-rnapro-inference) !).\n\nThese are my thoughts and there could be some inaccuracies. Nonetheless, I hope these can help understanding what happened regarding your scores.\n\n## The rerun\n\nThere was a rerun, similar to the reruns that happen in forecasting competitions ([example](https://www.kaggle.com/competitions/predict-energy-behavior-of-prosumers/leaderboard), [other example](https://www.kaggle.com/competitions/dfl-bundesliga-data-shootout)).\n\nWhich means the private leaderboard gets updated, and that both public and temporary-private leaderboard were populated with data **not necessarily** representative of what will be in the final-private dataset.\n\nThese datasets are populated to provide Kagglers an environment to debug their submissions against. They usually are from the same data source, and for tasks where time is important, hold **weaker guarantees** than the real test data. Public LB can often be used as a way to compare yourself with other teams, but the best performance indicator should (most of the time) be your local validation.\nFor forecasting competitions, it is important to update the hidden test dataset with unseen data, to ensure there is no leakage in the test evaluation. Kaggle competitions should sustain the best standard regarding model evaluation.\n\nIt is the same with RNA competitions: new targets get released as they get available.\nLast year the dataset was curated after the end of the competition. This year it was curated during the competition to make the experience better for Kagglers (having to wait 3 months for the results is not ideal), and because the hypothesis that these sequence will not leak into the training sets during the competition timeline is strong enough.\n\n**Because private LB data is collected after the public LB data, a shift has to be expected.**\n\nUsually the public LB is discarded (to save costs or for simplicity maybe), but here it is still available which I think is a better choice, although it can be confusing.\nI'm 90% sure it contains the same sequences as before - what should explain public LB variations is that\n- Public LB only contains the scores of the 2 (or 1) models selected as final submission\n- Randomness, sequence order, and other not-obvious random things such as hardware variation\n- The MSA directory was updated, and maybe some other files as well (?)\n\nIf there is an issue with the Kaggle rerun (which is not impossible), then let's hope it does not effect private LB standings which are what matters.\n\nNote that there are new answers from the hosts in [this thread](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/686938). \n\n## Performance drops\n\nIn a lot of competitions, you will notice performance shifts between public and private ([example](https://www.kaggle.com/competitions/birdclef-2024), [example](https://www.kaggle.com/competitions/rsna-breast-cancer-detection)). This happens a lot because:\n- Private data can be harder, whether it is intended or not\n- Public kernels are often overfitted to public data\n\nAs a competitor, it is one of your goals to ensure robustness against public / private variations. Note that hosts are encouraged to not share information about the private data by the Kaggle team, hence the oftentimes observed lack of transparency. It is better not to share than to share information that will ruin the competition setup.\n\nThe way I see it as a veteran Kaggler is that it's my goal to gather all the information I can to solve the private LB puzzle.\nSometimes the puzzle is trivial and private LB performance is stable. Sometimes the hosts do something unpredictable. Most of the time it's something in-between, and the solution can be figured using a mix of EDA, public score analysis, and carefully reading all information available.\n\nNot figuring out this puzzle can be frustrating - and I can ensure you that noone has a 100% success rate at that.\n\n## Solving this year's puzzle\n\nFrom the competition page :\n\n> Now, you’ll face even more complex targets, including ones with no structural templates [...]\n\nThis is a big hint towards what your model should be robust against. You should build a solution that's good when there are no strong templates. Which means template-based approaches will very likely not be enough for the task. \n\nHowever, template based approaches worked well on the public dataset. In fact, templates found by top public kernels were so strong that plugging [RNAPro](https://www.kaggle.com/code/jaejohn/rnapro-inference-with-tbm) on top of them did not bring any improvement. This is an indicator that public LB sequences might not meet the \"no strong template\" criterion, or that public kernels performance is masking improvements. And both are hints that a shake-up is likely to happen when newer sequences are collected.\n\nThis is further amplified by the fact that public kernels are crowded with TBM solutions tweaked on the 3rd public LB digit, and that the public test set is quite small. A solution ranked 100th actually performs very similarly to a solution ranked 1000th. This is especially true for the public LB, but holds for the private LB since solutions adapted from public kernels will perform similarly. From then it is a coin flip that decides whether you get a silver medal or end up 500th.\n\nNote that many teams did a brilliant job at addressing this challenge. For instance 6 of the top 8 teams stayed in the gold zone, and last year's 1st and 2nd place also grabed a gold medal this year. \n\n## TLDR\n\n- There was a rerun, only scores of submissions that were selected are relevant.\n- Public LB contains the same sequences as before, but only shows scores of your 2 selected submissions.\n- Private LB contains newer sequences, on which widely used public template-based approaches turned out to be weaker. I do not think this is intended, but there were hints this was going to happen (the general expected difficulty was stated in the competition description page, and public notebook were overfitted). It is one of your (toughest !) challenges as a competitor to develop solutions robust against public/private shift.",
    "3434646": "@theoviel Thanks for the post. \n\nAnd everyone else: thanks for flagging these complexities. \n\nThis Stanford RNA 3D Folding Part 2 competition was meant to be quite similar to Part 1, and our checks before release did not reveal these complexities. \n\nWe competition hosts and the Kaggle developers and competition hosts were caught off guard, and we could have done better on responding earlier that we'd need time and your help to track down.\n\nFor a couple of Theo's points above, the Kaggle devs have agreed to let me describe some additional results that we hope can illustrate what happened.\n\n**Public LB variations**\n\nIt was a surprise to many participants that their Public LB rankings shifted. \n\nThe rerun of notebooks on those Public LB targets was unavoidable with current Kaggle infrastructure, but we didn't expect such a disruption in ranking. \n\nHere's why – the scores of notebooks actually changed very little: \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5659684%2F269d9dbbbbe339d2a85c972ed7455fef%2Fkaggle_part2_pub_origpub_scatterplot.png?generation=1775187253196596&alt=media)\n\nThis plot contains all the original and final Public LB scores for the two selected notebooks for each team. The reproducibility seemed very good, e.g., the mean absolute deviation between scores in the two runs was low (~0.007).  The Pearson correlation was high (r=0.997). This was similar or actually slightly better than the reproducibility in notebook scores that we saw in Part 1....\n\nHowever, there was a crowding of the scores, boxed in red, whose magnitude we didn't notice due to the overlap of points. As @theoviel and others have pointed out, random fluctuations produced small shifts in the score that then led to extremely large changes in the *rankings* within this crowd. \n\nThe hosts should have realized this ahead of time. But at least I (now host of 4 Kaggle competitions) had not seen such level of crowding before, even in Part 1 of this competition. \n\nWe are still very interested in how some teams were able to avoid this shakeup. The solution writeups so far have been illuminating -- some top teams set up their own internal validation sets and explored strategies distinct from the most widely shared public notebooks. We look forward to more writeups, more discussion, and more learning. \n\nAnd if there are future RNA competitions, we'll need to find a way to prevent reruns of the Public LB.\n\n**Performance drops**\n\nAs mentioned elsewhere, it was not known to us ahead of time whether the final private leaderboard would be easier or harder than the Public leaderboard. In the end, it turned out the `best_template_oracle` score was lower in Private than Public.  \n\nBut an important thing to point out is that many notebooks got *higher* scores on the Private leaderboard than Public:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5659684%2F091499c8797533e19b7fbcc7d352944c%2Fkaggle_part2_score_pub_priv_scatterplot.png?generation=1775188453951595&alt=media) \n\nIn terms of raw scores, the Pearson correlation between Private and Public leaderboard scores was actually pretty good (r=0.955). However, we overlooked the same crowding phenomenon (red box) and did not anticipate the subsequent shakeup of rankings. \n\nIt will be interesting to again hear from the teams that managed to retain good scores in both Public and Private leaderboards. Also if Kagglers are familiar with other competitions that had this same crowding, please do post.\n\n**Placeholder Private leaderboard**\n\nAlthough this was not mentioned in @theoviel's post, others, including @hongan @sheriytm in this thread, have noted that it was very confusing to have released the scores for an internal 'placeholder' Private LB test set that was updated at the end with a larger test set.  \n\nThe same thing actually happened in Part 1, but the placeholder Private LB set happened to give lower scores than the final Private LB, so not too many people noted the change. If there are future RNA competitions, hosts and devs now know that we should communicate better about any placeholder Private LB sets -- or avoid them altogether.\n\nDespite all of this, we hope everyone has learned a lot from their experience and from each other in this competition. The scientific impact of this competition's methods for RNA structure prediction in medicine and biology are likely to be quite high. \n\nAnd while we thought this was going to be a straightforward sequel to Part 1 of the Stanford RNA Folding Challenge, it turned out to be more complex -- it makes us hosts appreciate even more the ingenuity of the teams who 'solved the puzzle'.\n",
    "3434119": "The confusion comes from that \n\n- public LB was cleared and when it was then revealed, the scores are not same (host explains that because the order of the sequences has changed, and that is why the public LB score has changed).\n- there are two private scores. I understand that there was an update on private test set. But still usually in rerun you only see the final private score, not two different scores with huge differences.\n\nIt is not common to see these practices in competitions. So naturally questions were asked! But now I think it is explained by the hosts in a more transparent way that at least cleared all my doubt. I just hope they can explain what happened in the first place.",
    "3434358": "@theoviel, do you really think by them ignoring answering our genuine be it uncomfortable questions, they are doing us as a community and Kaggle's reputation a favor?"
  }
}