{
  "id": 674343,
  "title": "Question Regarding the Usage Rules for protenix_base_20250630_v1.0.0",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/674343",
  "author_name": "Makotu",
  "post_date": "2026-02-19T23:14:11.117000",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>  ,</p>\n<p>I have noticed that some public notebooks are using models based on protenix_base_20250630_v1.0.0, and I would like to clarify the rules around its usage.</p>\n<p>Specifically, based on the dataset description, the test data appears to include sequences released on or after May 29, 2025. At the same time, protenix_base_20250630_v1.0.0 seems to have been trained on data up to June 30, 2025.</p>\n<p>This raises the concern that the model may have been exposed to some of the test data during training. Therefore, my question is: <strong>would using protenix_base_20250630_v1.0.0 to predict sequences released between May 29 and June 30, 2025 constitute a violation of the competition rules?</strong></p>\n<p>To mitigate this risk, I have currently implemented a workaround that uses a different model for sequences released within that date range. I would like to ask the hosts whether such a workaround is actually necessary.</p>\n<p>Thank you very much for your time and assistance.</p>",
  "messages": [
    {
      "id": 3408163,
      "postDate": "2026-02-19T23:34:07.743Z",
      "content": "<p>Good question! </p>\n<p>There's no risk that using Protenix – or even a Protenix version retrained with all publicly available sequences and structures today -- will leak into the actual hidden test set that is used for the competition Public or Private leaderboards. </p>\n<p>That's because all of the test RNA molecules involve experimentally solved 3D structures that have not been made publicly available!  </p>\n<p>However, for teams that are hill-climbing with the validation set (which is the same as the test_sequences.csv provided publicly <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/data?select=test_sequences.csv\" target=\"_blank\">here</a>), there is a risk of leakage that will affect decisions during algorithm development.  </p>\n<p>So it's a great idea to have the workaround in place, but it's not necessary to satisfy the competition rules.</p>",
      "rawMarkdown": "Good question! \n\nThere's no risk that using Protenix – or even a Protenix version retrained with all publicly available sequences and structures today -- will leak into the actual hidden test set that is used for the competition Public or Private leaderboards. \n\nThat's because all of the test RNA molecules involve experimentally solved 3D structures that have not been made publicly available!  \n\nHowever, for teams that are hill-climbing with the validation set (which is the same as the test_sequences.csv provided publicly [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/data?select=test_sequences.csv)), there is a risk of leakage that will affect decisions during algorithm development.  \n\nSo it's a great idea to have the workaround in place, but it's not necessary to satisfy the competition rules.",
      "votes": 5,
      "replies": [
        {
          "id": 3408168,
          "postDate": "2026-02-19T23:45:43.073Z",
          "content": "<p>Thank you for the quick response! \nThe rule is now clear to me. I can now use it with confidence!</p>",
          "rawMarkdown": "Thank you for the quick response! \nThe rule is now clear to me. I can now use it with confidence!"
        }
      ]
    },
    {
      "id": 3408159,
      "postDate": "2026-02-19T23:14:11.117Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>  ,</p>\n<p>I have noticed that some public notebooks are using models based on protenix_base_20250630_v1.0.0, and I would like to clarify the rules around its usage.</p>\n<p>Specifically, based on the dataset description, the test data appears to include sequences released on or after May 29, 2025. At the same time, protenix_base_20250630_v1.0.0 seems to have been trained on data up to June 30, 2025.</p>\n<p>This raises the concern that the model may have been exposed to some of the test data during training. Therefore, my question is: <strong>would using protenix_base_20250630_v1.0.0 to predict sequences released between May 29 and June 30, 2025 constitute a violation of the competition rules?</strong></p>\n<p>To mitigate this risk, I have currently implemented a workaround that uses a different model for sequences released within that date range. I would like to ask the hosts whether such a workaround is actually necessary.</p>\n<p>Thank you very much for your time and assistance.</p>",
      "rawMarkdown": "Hi @rhijudas  ,\n\nI have noticed that some public notebooks are using models based on protenix_base_20250630_v1.0.0, and I would like to clarify the rules around its usage.\n\nSpecifically, based on the dataset description, the test data appears to include sequences released on or after May 29, 2025. At the same time, protenix_base_20250630_v1.0.0 seems to have been trained on data up to June 30, 2025.\n\nThis raises the concern that the model may have been exposed to some of the test data during training. Therefore, my question is: **would using protenix_base_20250630_v1.0.0 to predict sequences released between May 29 and June 30, 2025 constitute a violation of the competition rules?**\n\nTo mitigate this risk, I have currently implemented a workaround that uses a different model for sequences released within that date range. I would like to ask the hosts whether such a workaround is actually necessary.\n\nThank you very much for your time and assistance.\n\n",
      "votes": 4
    },
    {
      "id": 3410575,
      "postDate": "2026-02-23T09:54:39.487Z",
      "content": "<p>Does the provided training data include all the publicly available data? <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> s </p>",
      "rawMarkdown": "Does the provided training data include all the publicly available data? @rhijudas s ",
      "replies": [
        {
          "id": 3410840,
          "postDate": "2026-02-23T18:40:06.677Z",
          "content": "<p>No, we do not provide all publicly available data. We used <strong>Dec. 17th 2025</strong> cutoff for all entries and we do additional filtering and remove some entries from the train and validation datasets.</p>\n<p>Here is a table that summarizes different datasets that are part of this competition:</p>\n<table>\n<thead>\n<tr>\n<th>Dataset / file</th>\n<th>Number of entries</th>\n<th>Selection criteria</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>PDB_RNA</code></td>\n<td>9564</td>\n<td>All PDB entries released before <strong>Dec. 17th 2025</strong> that contain RNA or RNA-DNA hybrid</td>\n</tr>\n<tr>\n<td><code>train_sequences.csv</code></td>\n<td>5716</td>\n<td>Entries that fall into 30% identity clusters for which any entry was released before <strong>May 29th, 2025</strong>. Filtered according to quality criteria. See additional notes in updated <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/data\" target=\"_blank\">data description</a> and <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/674359#3408505\" target=\"_blank\">this post</a></td>\n</tr>\n<tr>\n<td><code>{validation,test}_sequences.csv</code></td>\n<td>28</td>\n<td>Entries that do not have &gt;30% identity homologs released before <strong>May 29th, 2025</strong>, filtered according to quality criteria, with redundancy removed at 90% identity</td>\n</tr>\n</tbody>\n</table>\n<p>For comparison PDB repository as of Feb. 23rd contains 9803 entries. </p>",
          "rawMarkdown": "No, we do not provide all publicly available data. We used **Dec. 17th 2025** cutoff for all entries and we do additional filtering and remove some entries from the train and validation datasets.\n\nHere is a table that summarizes different datasets that are part of this competition:\n| Dataset / file | Number of entries | Selection criteria\n| --- | --- |\n| `PDB_RNA` | 9564 | All PDB entries released before **Dec. 17th 2025** that contain RNA or RNA-DNA hybrid | \n| `train_sequences.csv` | 5716 | Entries that fall into 30% identity clusters for which any entry was released before **May 29th, 2025**. Filtered according to quality criteria. See additional notes in updated [data description](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/data) and [this post](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/674359#3408505) |\n| `{validation,test}_sequences.csv` | 28 | Entries that do not have >30% identity homologs released before **May 29th, 2025**, filtered according to quality criteria, with redundancy removed at 90% identity\n\n\n\nFor comparison PDB repository as of Feb. 23rd contains 9803 entries. \n",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3408163,
      "author_name": "Rhiju Das",
      "author_url": "",
      "post_date": "2026-02-19T23:34:07.743000",
      "content": "<p>Good question! </p>\n<p>There's no risk that using Protenix – or even a Protenix version retrained with all publicly available sequences and structures today -- will leak into the actual hidden test set that is used for the competition Public or Private leaderboards. </p>\n<p>That's because all of the test RNA molecules involve experimentally solved 3D structures that have not been made publicly available!  </p>\n<p>However, for teams that are hill-climbing with the validation set (which is the same as the test_sequences.csv provided publicly <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/data?select=test_sequences.csv\" target=\"_blank\">here</a>), there is a risk of leakage that will affect decisions during algorithm development.  </p>\n<p>So it's a great idea to have the workaround in place, but it's not necessary to satisfy the competition rules.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 3408168,
          "author_name": "Makotu",
          "author_url": "",
          "post_date": "2026-02-19T23:45:43.073000",
          "content": "<p>Thank you for the quick response! \nThe rule is now clear to me. I can now use it with confidence!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3410575,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2026-02-23T09:54:39.487000",
      "content": "<p>Does the provided training data include all the publicly available data? <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> s </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3410840,
          "author_name": "Przemek Porebski",
          "author_url": "",
          "post_date": "2026-02-23T18:40:06.677000",
          "content": "<p>No, we do not provide all publicly available data. We used <strong>Dec. 17th 2025</strong> cutoff for all entries and we do additional filtering and remove some entries from the train and validation datasets.</p>\n<p>Here is a table that summarizes different datasets that are part of this competition:</p>\n<table>\n<thead>\n<tr>\n<th>Dataset / file</th>\n<th>Number of entries</th>\n<th>Selection criteria</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>PDB_RNA</code></td>\n<td>9564</td>\n<td>All PDB entries released before <strong>Dec. 17th 2025</strong> that contain RNA or RNA-DNA hybrid</td>\n</tr>\n<tr>\n<td><code>train_sequences.csv</code></td>\n<td>5716</td>\n<td>Entries that fall into 30% identity clusters for which any entry was released before <strong>May 29th, 2025</strong>. Filtered according to quality criteria. See additional notes in updated <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/data\" target=\"_blank\">data description</a> and <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/674359#3408505\" target=\"_blank\">this post</a></td>\n</tr>\n<tr>\n<td><code>{validation,test}_sequences.csv</code></td>\n<td>28</td>\n<td>Entries that do not have &gt;30% identity homologs released before <strong>May 29th, 2025</strong>, filtered according to quality criteria, with redundancy removed at 90% identity</td>\n</tr>\n</tbody>\n</table>\n<p>For comparison PDB repository as of Feb. 23rd contains 9803 entries. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3408163": "Good question! \n\nThere's no risk that using Protenix – or even a Protenix version retrained with all publicly available sequences and structures today -- will leak into the actual hidden test set that is used for the competition Public or Private leaderboards. \n\nThat's because all of the test RNA molecules involve experimentally solved 3D structures that have not been made publicly available!  \n\nHowever, for teams that are hill-climbing with the validation set (which is the same as the test_sequences.csv provided publicly [here](https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/data?select=test_sequences.csv)), there is a risk of leakage that will affect decisions during algorithm development.  \n\nSo it's a great idea to have the workaround in place, but it's not necessary to satisfy the competition rules.",
    "3408159": "Hi @rhijudas  ,\n\nI have noticed that some public notebooks are using models based on protenix_base_20250630_v1.0.0, and I would like to clarify the rules around its usage.\n\nSpecifically, based on the dataset description, the test data appears to include sequences released on or after May 29, 2025. At the same time, protenix_base_20250630_v1.0.0 seems to have been trained on data up to June 30, 2025.\n\nThis raises the concern that the model may have been exposed to some of the test data during training. Therefore, my question is: **would using protenix_base_20250630_v1.0.0 to predict sequences released between May 29 and June 30, 2025 constitute a violation of the competition rules?**\n\nTo mitigate this risk, I have currently implemented a workaround that uses a different model for sequences released within that date range. I would like to ask the hosts whether such a workaround is actually necessary.\n\nThank you very much for your time and assistance.\n\n",
    "3410575": "Does the provided training data include all the publicly available data? @rhijudas s "
  }
}