{
  "id": 669710,
  "title": "What is exactly temporal cutoff?",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/669710",
  "author_name": "Bridgeport",
  "post_date": "2026-01-23T21:34:30.283000",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Here are two identical sequences with different target_id and temporal cutoff: \n'5EPI' and '8RN1'. Can somebody clarify meaning of temporal cutoff using these two target_ids?</p>\n<p>I would like to compute TM-score for them (since they are perfectly aligned), but cannot understand over what we need to take maximum. </p>\n<p>In addition, is there python utility to compute TM-score?</p>",
  "messages": [
    {
      "id": 3397193,
      "postDate": "2026-01-26T16:58:01.043Z",
      "content": "<p>Hello, </p>\n<p>The main purpose of the <code>temporal_cutoff</code> is to provide an easy way of filtering the RNA that may have been previously used to train a model, to limit the data leakage. This is essentially a date when a given structure was released publicly. Entries such as <code>5EPI</code> and <code>8RN1</code> represent the same RNA but with structures determined at different time, that may differ in method, experimental conditions, additional molecules etc. </p>\n<p>There is slight difference how we assign temporal cutoff for the structures in the train and validation data.\nThe <code>train_sequences.csv</code> are not aggregated in any way. We leave to competitors discretion how to select appropriate entries. In this dataset you will see redundant entries such as <code>5EPI</code> and <code>8RN1</code>, that were released at different times (hence different <code>temporal_cutoff</code>). Depending on the application you may want to select one or the other. If it is desired to remove redundancy, the <code>extra/rna_metadata.csv</code> contains additional information about the experiment (such as resolution) and sequence based clusters (please take a look <code>extra/README.md</code> for description of additional columns).</p>\n<p>For the  <code>{validation,test}_sequences.csv</code>  we removed redundancy based on MMseqs2 clustering (mmseqs_0.900 column in <code>extra/rna_metadata.csv</code>. In this case the representative entry gets assigned the earliest <code>temporal_cutoff</code> from a set of sequences having the same ID in a clustering column.</p>\n<p>Because the <code>train_sequences.csv</code> are not aggregated, <code>train_labels.csv</code> have only one set of coordinates (<code>x_1, y_1, z_1</code>) for each target.  <code>validation_labels.csv</code> can have multiple coordinates, each set reflecting different PDB entries, with the same sequence and RNA composition. </p>\n<p>If you would like to calculate the TM-score between different entries in <code>train_sequences.csv</code> you will have to take coordinates for each entry and compare them to each other. Depending on number of entries with the same sequence (or within the same cluster) and your use case, you may need to do all-vs-all comparison and for example calculate the maximum TM-score  to represent maximum structural similarity within that cluster. For two entries you will have just one score comparing them directly.</p>\n<p>For the purpose of scoring submissions we calculate all-vs-all TM-scores between 5 predicted conformations and ground truth conformations, and then take the maximum as the final TM-score for that entry. You can replicate that when scoring you predictions against <code>validation_labels.csv</code>.</p>\n<p>To do that in pure python you can look at <a href=\"https://www.biotite-python.org/\" target=\"_blank\">Biotite</a> that has modules to <a href=\"https://www.biotite-python.org/latest/apidoc/biotite.structure.superimpose.html\" target=\"_blank\">superimpose structures</a> and calculate <a href=\"https://www.biotite-python.org/latest/apidoc/biotite.structure.tm_score.html\" target=\"_blank\">TM-score</a>. These results may be different though from the official scoring, therefore I recommend using the official scoring metric that <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> posted here: <a href=\"https://www.kaggle.com/code/rhijudas/tm-score-permutechains\" target=\"_blank\">https://www.kaggle.com/code/rhijudas/tm-score-permutechains</a> and described how to add to your notebooks here:  <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/667106\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/667106</a> </p>\n<p>Hope that helps!</p>",
      "rawMarkdown": "Hello, \n\nThe main purpose of the `temporal_cutoff` is to provide an easy way of filtering the RNA that may have been previously used to train a model, to limit the data leakage. This is essentially a date when a given structure was released publicly. Entries such as `5EPI` and `8RN1` represent the same RNA but with structures determined at different time, that may differ in method, experimental conditions, additional molecules etc. \n\nThere is slight difference how we assign temporal cutoff for the structures in the train and validation data.\nThe `train_sequences.csv` are not aggregated in any way. We leave to competitors discretion how to select appropriate entries. In this dataset you will see redundant entries such as `5EPI` and `8RN1`, that were released at different times (hence different `temporal_cutoff`). Depending on the application you may want to select one or the other. If it is desired to remove redundancy, the `extra/rna_metadata.csv` contains additional information about the experiment (such as resolution) and sequence based clusters (please take a look `extra/README.md` for description of additional columns).\n\nFor the  `{validation,test}_sequences.csv`  we removed redundancy based on MMseqs2 clustering (mmseqs_0.900 column in `extra/rna_metadata.csv`. In this case the representative entry gets assigned the earliest `temporal_cutoff` from a set of sequences having the same ID in a clustering column.\n\nBecause the `train_sequences.csv` are not aggregated, `train_labels.csv` have only one set of coordinates (`x_1, y_1, z_1`) for each target.  `validation_labels.csv` can have multiple coordinates, each set reflecting different PDB entries, with the same sequence and RNA composition. \n\nIf you would like to calculate the TM-score between different entries in `train_sequences.csv` you will have to take coordinates for each entry and compare them to each other. Depending on number of entries with the same sequence (or within the same cluster) and your use case, you may need to do all-vs-all comparison and for example calculate the maximum TM-score  to represent maximum structural similarity within that cluster. For two entries you will have just one score comparing them directly.\n\nFor the purpose of scoring submissions we calculate all-vs-all TM-scores between 5 predicted conformations and ground truth conformations, and then take the maximum as the final TM-score for that entry. You can replicate that when scoring you predictions against `validation_labels.csv`.\n\nTo do that in pure python you can look at [Biotite](https://www.biotite-python.org/) that has modules to [superimpose structures](https://www.biotite-python.org/latest/apidoc/biotite.structure.superimpose.html) and calculate [TM-score](https://www.biotite-python.org/latest/apidoc/biotite.structure.tm_score.html). These results may be different though from the official scoring, therefore I recommend using the official scoring metric that @rhijudas posted here: https://www.kaggle.com/code/rhijudas/tm-score-permutechains and described how to add to your notebooks here:  https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/667106 \n\nHope that helps!",
      "votes": 3,
      "replies": [
        {
          "id": 3400907,
          "postDate": "2026-02-02T15:29:44.620Z",
          "content": "<p>Thank you so much for the clarifications! For the same sequence with IDs <code>5EPI</code> and <code>8RN1</code> I get score of 0.54. Would you say this is normal score, or is it too low or too high?</p>\n<p>Here is my notebook with score computations.\n<a href=\"https://www.kaggle.com/code/bridgeport/sam-sequence-different-ids\" target=\"_blank\">https://www.kaggle.com/code/bridgeport/sam-sequence-different-ids</a></p>",
          "rawMarkdown": "Thank you so much for the clarifications! For the same sequence with IDs `5EPI` and `8RN1` I get score of 0.54. Would you say this is normal score, or is it too low or too high?\n\nHere is my notebook with score computations.\nhttps://www.kaggle.com/code/bridgeport/sam-sequence-different-ids",
          "votes": 1,
          "replies": [
            {
              "id": 3401576,
              "postDate": "2026-02-04T03:49:36.883Z",
              "content": "<p>This is the correct score for these structures. </p>",
              "rawMarkdown": "This is the correct score for these structures. ",
              "votes": 2
            },
            {
              "id": 3401915,
              "postDate": "2026-02-04T20:27:00.517Z",
              "content": "<p>Thank you confirmation. My question was more about parameters of distribution of scores for the case of identical sequences and different ids. Something like: for any sequence with two different IDs we can expect on average score of 0.6 with standard deviation of 0.2. I am trying to compute it right now myself, but it is taking longer than I expected.</p>",
              "rawMarkdown": "Thank you confirmation. My question was more about parameters of distribution of scores for the case of identical sequences and different ids. Something like: for any sequence with two different IDs we can expect on average score of 0.6 with standard deviation of 0.2. I am trying to compute it right now myself, but it is taking longer than I expected."
            },
            {
              "id": 3404043,
              "postDate": "2026-02-09T19:06:30.677Z",
              "content": "<p>We do not have such statistics. The results will vary significantly depending on experimental methods, resolution and conditions. For better context, you can refer to <a href=\"https://academic.oup.com/bioinformatics/article/35/21/4459/5480133\" target=\"_blank\">RNAalign publication</a> and <a href=\"https://onlinelibrary.wiley.com/doi/10.1002/prot.70072\" target=\"_blank\">CASP16 RNA findings</a> (esp. <a href=\"https://onlinelibrary.wiley.com/doi/10.1002/prot.70072#prot70072-fig-0007\" target=\"_blank\">Figure 7</a>)</p>",
              "rawMarkdown": "We do not have such statistics. The results will vary significantly depending on experimental methods, resolution and conditions. For better context, you can refer to [RNAalign publication](https://academic.oup.com/bioinformatics/article/35/21/4459/5480133) and [CASP16 RNA findings] (https://onlinelibrary.wiley.com/doi/10.1002/prot.70072) (esp. [Figure 7](https://onlinelibrary.wiley.com/doi/10.1002/prot.70072#prot70072-fig-0007))",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3395890,
      "postDate": "2026-01-23T21:34:30.283Z",
      "content": "<p>Here are two identical sequences with different target_id and temporal cutoff: \n'5EPI' and '8RN1'. Can somebody clarify meaning of temporal cutoff using these two target_ids?</p>\n<p>I would like to compute TM-score for them (since they are perfectly aligned), but cannot understand over what we need to take maximum. </p>\n<p>In addition, is there python utility to compute TM-score?</p>",
      "rawMarkdown": "Here are two identical sequences with different target_id and temporal cutoff: \n'5EPI' and '8RN1'. Can somebody clarify meaning of temporal cutoff using these two target_ids?\n\nI would like to compute TM-score for them (since they are perfectly aligned), but cannot understand over what we need to take maximum. \n\nIn addition, is there python utility to compute TM-score?\n",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 3397193,
      "author_name": "Przemek Porebski",
      "author_url": "",
      "post_date": "2026-01-26T16:58:01.043000",
      "content": "<p>Hello, </p>\n<p>The main purpose of the <code>temporal_cutoff</code> is to provide an easy way of filtering the RNA that may have been previously used to train a model, to limit the data leakage. This is essentially a date when a given structure was released publicly. Entries such as <code>5EPI</code> and <code>8RN1</code> represent the same RNA but with structures determined at different time, that may differ in method, experimental conditions, additional molecules etc. </p>\n<p>There is slight difference how we assign temporal cutoff for the structures in the train and validation data.\nThe <code>train_sequences.csv</code> are not aggregated in any way. We leave to competitors discretion how to select appropriate entries. In this dataset you will see redundant entries such as <code>5EPI</code> and <code>8RN1</code>, that were released at different times (hence different <code>temporal_cutoff</code>). Depending on the application you may want to select one or the other. If it is desired to remove redundancy, the <code>extra/rna_metadata.csv</code> contains additional information about the experiment (such as resolution) and sequence based clusters (please take a look <code>extra/README.md</code> for description of additional columns).</p>\n<p>For the  <code>{validation,test}_sequences.csv</code>  we removed redundancy based on MMseqs2 clustering (mmseqs_0.900 column in <code>extra/rna_metadata.csv</code>. In this case the representative entry gets assigned the earliest <code>temporal_cutoff</code> from a set of sequences having the same ID in a clustering column.</p>\n<p>Because the <code>train_sequences.csv</code> are not aggregated, <code>train_labels.csv</code> have only one set of coordinates (<code>x_1, y_1, z_1</code>) for each target.  <code>validation_labels.csv</code> can have multiple coordinates, each set reflecting different PDB entries, with the same sequence and RNA composition. </p>\n<p>If you would like to calculate the TM-score between different entries in <code>train_sequences.csv</code> you will have to take coordinates for each entry and compare them to each other. Depending on number of entries with the same sequence (or within the same cluster) and your use case, you may need to do all-vs-all comparison and for example calculate the maximum TM-score  to represent maximum structural similarity within that cluster. For two entries you will have just one score comparing them directly.</p>\n<p>For the purpose of scoring submissions we calculate all-vs-all TM-scores between 5 predicted conformations and ground truth conformations, and then take the maximum as the final TM-score for that entry. You can replicate that when scoring you predictions against <code>validation_labels.csv</code>.</p>\n<p>To do that in pure python you can look at <a href=\"https://www.biotite-python.org/\" target=\"_blank\">Biotite</a> that has modules to <a href=\"https://www.biotite-python.org/latest/apidoc/biotite.structure.superimpose.html\" target=\"_blank\">superimpose structures</a> and calculate <a href=\"https://www.biotite-python.org/latest/apidoc/biotite.structure.tm_score.html\" target=\"_blank\">TM-score</a>. These results may be different though from the official scoring, therefore I recommend using the official scoring metric that <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> posted here: <a href=\"https://www.kaggle.com/code/rhijudas/tm-score-permutechains\" target=\"_blank\">https://www.kaggle.com/code/rhijudas/tm-score-permutechains</a> and described how to add to your notebooks here:  <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/667106\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/667106</a> </p>\n<p>Hope that helps!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3400907,
          "author_name": "Bridgeport",
          "author_url": "",
          "post_date": "2026-02-02T15:29:44.620000",
          "content": "<p>Thank you so much for the clarifications! For the same sequence with IDs <code>5EPI</code> and <code>8RN1</code> I get score of 0.54. Would you say this is normal score, or is it too low or too high?</p>\n<p>Here is my notebook with score computations.\n<a href=\"https://www.kaggle.com/code/bridgeport/sam-sequence-different-ids\" target=\"_blank\">https://www.kaggle.com/code/bridgeport/sam-sequence-different-ids</a></p>",
          "votes": 1,
          "replies": [
            {
              "id": 3401576,
              "author_name": "Przemek Porebski",
              "author_url": "",
              "post_date": "2026-02-04T03:49:36.883000",
              "content": "<p>This is the correct score for these structures. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3401915,
              "author_name": "Bridgeport",
              "author_url": "",
              "post_date": "2026-02-04T20:27:00.517000",
              "content": "<p>Thank you confirmation. My question was more about parameters of distribution of scores for the case of identical sequences and different ids. Something like: for any sequence with two different IDs we can expect on average score of 0.6 with standard deviation of 0.2. I am trying to compute it right now myself, but it is taking longer than I expected.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3404043,
              "author_name": "Przemek Porebski",
              "author_url": "",
              "post_date": "2026-02-09T19:06:30.677000",
              "content": "<p>We do not have such statistics. The results will vary significantly depending on experimental methods, resolution and conditions. For better context, you can refer to <a href=\"https://academic.oup.com/bioinformatics/article/35/21/4459/5480133\" target=\"_blank\">RNAalign publication</a> and <a href=\"https://onlinelibrary.wiley.com/doi/10.1002/prot.70072\" target=\"_blank\">CASP16 RNA findings</a> (esp. <a href=\"https://onlinelibrary.wiley.com/doi/10.1002/prot.70072#prot70072-fig-0007\" target=\"_blank\">Figure 7</a>)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3397193": "Hello, \n\nThe main purpose of the `temporal_cutoff` is to provide an easy way of filtering the RNA that may have been previously used to train a model, to limit the data leakage. This is essentially a date when a given structure was released publicly. Entries such as `5EPI` and `8RN1` represent the same RNA but with structures determined at different time, that may differ in method, experimental conditions, additional molecules etc. \n\nThere is slight difference how we assign temporal cutoff for the structures in the train and validation data.\nThe `train_sequences.csv` are not aggregated in any way. We leave to competitors discretion how to select appropriate entries. In this dataset you will see redundant entries such as `5EPI` and `8RN1`, that were released at different times (hence different `temporal_cutoff`). Depending on the application you may want to select one or the other. If it is desired to remove redundancy, the `extra/rna_metadata.csv` contains additional information about the experiment (such as resolution) and sequence based clusters (please take a look `extra/README.md` for description of additional columns).\n\nFor the  `{validation,test}_sequences.csv`  we removed redundancy based on MMseqs2 clustering (mmseqs_0.900 column in `extra/rna_metadata.csv`. In this case the representative entry gets assigned the earliest `temporal_cutoff` from a set of sequences having the same ID in a clustering column.\n\nBecause the `train_sequences.csv` are not aggregated, `train_labels.csv` have only one set of coordinates (`x_1, y_1, z_1`) for each target.  `validation_labels.csv` can have multiple coordinates, each set reflecting different PDB entries, with the same sequence and RNA composition. \n\nIf you would like to calculate the TM-score between different entries in `train_sequences.csv` you will have to take coordinates for each entry and compare them to each other. Depending on number of entries with the same sequence (or within the same cluster) and your use case, you may need to do all-vs-all comparison and for example calculate the maximum TM-score  to represent maximum structural similarity within that cluster. For two entries you will have just one score comparing them directly.\n\nFor the purpose of scoring submissions we calculate all-vs-all TM-scores between 5 predicted conformations and ground truth conformations, and then take the maximum as the final TM-score for that entry. You can replicate that when scoring you predictions against `validation_labels.csv`.\n\nTo do that in pure python you can look at [Biotite](https://www.biotite-python.org/) that has modules to [superimpose structures](https://www.biotite-python.org/latest/apidoc/biotite.structure.superimpose.html) and calculate [TM-score](https://www.biotite-python.org/latest/apidoc/biotite.structure.tm_score.html). These results may be different though from the official scoring, therefore I recommend using the official scoring metric that @rhijudas posted here: https://www.kaggle.com/code/rhijudas/tm-score-permutechains and described how to add to your notebooks here:  https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/667106 \n\nHope that helps!",
    "3395890": "Here are two identical sequences with different target_id and temporal cutoff: \n'5EPI' and '8RN1'. Can somebody clarify meaning of temporal cutoff using these two target_ids?\n\nI would like to compute TM-score for them (since they are perfectly aligned), but cannot understand over what we need to take maximum. \n\nIn addition, is there python utility to compute TM-score?\n"
  }
}