{
  "id": 679781,
  "title": "What Metric Should I Use for Cross-Validation?",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/679781",
  "author_name": "Ananya Dwivedi",
  "post_date": "2026-03-03T18:26:47.332000",
  "votes": 0,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi everyone, I’m new to structural prediction competitions and a bit confused about validation.</p>\n<p>Since the official metric is alignment-based (TM-score via USalign), should I use the exact same metric for cross-validation? Making the full TM-score validation pipeline (as mentioned in one of the discussions) seems quite hard.</p>\n<p>I also saw a combination of multiple metrics used in a previous competition notebook, which made me even more confused about how to properly design validation.</p>\n<p>Also, how are you structuring folds to avoid leakage (by target, sequence similarity, or time)?</p>\n<p>Would really appreciate some beginner-friendly guidance 🙏</p>",
  "messages": [
    {
      "id": 3417807,
      "postDate": "2026-03-06T09:11:55.170Z",
      "content": "<p>Hi,</p>\n<p>if you check my notebook, I uploaded the metric.py file as a dataset:\n<a href=\"https://www.kaggle.com/code/andreshzapke/stanford-rna-folding-2-template-based-approach/edit\" target=\"_blank\">https://www.kaggle.com/code/andreshzapke/stanford-rna-folding-2-template-based-approach/edit</a></p>\n<p>then you can simply:</p>\n<p>`sys.path.append(\"/kaggle/input/datasets/andreshzapke/metric-score\")</p>\n<p>from metric import score`</p>\n<p>and then use score() as explained here:\n<a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/667106\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/667106</a></p>\n<p>As to cross-validation, I am unsure what you mean by it, unless you are training a ML model, I would always use the same metric for both training and validating.</p>\n<p>In any case the metric.py file is quite complex, but I could learn a lot just from reading this (without ML) approach:</p>\n<p><a href=\"https://www.kaggle.com/code/nihilisticneuralnet/stanford-rna-folding-2-template-based-approach\" target=\"_blank\">https://www.kaggle.com/code/nihilisticneuralnet/stanford-rna-folding-2-template-based-approach</a></p>",
      "rawMarkdown": "Hi,\n\nif you check my notebook, I uploaded the metric.py file as a dataset:\nhttps://www.kaggle.com/code/andreshzapke/stanford-rna-folding-2-template-based-approach/edit\n\nthen you can simply:\n\n`sys.path.append(\"/kaggle/input/datasets/andreshzapke/metric-score\")\n\nfrom metric import score`\n\nand then use score() as explained here:\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/667106\n\nAs to cross-validation, I am unsure what you mean by it, unless you are training a ML model, I would always use the same metric for both training and validating.\n\nIn any case the metric.py file is quite complex, but I could learn a lot just from reading this (without ML) approach:\n\nhttps://www.kaggle.com/code/nihilisticneuralnet/stanford-rna-folding-2-template-based-approach",
      "votes": 1
    },
    {
      "id": 3416786,
      "postDate": "2026-03-03T18:26:47.333Z",
      "content": "<p>Hi everyone, I’m new to structural prediction competitions and a bit confused about validation.</p>\n<p>Since the official metric is alignment-based (TM-score via USalign), should I use the exact same metric for cross-validation? Making the full TM-score validation pipeline (as mentioned in one of the discussions) seems quite hard.</p>\n<p>I also saw a combination of multiple metrics used in a previous competition notebook, which made me even more confused about how to properly design validation.</p>\n<p>Also, how are you structuring folds to avoid leakage (by target, sequence similarity, or time)?</p>\n<p>Would really appreciate some beginner-friendly guidance 🙏</p>",
      "rawMarkdown": "Hi everyone, I’m new to structural prediction competitions and a bit confused about validation.\n\nSince the official metric is alignment-based (TM-score via USalign), should I use the exact same metric for cross-validation? Making the full TM-score validation pipeline (as mentioned in one of the discussions) seems quite hard.\n\nI also saw a combination of multiple metrics used in a previous competition notebook, which made me even more confused about how to properly design validation.\n\nAlso, how are you structuring folds to avoid leakage (by target, sequence similarity, or time)?\n\nWould really appreciate some beginner-friendly guidance 🙏"
    }
  ],
  "comments": [
    {
      "id": 3417807,
      "author_name": "Andres H. Zapke",
      "author_url": "",
      "post_date": "2026-03-06T09:11:55.170000",
      "content": "<p>Hi,</p>\n<p>if you check my notebook, I uploaded the metric.py file as a dataset:\n<a href=\"https://www.kaggle.com/code/andreshzapke/stanford-rna-folding-2-template-based-approach/edit\" target=\"_blank\">https://www.kaggle.com/code/andreshzapke/stanford-rna-folding-2-template-based-approach/edit</a></p>\n<p>then you can simply:</p>\n<p>`sys.path.append(\"/kaggle/input/datasets/andreshzapke/metric-score\")</p>\n<p>from metric import score`</p>\n<p>and then use score() as explained here:\n<a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/667106\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/667106</a></p>\n<p>As to cross-validation, I am unsure what you mean by it, unless you are training a ML model, I would always use the same metric for both training and validating.</p>\n<p>In any case the metric.py file is quite complex, but I could learn a lot just from reading this (without ML) approach:</p>\n<p><a href=\"https://www.kaggle.com/code/nihilisticneuralnet/stanford-rna-folding-2-template-based-approach\" target=\"_blank\">https://www.kaggle.com/code/nihilisticneuralnet/stanford-rna-folding-2-template-based-approach</a></p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3417807": "Hi,\n\nif you check my notebook, I uploaded the metric.py file as a dataset:\nhttps://www.kaggle.com/code/andreshzapke/stanford-rna-folding-2-template-based-approach/edit\n\nthen you can simply:\n\n`sys.path.append(\"/kaggle/input/datasets/andreshzapke/metric-score\")\n\nfrom metric import score`\n\nand then use score() as explained here:\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding-2/discussion/667106\n\nAs to cross-validation, I am unsure what you mean by it, unless you are training a ML model, I would always use the same metric for both training and validating.\n\nIn any case the metric.py file is quite complex, but I could learn a lot just from reading this (without ML) approach:\n\nhttps://www.kaggle.com/code/nihilisticneuralnet/stanford-rna-folding-2-template-based-approach",
    "3416786": "Hi everyone, I’m new to structural prediction competitions and a bit confused about validation.\n\nSince the official metric is alignment-based (TM-score via USalign), should I use the exact same metric for cross-validation? Making the full TM-score validation pipeline (as mentioned in one of the discussions) seems quite hard.\n\nI also saw a combination of multiple metrics used in a previous competition notebook, which made me even more confused about how to properly design validation.\n\nAlso, how are you structuring folds to avoid leakage (by target, sequence similarity, or time)?\n\nWould really appreciate some beginner-friendly guidance 🙏"
  }
}