{
  "id": 670892,
  "title": "Ribonanza QuickStart",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/670892",
  "author_name": "Rhiju Das",
  "post_date": "2026-01-30T04:11:05.795000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Some folks have asked about auxiliary data sources for this challenge's 3D RNA structure prediction problem. </p>\n<p>We wanted to remind everyone that, in addition to this competition's <code>train_labels.csv</code> with 3D coordinates, there are also chemical mapping data that are sensitive to RNA structure and can be acquired at large scale, but remain underexplored for boosting 3D modeling. </p>\n<p>For example, a previous <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding\" target=\"_blank\">Ribonanza Kaggle challenge</a> made publicly available these kinds of experimental data for &gt;2M RNA sequences -- but it's an open question as to whether these data can improve modeling of RNA 3D structure.</p>\n<p>Interested teams might get started with the Ribonanza 'QUICK_START' data set:</p>\n<p><a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data?select=train_data_QUICK_START.csv\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data?select=train_data_QUICK_START.csv</a></p>\n<p>This data set has been curated to hold 2A3 and DMS chemical mapping profiles with reasonable signal to noise on ~170k RNA's with diverse sequences and structural ensembles.</p>\n<p>In addition, we know that many teams are currently incorporating template-based modeling into their methods. </p>\n<p>To help those efforts make contact with the Ribonanza data, <a href=\"https://www.kaggle.com/ckjoshi9\" target=\"_blank\">@ckjoshi9</a> – who is joining the HHMI &amp; Stanford teams co-organizing this competition – has recently made available 3D templates corresponding to these sequences here:</p>\n<p><a href=\"https://www.kaggle.com/datasets/ckjoshi9/ribonanza-quickstart-3d-templates\" target=\"_blank\">https://www.kaggle.com/datasets/ckjoshi9/ribonanza-quickstart-3d-templates</a></p>\n<p>If you find useful signal from those data, note that there's much more publicly available, including the full set of training profiles and test data from the Ribonanza competition (see links in the Supplemental Table in the <a href=\"https://www.biorxiv.org/content/10.1101/2024.02.24.581671v2\" target=\"_blank\">Ribonanza preprint</a>. </p>\n<p>Good luck and we look forward to more discussions!</p>",
  "messages": [
    {
      "id": 3399018,
      "postDate": "2026-01-30T04:11:05.797Z",
      "content": "<p>Some folks have asked about auxiliary data sources for this challenge's 3D RNA structure prediction problem. </p>\n<p>We wanted to remind everyone that, in addition to this competition's <code>train_labels.csv</code> with 3D coordinates, there are also chemical mapping data that are sensitive to RNA structure and can be acquired at large scale, but remain underexplored for boosting 3D modeling. </p>\n<p>For example, a previous <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding\" target=\"_blank\">Ribonanza Kaggle challenge</a> made publicly available these kinds of experimental data for &gt;2M RNA sequences -- but it's an open question as to whether these data can improve modeling of RNA 3D structure.</p>\n<p>Interested teams might get started with the Ribonanza 'QUICK_START' data set:</p>\n<p><a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data?select=train_data_QUICK_START.csv\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data?select=train_data_QUICK_START.csv</a></p>\n<p>This data set has been curated to hold 2A3 and DMS chemical mapping profiles with reasonable signal to noise on ~170k RNA's with diverse sequences and structural ensembles.</p>\n<p>In addition, we know that many teams are currently incorporating template-based modeling into their methods. </p>\n<p>To help those efforts make contact with the Ribonanza data, <a href=\"https://www.kaggle.com/ckjoshi9\" target=\"_blank\">@ckjoshi9</a> – who is joining the HHMI &amp; Stanford teams co-organizing this competition – has recently made available 3D templates corresponding to these sequences here:</p>\n<p><a href=\"https://www.kaggle.com/datasets/ckjoshi9/ribonanza-quickstart-3d-templates\" target=\"_blank\">https://www.kaggle.com/datasets/ckjoshi9/ribonanza-quickstart-3d-templates</a></p>\n<p>If you find useful signal from those data, note that there's much more publicly available, including the full set of training profiles and test data from the Ribonanza competition (see links in the Supplemental Table in the <a href=\"https://www.biorxiv.org/content/10.1101/2024.02.24.581671v2\" target=\"_blank\">Ribonanza preprint</a>. </p>\n<p>Good luck and we look forward to more discussions!</p>",
      "rawMarkdown": "Some folks have asked about auxiliary data sources for this challenge's 3D RNA structure prediction problem. \n\nWe wanted to remind everyone that, in addition to this competition's `train_labels.csv` with 3D coordinates, there are also chemical mapping data that are sensitive to RNA structure and can be acquired at large scale, but remain underexplored for boosting 3D modeling. \n\nFor example, a previous [Ribonanza Kaggle challenge](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding) made publicly available these kinds of experimental data for >2M RNA sequences -- but it's an open question as to whether these data can improve modeling of RNA 3D structure.\n\nInterested teams might get started with the Ribonanza 'QUICK_START' data set:\n\nhttps://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data?select=train_data_QUICK_START.csv\n\nThis data set has been curated to hold 2A3 and DMS chemical mapping profiles with reasonable signal to noise on ~170k RNA's with diverse sequences and structural ensembles.\n\nIn addition, we know that many teams are currently incorporating template-based modeling into their methods. \n\nTo help those efforts make contact with the Ribonanza data, @ckjoshi9 – who is joining the HHMI & Stanford teams co-organizing this competition – has recently made available 3D templates corresponding to these sequences here:\n\nhttps://www.kaggle.com/datasets/ckjoshi9/ribonanza-quickstart-3d-templates\n\nIf you find useful signal from those data, note that there's much more publicly available, including the full set of training profiles and test data from the Ribonanza competition (see links in the Supplemental Table in the [Ribonanza preprint](https://www.biorxiv.org/content/10.1101/2024.02.24.581671v2). \n\nGood luck and we look forward to more discussions!",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3399018": "Some folks have asked about auxiliary data sources for this challenge's 3D RNA structure prediction problem. \n\nWe wanted to remind everyone that, in addition to this competition's `train_labels.csv` with 3D coordinates, there are also chemical mapping data that are sensitive to RNA structure and can be acquired at large scale, but remain underexplored for boosting 3D modeling. \n\nFor example, a previous [Ribonanza Kaggle challenge](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding) made publicly available these kinds of experimental data for >2M RNA sequences -- but it's an open question as to whether these data can improve modeling of RNA 3D structure.\n\nInterested teams might get started with the Ribonanza 'QUICK_START' data set:\n\nhttps://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/data?select=train_data_QUICK_START.csv\n\nThis data set has been curated to hold 2A3 and DMS chemical mapping profiles with reasonable signal to noise on ~170k RNA's with diverse sequences and structural ensembles.\n\nIn addition, we know that many teams are currently incorporating template-based modeling into their methods. \n\nTo help those efforts make contact with the Ribonanza data, @ckjoshi9 – who is joining the HHMI & Stanford teams co-organizing this competition – has recently made available 3D templates corresponding to these sequences here:\n\nhttps://www.kaggle.com/datasets/ckjoshi9/ribonanza-quickstart-3d-templates\n\nIf you find useful signal from those data, note that there's much more publicly available, including the full set of training profiles and test data from the Ribonanza competition (see links in the Supplemental Table in the [Ribonanza preprint](https://www.biorxiv.org/content/10.1101/2024.02.24.581671v2). \n\nGood luck and we look forward to more discussions!"
  }
}