{
  "id": 681153,
  "title": "Inquiry about the dataset",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/681153",
  "author_name": "Zelin Huang",
  "post_date": "2026-03-13T06:03:40.123000",
  "votes": 0,
  "comment_count": 1,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> <a href=\"https://www.kaggle.com/przemekporebski\" target=\"_blank\">@przemekporebski</a> I have a question about the dataset. Are we only allowed to use the dataset provided in this competition, or can we obtain additional datasets from other sources? If we use external datasets, could that potentially cause data leakage—for example, overlapping with the private test set used for the competition leaderboard?</p>",
  "messages": [
    {
      "id": 3420516,
      "postDate": "2026-03-13T13:29:11.333Z",
      "content": "<p>You can use any external data set. </p>\n<p>The leaderboard targets are all confidential RNAs whose sequences and coordinates are known only to hosts and our collaborators (and very briefly to your notebooks when they run behind the Kaggle firewall!). So there should not be leakage.</p>\n<p>One thing to reiterate - your notebooks <em>can</em> access any of the LLMs available via Kaggle API or that you attach as a dataset or model during development. </p>\n<p>Even if those LLMs have memorized whatever information is publicly available in, for example, the scientific literature or online databases, that is totally fine for this competition — and I believe that accessing that information should give your notebooks an edge.</p>",
      "rawMarkdown": "You can use any external data set. \n\nThe leaderboard targets are all confidential RNAs whose sequences and coordinates are known only to hosts and our collaborators (and very briefly to your notebooks when they run behind the Kaggle firewall!). So there should not be leakage.\n\nOne thing to reiterate - your notebooks *can* access any of the LLMs available via Kaggle API or that you attach as a dataset or model during development. \n\nEven if those LLMs have memorized whatever information is publicly available in, for example, the scientific literature or online databases, that is totally fine for this competition — and I believe that accessing that information should give your notebooks an edge."
    },
    {
      "id": 3420374,
      "postDate": "2026-03-13T06:03:40.123Z",
      "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> <a href=\"https://www.kaggle.com/przemekporebski\" target=\"_blank\">@przemekporebski</a> I have a question about the dataset. Are we only allowed to use the dataset provided in this competition, or can we obtain additional datasets from other sources? If we use external datasets, could that potentially cause data leakage—for example, overlapping with the private test set used for the competition leaderboard?</p>",
      "rawMarkdown": "@rhijudas @przemekporebski I have a question about the dataset. Are we only allowed to use the dataset provided in this competition, or can we obtain additional datasets from other sources? If we use external datasets, could that potentially cause data leakage—for example, overlapping with the private test set used for the competition leaderboard?"
    }
  ],
  "comments": [
    {
      "id": 3420516,
      "author_name": "Rhiju Das",
      "author_url": "",
      "post_date": "2026-03-13T13:29:11.333000",
      "content": "<p>You can use any external data set. </p>\n<p>The leaderboard targets are all confidential RNAs whose sequences and coordinates are known only to hosts and our collaborators (and very briefly to your notebooks when they run behind the Kaggle firewall!). So there should not be leakage.</p>\n<p>One thing to reiterate - your notebooks <em>can</em> access any of the LLMs available via Kaggle API or that you attach as a dataset or model during development. </p>\n<p>Even if those LLMs have memorized whatever information is publicly available in, for example, the scientific literature or online databases, that is totally fine for this competition — and I believe that accessing that information should give your notebooks an edge.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3420516": "You can use any external data set. \n\nThe leaderboard targets are all confidential RNAs whose sequences and coordinates are known only to hosts and our collaborators (and very briefly to your notebooks when they run behind the Kaggle firewall!). So there should not be leakage.\n\nOne thing to reiterate - your notebooks *can* access any of the LLMs available via Kaggle API or that you attach as a dataset or model during development. \n\nEven if those LLMs have memorized whatever information is publicly available in, for example, the scientific literature or online databases, that is totally fine for this competition — and I believe that accessing that information should give your notebooks an edge.",
    "3420374": "@rhijudas @przemekporebski I have a question about the dataset. Are we only allowed to use the dataset provided in this competition, or can we obtain additional datasets from other sources? If we use external datasets, could that potentially cause data leakage—for example, overlapping with the private test set used for the competition leaderboard?"
  }
}