{
  "id": 495129,
  "title": "A theoretically interesting approach",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/495129",
  "author_name": "asarvazyan",
  "post_date": "2024-04-19T19:05:04.568000",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Seeing the large scale of the dataset I was wondering what would be the score of a solution that predicts the target values of the row with most similar features in the training set. </p>\n<p>What's stopping someone from running something like this for the whole dataset (or, e.g. using only on the first 1M training samples) in the background until the competition is over? Granted it could take a while, but would be interesting to see if a naive solution without learning can surpass some learned baselines</p>",
  "messages": [
    {
      "id": 2761308,
      "postDate": "2024-04-19T19:05:04.570Z",
      "content": "<p>Seeing the large scale of the dataset I was wondering what would be the score of a solution that predicts the target values of the row with most similar features in the training set. </p>\n<p>What's stopping someone from running something like this for the whole dataset (or, e.g. using only on the first 1M training samples) in the background until the competition is over? Granted it could take a while, but would be interesting to see if a naive solution without learning can surpass some learned baselines</p>",
      "rawMarkdown": "Seeing the large scale of the dataset I was wondering what would be the score of a solution that predicts the target values of the row with most similar features in the training set. \n\nWhat's stopping someone from running something like this for the whole dataset (or, e.g. using only on the first 1M training samples) in the background until the competition is over? Granted it could take a while, but would be interesting to see if a naive solution without learning can surpass some learned baselines",
      "votes": 6
    },
    {
      "id": 2767466,
      "postDate": "2024-04-22T10:27:06.903Z",
      "content": "<blockquote>\n  <p>Seeing the large scale of the dataset I was wondering what would be the score of a solution that predicts the target values of the row with most similar features in the training set.</p>\n</blockquote>\n<p>So something like a K Nearest Neighbors algorithm, chunk by chunk? It did come to my mind, but I've tried this approach in a previous competition and the performance was just alright (<a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/leaderboard\" target=\"_blank\">Ranked 80th out of 500~</a>).The size of the dataset in this contest is even more extreme. One could still try it though.</p>",
      "rawMarkdown": "> Seeing the large scale of the dataset I was wondering what would be the score of a solution that predicts the target values of the row with most similar features in the training set.\n\nSo something like a K Nearest Neighbors algorithm, chunk by chunk? It did come to my mind, but I've tried this approach in a previous competition and the performance was just alright ([Ranked 80th out of 500~](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/leaderboard)).The size of the dataset in this contest is even more extreme. One could still try it though."
    },
    {
      "id": 2761579,
      "postDate": "2024-04-19T23:54:23.427Z",
      "content": "<p>With 25 variables, even if we bin each variable only to 5 values, it's already 3e17 different combinations, i.e., train data may be large, but the possible combinations space is much much, much larger…so good luck trying this trick, miracles are always possible haha</p>",
      "rawMarkdown": "With 25 variables, even if we bin each variable only to 5 values, it's already 3e17 different combinations, i.e., train data may be large, but the possible combinations space is much much, much larger...so good luck trying this trick, miracles are always possible haha"
    }
  ],
  "comments": [
    {
      "id": 2767466,
      "author_name": "Aatif Fraz",
      "author_url": "",
      "post_date": "2024-04-22T10:27:06.903000",
      "content": "<blockquote>\n  <p>Seeing the large scale of the dataset I was wondering what would be the score of a solution that predicts the target values of the row with most similar features in the training set.</p>\n</blockquote>\n<p>So something like a K Nearest Neighbors algorithm, chunk by chunk? It did come to my mind, but I've tried this approach in a previous competition and the performance was just alright (<a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/leaderboard\" target=\"_blank\">Ranked 80th out of 500~</a>).The size of the dataset in this contest is even more extreme. One could still try it though.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2761579,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-04-19T23:54:23.427000",
      "content": "<p>With 25 variables, even if we bin each variable only to 5 values, it's already 3e17 different combinations, i.e., train data may be large, but the possible combinations space is much much, much larger…so good luck trying this trick, miracles are always possible haha</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2761308": "Seeing the large scale of the dataset I was wondering what would be the score of a solution that predicts the target values of the row with most similar features in the training set. \n\nWhat's stopping someone from running something like this for the whole dataset (or, e.g. using only on the first 1M training samples) in the background until the competition is over? Granted it could take a while, but would be interesting to see if a naive solution without learning can surpass some learned baselines",
    "2767466": "> Seeing the large scale of the dataset I was wondering what would be the score of a solution that predicts the target values of the row with most similar features in the training set.\n\nSo something like a K Nearest Neighbors algorithm, chunk by chunk? It did come to my mind, but I've tried this approach in a previous competition and the performance was just alright ([Ranked 80th out of 500~](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/leaderboard)).The size of the dataset in this contest is even more extreme. One could still try it though.",
    "2761579": "With 25 variables, even if we bin each variable only to 5 values, it's already 3e17 different combinations, i.e., train data may be large, but the possible combinations space is much much, much larger...so good luck trying this trick, miracles are always possible haha"
  }
}