{
  "id": 420532,
  "title": "Does test set have different distribution with train/validation set?",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/420532",
  "author_name": "Tawara",
  "post_date": "2023-07-01T08:14:33.290000",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi <a href=\"https://www.kaggle.com/aaronsarna\" target=\"_blank\">@aaronsarna</a>, I have a question about test set.</p>\n<p>According to the hosts' <a href=\"https://arxiv.org/abs/2304.02122\" target=\"_blank\">preprint</a> (p.4),</p>\n<blockquote>\n  <p>The full dataset contains 20,544 examples in the trainset and 1,866 examples in the validation set.</p>\n</blockquote>\n<p>and this competiton gives us 20,529 examples in train folder and 1,856 examples in  validation foloder, therefore dataset in preprint and competition seem almost same.</p>\n<p>My concern here is where and when was the test set collected? The preprint didn't mention test sets.</p>\n<p>Train set has more examples having contrails than validation set but where and when they were collected are almost same (we can check this by metadata.)</p>\n<p>I'm afraid that big shake will be coming in the end of competition If test set distribution is much different from train/validation.</p>\n<p>I'd like to know if there is a much difference in test set's distribution with train/validation set, <strong>not details of test set's metadata</strong>.</p>",
  "messages": [
    {
      "id": 2325273,
      "postDate": "2023-07-01T08:14:33.290Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/aaronsarna\" target=\"_blank\">@aaronsarna</a>, I have a question about test set.</p>\n<p>According to the hosts' <a href=\"https://arxiv.org/abs/2304.02122\" target=\"_blank\">preprint</a> (p.4),</p>\n<blockquote>\n  <p>The full dataset contains 20,544 examples in the trainset and 1,866 examples in the validation set.</p>\n</blockquote>\n<p>and this competiton gives us 20,529 examples in train folder and 1,856 examples in  validation foloder, therefore dataset in preprint and competition seem almost same.</p>\n<p>My concern here is where and when was the test set collected? The preprint didn't mention test sets.</p>\n<p>Train set has more examples having contrails than validation set but where and when they were collected are almost same (we can check this by metadata.)</p>\n<p>I'm afraid that big shake will be coming in the end of competition If test set distribution is much different from train/validation.</p>\n<p>I'd like to know if there is a much difference in test set's distribution with train/validation set, <strong>not details of test set's metadata</strong>.</p>",
      "rawMarkdown": "Hi @aaronsarna, I have a question about test set.\n\nAccording to the hosts' [preprint](https://arxiv.org/abs/2304.02122) (p.4),\n\n>The full dataset contains 20,544 examples in the trainset and 1,866 examples in the validation set.\n\nand this competiton gives us 20,529 examples in train folder and 1,856 examples in  validation foloder, therefore dataset in preprint and competition seem almost same.\n\nMy concern here is where and when was the test set collected? The preprint didn't mention test sets.\n\nTrain set has more examples having contrails than validation set but where and when they were collected are almost same (we can check this by metadata.)\n\nI'm afraid that big shake will be coming in the end of competition If test set distribution is much different from train/validation.\n\nI'd like to know if there is a much difference in test set's distribution with train/validation set, **not details of test set's metadata**.",
      "votes": 3
    },
    {
      "id": 2325323,
      "postDate": "2023-07-01T08:52:27.993Z",
      "content": "<p>Just my random thought but the host might want to hide the detail of the test dataset if they are collected from public dataset, and label by themselves. Satellite images are difficult to collect by individuals/companies, so I guess test dataset is also taken from the public dataset.</p>",
      "rawMarkdown": "Just my random thought but the host might want to hide the detail of the test dataset if they are collected from public dataset, and label by themselves. Satellite images are difficult to collect by individuals/companies, so I guess test dataset is also taken from the public dataset.",
      "votes": 1,
      "replies": [
        {
          "id": 2325357,
          "postDate": "2023-07-01T09:19:20.537Z",
          "content": "<p>Thank you for your thought. If that were true, big shake may be coming…</p>",
          "rawMarkdown": "Thank you for your thought. If that were true, big shake may be coming...",
          "replies": [
            {
              "id": 2325385,
              "postDate": "2023-07-01T09:44:50.240Z",
              "content": "<p>You can just LB probe if you want to check if test set distribution is close to train set. Besides, since public test set is very small (few hundreds), we should more or less build model with proper generalization whether test data information is provided or not.</p>",
              "rawMarkdown": "You can just LB probe if you want to check if test set distribution is close to train set. Besides, since public test set is very small (few hundreds), we should more or less build model with proper generalization whether test data information is provided or not.",
              "votes": 1
            },
            {
              "id": 2325508,
              "postDate": "2023-07-01T12:22:53.120Z",
              "content": "<p>I did that and came up with a 20-30% distribution in the test set (of images containing at least one positive pixel)</p>",
              "rawMarkdown": "I did that and came up with a 20-30% distribution in the test set (of images containing at least one positive pixel)",
              "votes": 3
            },
            {
              "id": 2325711,
              "postDate": "2023-07-01T14:51:16.200Z",
              "content": "<p><a href=\"https://www.kaggle.com/janmpia\" target=\"_blank\">@janmpia</a> That's an interesting result. However, how did you obtain this? In other words, how were you able to determine whether the image contains a positive pixel or not, despite no labels are given?</p>",
              "rawMarkdown": "@janmpia That's an interesting result. However, how did you obtain this? In other words, how were you able to determine whether the image contains a positive pixel or not, despite no labels are given?"
            },
            {
              "id": 2325745,
              "postDate": "2023-07-01T15:11:33.873Z",
              "content": "<p>I trained a classifier that has 90% accuracy on CV, and did some submission time tricks. </p>",
              "rawMarkdown": "I trained a classifier that has 90% accuracy on CV, and did some submission time tricks. ",
              "votes": 2
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2325323,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2023-07-01T08:52:27.993000",
      "content": "<p>Just my random thought but the host might want to hide the detail of the test dataset if they are collected from public dataset, and label by themselves. Satellite images are difficult to collect by individuals/companies, so I guess test dataset is also taken from the public dataset.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2325357,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2023-07-01T09:19:20.537000",
          "content": "<p>Thank you for your thought. If that were true, big shake may be coming…</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2325385,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2023-07-01T09:44:50.240000",
              "content": "<p>You can just LB probe if you want to check if test set distribution is close to train set. Besides, since public test set is very small (few hundreds), we should more or less build model with proper generalization whether test data information is provided or not.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2325508,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-07-01T12:22:53.120000",
              "content": "<p>I did that and came up with a 20-30% distribution in the test set (of images containing at least one positive pixel)</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2325711,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2023-07-01T14:51:16.200000",
              "content": "<p><a href=\"https://www.kaggle.com/janmpia\" target=\"_blank\">@janmpia</a> That's an interesting result. However, how did you obtain this? In other words, how were you able to determine whether the image contains a positive pixel or not, despite no labels are given?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2325745,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-07-01T15:11:33.873000",
              "content": "<p>I trained a classifier that has 90% accuracy on CV, and did some submission time tricks. </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2325273": "Hi @aaronsarna, I have a question about test set.\n\nAccording to the hosts' [preprint](https://arxiv.org/abs/2304.02122) (p.4),\n\n>The full dataset contains 20,544 examples in the trainset and 1,866 examples in the validation set.\n\nand this competiton gives us 20,529 examples in train folder and 1,856 examples in  validation foloder, therefore dataset in preprint and competition seem almost same.\n\nMy concern here is where and when was the test set collected? The preprint didn't mention test sets.\n\nTrain set has more examples having contrails than validation set but where and when they were collected are almost same (we can check this by metadata.)\n\nI'm afraid that big shake will be coming in the end of competition If test set distribution is much different from train/validation.\n\nI'd like to know if there is a much difference in test set's distribution with train/validation set, **not details of test set's metadata**.",
    "2325323": "Just my random thought but the host might want to hide the detail of the test dataset if they are collected from public dataset, and label by themselves. Satellite images are difficult to collect by individuals/companies, so I guess test dataset is also taken from the public dataset."
  }
}