{
  "id": 415764,
  "title": "Did the labelers know about the test set ?",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/415764",
  "author_name": "JEANMPIA",
  "post_date": "2023-06-08T00:21:44.981000",
  "votes": 11,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello <a href=\"https://www.kaggle.com/aaronsarna\" target=\"_blank\">@aaronsarna</a>,<br>\nI have a question regarding the train/test split. (not train/valid, I count valid as part of train)<br>\n<strong>Were the labellers informed that what they were labelling was either the train or the test set ?</strong><br>\n<em>My question might sound anecdotal but take a look:</em><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2Fea0fa6bd03a87e1da55d1101c7fddcbd%2FScreenshot_101.jpg?generation=1686183194357307&amp;alt=media\" alt=\"\"><br>\n<em>the 4 first labels are layers of the human_masks and the last is the pixel_mask</em></p>\n<p>I don't know for sure those are con-trails, but thats really not the only image where some of the labellers didn't label some con-trails where other did, often resulting in no-contrail labels.<br>\nI plan on changing the pixel mask for this image to one of the ones that saw the con-trail, would that be foolish since the same amounth of care was given to the test set, or did they know that this set was going to be important so they took extra attention ? (no disrespect to the labellers of course I get that this job requires a lot of focus and I'm glad I'm not on this side of the pipeline)</p>",
  "messages": [
    {
      "id": 2291944,
      "postDate": "2023-06-08T01:10:55.887Z",
      "content": "<p>The labelers did not know if a sample was going to be in train, test, or validation. The splits were determined using random sampling after the labeling was completed.</p>",
      "rawMarkdown": "The labelers did not know if a sample was going to be in train, test, or validation. The splits were determined using random sampling after the labeling was completed.",
      "votes": 10,
      "replies": [
        {
          "id": 2291945,
          "postDate": "2023-06-08T01:16:55.217Z",
          "content": "<p>Thx for your fast answer ! really appreciated,<br>\ngood to know as well. Noted.</p>",
          "rawMarkdown": "Thx for your fast answer ! really appreciated,\ngood to know as well. Noted."
        }
      ]
    },
    {
      "id": 2291920,
      "postDate": "2023-06-08T00:21:44.983Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/aaronsarna\" target=\"_blank\">@aaronsarna</a>,<br>\nI have a question regarding the train/test split. (not train/valid, I count valid as part of train)<br>\n<strong>Were the labellers informed that what they were labelling was either the train or the test set ?</strong><br>\n<em>My question might sound anecdotal but take a look:</em><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2Fea0fa6bd03a87e1da55d1101c7fddcbd%2FScreenshot_101.jpg?generation=1686183194357307&amp;alt=media\" alt=\"\"><br>\n<em>the 4 first labels are layers of the human_masks and the last is the pixel_mask</em></p>\n<p>I don't know for sure those are con-trails, but thats really not the only image where some of the labellers didn't label some con-trails where other did, often resulting in no-contrail labels.<br>\nI plan on changing the pixel mask for this image to one of the ones that saw the con-trail, would that be foolish since the same amounth of care was given to the test set, or did they know that this set was going to be important so they took extra attention ? (no disrespect to the labellers of course I get that this job requires a lot of focus and I'm glad I'm not on this side of the pipeline)</p>",
      "rawMarkdown": "Hello @aaronsarna,\nI have a question regarding the train/test split. (not train/valid, I count valid as part of train)\n**Were the labellers informed that what they were labelling was either the train or the test set ?**\n*My question might sound anecdotal but take a look:*\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2Fea0fa6bd03a87e1da55d1101c7fddcbd%2FScreenshot_101.jpg?generation=1686183194357307&alt=media)\n*the 4 first labels are layers of the human_masks and the last is the pixel_mask*\n\nI don't know for sure those are con-trails, but thats really not the only image where some of the labellers didn't label some con-trails where other did, often resulting in no-contrail labels.\nI plan on changing the pixel mask for this image to one of the ones that saw the con-trail, would that be foolish since the same amounth of care was given to the test set, or did they know that this set was going to be important so they took extra attention ? (no disrespect to the labellers of course I get that this job requires a lot of focus and I'm glad I'm not on this side of the pipeline)\n",
      "votes": 10
    },
    {
      "id": 2291923,
      "postDate": "2023-06-08T00:28:07.040Z",
      "content": "<p><em>other examples of mask that might prevent the model from training correctely:</em></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F61d08133cd8eab8ea3462649d26e5263%2FScreenshot_103.jpg?generation=1686184278301327&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F458a6b124edc9c83af4440b09808ab34%2FScreenshot_102.jpg?generation=1686184063362719&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "*other examples of mask that might prevent the model from training correctely:*\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F61d08133cd8eab8ea3462649d26e5263%2FScreenshot_103.jpg?generation=1686184278301327&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F458a6b124edc9c83af4440b09808ab34%2FScreenshot_102.jpg?generation=1686184063362719&alt=media)",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2291944,
      "author_name": "Aaron Sarna",
      "author_url": "",
      "post_date": "2023-06-08T01:10:55.887000",
      "content": "<p>The labelers did not know if a sample was going to be in train, test, or validation. The splits were determined using random sampling after the labeling was completed.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 2291945,
          "author_name": "JEANMPIA",
          "author_url": "",
          "post_date": "2023-06-08T01:16:55.217000",
          "content": "<p>Thx for your fast answer ! really appreciated,<br>\ngood to know as well. Noted.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2291923,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2023-06-08T00:28:07.040000",
      "content": "<p><em>other examples of mask that might prevent the model from training correctely:</em></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F61d08133cd8eab8ea3462649d26e5263%2FScreenshot_103.jpg?generation=1686184278301327&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F458a6b124edc9c83af4440b09808ab34%2FScreenshot_102.jpg?generation=1686184063362719&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2291944": "The labelers did not know if a sample was going to be in train, test, or validation. The splits were determined using random sampling after the labeling was completed.",
    "2291920": "Hello @aaronsarna,\nI have a question regarding the train/test split. (not train/valid, I count valid as part of train)\n**Were the labellers informed that what they were labelling was either the train or the test set ?**\n*My question might sound anecdotal but take a look:*\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2Fea0fa6bd03a87e1da55d1101c7fddcbd%2FScreenshot_101.jpg?generation=1686183194357307&alt=media)\n*the 4 first labels are layers of the human_masks and the last is the pixel_mask*\n\nI don't know for sure those are con-trails, but thats really not the only image where some of the labellers didn't label some con-trails where other did, often resulting in no-contrail labels.\nI plan on changing the pixel mask for this image to one of the ones that saw the con-trail, would that be foolish since the same amounth of care was given to the test set, or did they know that this set was going to be important so they took extra attention ? (no disrespect to the labellers of course I get that this job requires a lot of focus and I'm glad I'm not on this side of the pipeline)\n",
    "2291923": "*other examples of mask that might prevent the model from training correctely:*\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F61d08133cd8eab8ea3462649d26e5263%2FScreenshot_103.jpg?generation=1686184278301327&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F458a6b124edc9c83af4440b09808ab34%2FScreenshot_102.jpg?generation=1686184063362719&alt=media)"
  }
}