{
  "id": 430639,
  "title": "I haven't seen this discussed much, but what's up with the difference between train and validation",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/430639",
  "author_name": "DennisSakva",
  "post_date": "2023-08-10T15:01:51.450000",
  "votes": 6,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Here are the average masks for the training and validation sets. Why are the masks for the training set so centered, while those for the validation set are roughly equally distributed across the field of view?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2F340f7f9e04705c9b4bc3f784283881af%2Ftvmasks.png?generation=1691679641521578&amp;alt=media\" alt=\"\"><br>\nCourtesy of Eugene Khvedchenya</p>",
  "messages": [
    {
      "id": 2383792,
      "postDate": "2023-08-10T15:01:51.450Z",
      "content": "<p>Here are the average masks for the training and validation sets. Why are the masks for the training set so centered, while those for the validation set are roughly equally distributed across the field of view?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2F340f7f9e04705c9b4bc3f784283881af%2Ftvmasks.png?generation=1691679641521578&amp;alt=media\" alt=\"\"><br>\nCourtesy of Eugene Khvedchenya</p>",
      "rawMarkdown": "Here are the average masks for the training and validation sets. Why are the masks for the training set so centered, while those for the validation set are roughly equally distributed across the field of view?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2F340f7f9e04705c9b4bc3f784283881af%2Ftvmasks.png?generation=1691679641521578&alt=media)\nCourtesy of Eugene Khvedchenya",
      "votes": 5
    },
    {
      "id": 2383819,
      "postDate": "2023-08-10T15:16:34.360Z",
      "content": "<p>I believe that that is from the google street images that got added in. Some of their images were randomly sampled but to get more positives they specifically tracked down some contrails from street view. Those were probably centered</p>",
      "rawMarkdown": "I believe that that is from the google street images that got added in. Some of their images were randomly sampled but to get more positives they specifically tracked down some contrails from street view. Those were probably centered",
      "votes": 4,
      "replies": [
        {
          "id": 2385319,
          "postDate": "2023-08-11T09:09:56.637Z",
          "content": "<p>I don't think we have google street images in either train or validation sets.</p>",
          "rawMarkdown": "I don't think we have google street images in either train or validation sets.",
          "replies": [
            {
              "id": 2385687,
              "postDate": "2023-08-11T13:55:49.307Z",
              "content": "<p>They found contrails from google street images and then used metadata to  locate the time and location and found them in the GOES images. We dont directly have google street images but we absolutely have them in the training set. That is why if you do validation with training results included your model will do slightly better than if you validate against the folder. </p>\n<blockquote>\n  <p>To further boost the number of positives in the dataset, we<br>\n  also included some GOES-16 ABI imagery at locations in the<br>\n  US where Google Street View images of the sky contained<br>\n  contrails. To define when a Street View image of the sky<br>\n  contained a contrail, we used 64-dimensional image feature<br>\n  vectors derived from image-text data, created with an approach<br>\n  similar to that used by Juan et al. [26]. We applied a threshold<br>\n  to the cosine similarity of the Street View image feature vector<br>\n  and that of a seed image of a contrail taken from the ground;<br>\n  if it was similar enough, GOES-16 imagery at that time and<br>\n  location were sampled for human labeling of contrails. These<br>\n  additional labeled images are only in the training set: because<br>\n  Street View cars operate on days with sunnier weather, it<br>\n  may be easier than usual to identify contrails in the GOES16 imagery of those locations. We found in our experiment<br>\n  that these additional training data slightly improve the contrail<br>\n  detection model performance on the validation set. Note that<br>\n  there are no Street View images in the dataset or used in<br>\n  models reported here: they were only used to identify contrailrich GOES-16 scenes for human labeling.</p>\n</blockquote>",
              "rawMarkdown": "They found contrails from google street images and then used metadata to  locate the time and location and found them in the GOES images. We dont directly have google street images but we absolutely have them in the training set. That is why if you do validation with training results included your model will do slightly better than if you validate against the folder. \n\n> To further boost the number of positives in the dataset, we\nalso included some GOES-16 ABI imagery at locations in the\nUS where Google Street View images of the sky contained\ncontrails. To define when a Street View image of the sky\ncontained a contrail, we used 64-dimensional image feature\nvectors derived from image-text data, created with an approach\nsimilar to that used by Juan et al. [26]. We applied a threshold\nto the cosine similarity of the Street View image feature vector\nand that of a seed image of a contrail taken from the ground;\nif it was similar enough, GOES-16 imagery at that time and\nlocation were sampled for human labeling of contrails. These\nadditional labeled images are only in the training set: because\nStreet View cars operate on days with sunnier weather, it\nmay be easier than usual to identify contrails in the GOES16 imagery of those locations. We found in our experiment\nthat these additional training data slightly improve the contrail\ndetection model performance on the validation set. Note that\nthere are no Street View images in the dataset or used in\nmodels reported here: they were only used to identify contrailrich GOES-16 scenes for human labeling.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2383825,
      "postDate": "2023-08-10T15:18:58.550Z",
      "content": "<p>If you take a random sample from the train set that is the same size as the validation set, does it end up looking more like the validation set? I'm wondering if what you're seeing is just a function of the validation set being a lot smaller.</p>",
      "rawMarkdown": "If you take a random sample from the train set that is the same size as the validation set, does it end up looking more like the validation set? I'm wondering if what you're seeing is just a function of the validation set being a lot smaller.",
      "votes": 1,
      "replies": [
        {
          "id": 2385314,
          "postDate": "2023-08-11T09:07:03.630Z",
          "content": "<p>No, not really. There's something train set specific is going on. Here is a random sample from the train set with a size equal to that of the validation set.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2Fdfd481cbe6e9e4a079b94a8653af4fdb%2Ftrainsample.png?generation=1691744857068589&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "No, not really. There's something train set specific is going on. Here is a random sample from the train set with a size equal to that of the validation set.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2Fdfd481cbe6e9e4a079b94a8653af4fdb%2Ftrainsample.png?generation=1691744857068589&alt=media)"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2383819,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2023-08-10T15:16:34.360000",
      "content": "<p>I believe that that is from the google street images that got added in. Some of their images were randomly sampled but to get more positives they specifically tracked down some contrails from street view. Those were probably centered</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2385319,
          "author_name": "DennisSakva",
          "author_url": "",
          "post_date": "2023-08-11T09:09:56.637000",
          "content": "<p>I don't think we have google street images in either train or validation sets.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2385687,
              "author_name": "ryches",
              "author_url": "",
              "post_date": "2023-08-11T13:55:49.307000",
              "content": "<p>They found contrails from google street images and then used metadata to  locate the time and location and found them in the GOES images. We dont directly have google street images but we absolutely have them in the training set. That is why if you do validation with training results included your model will do slightly better than if you validate against the folder. </p>\n<blockquote>\n  <p>To further boost the number of positives in the dataset, we<br>\n  also included some GOES-16 ABI imagery at locations in the<br>\n  US where Google Street View images of the sky contained<br>\n  contrails. To define when a Street View image of the sky<br>\n  contained a contrail, we used 64-dimensional image feature<br>\n  vectors derived from image-text data, created with an approach<br>\n  similar to that used by Juan et al. [26]. We applied a threshold<br>\n  to the cosine similarity of the Street View image feature vector<br>\n  and that of a seed image of a contrail taken from the ground;<br>\n  if it was similar enough, GOES-16 imagery at that time and<br>\n  location were sampled for human labeling of contrails. These<br>\n  additional labeled images are only in the training set: because<br>\n  Street View cars operate on days with sunnier weather, it<br>\n  may be easier than usual to identify contrails in the GOES16 imagery of those locations. We found in our experiment<br>\n  that these additional training data slightly improve the contrail<br>\n  detection model performance on the validation set. Note that<br>\n  there are no Street View images in the dataset or used in<br>\n  models reported here: they were only used to identify contrailrich GOES-16 scenes for human labeling.</p>\n</blockquote>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2383825,
      "author_name": "Aaron Sarna",
      "author_url": "",
      "post_date": "2023-08-10T15:18:58.550000",
      "content": "<p>If you take a random sample from the train set that is the same size as the validation set, does it end up looking more like the validation set? I'm wondering if what you're seeing is just a function of the validation set being a lot smaller.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2385314,
          "author_name": "DennisSakva",
          "author_url": "",
          "post_date": "2023-08-11T09:07:03.630000",
          "content": "<p>No, not really. There's something train set specific is going on. Here is a random sample from the train set with a size equal to that of the validation set.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2Fdfd481cbe6e9e4a079b94a8653af4fdb%2Ftrainsample.png?generation=1691744857068589&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2383792": "Here are the average masks for the training and validation sets. Why are the masks for the training set so centered, while those for the validation set are roughly equally distributed across the field of view?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2F340f7f9e04705c9b4bc3f784283881af%2Ftvmasks.png?generation=1691679641521578&alt=media)\nCourtesy of Eugene Khvedchenya",
    "2383819": "I believe that that is from the google street images that got added in. Some of their images were randomly sampled but to get more positives they specifically tracked down some contrails from street view. Those were probably centered",
    "2383825": "If you take a random sample from the train set that is the same size as the validation set, does it end up looking more like the validation set? I'm wondering if what you're seeing is just a function of the validation set being a lot smaller."
  }
}