{
  "id": 147603,
  "title": "How different is the train and test set?",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/147603",
  "author_name": "Nicholas Lyu",
  "post_date": "2020-05-01T07:57:49.894000",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>The data tab claimed that \"slightly different procedures were in place for the images used in the test set than the training set\". Here are some of my experiments:</p>\n\n<ol>\n<li>Thresholding to generate mask (threshold value 255-30=225) then perspective crop &amp; resize based on clustered aspect ratio (1:5 for all training data). This yields .83 CV but only .66 LB</li>\n<li>Replicate Iafoss's tile-pooling idea. .77CV as well as LB</li>\n</ol>\n\n<p>While thresholding with value 225 performed great locally, it was very bad on LB. It is possible that \n(1) test images have slightly different background colors (thus messing up the object-selection function), or\n (2) the test images have very different distribution of shape (thus messing up the aspect ratio).</p>\n\n<p>Please kindly share how your CV &amp; LB correlates with different approaches. With different preprocessing and data pipeline CV&amp;LB correlation is quite incomparable.</p>",
  "messages": [
    {
      "id": 828654,
      "postDate": "2020-05-01T07:57:49.893Z",
      "content": "<p>The data tab claimed that \"slightly different procedures were in place for the images used in the test set than the training set\". Here are some of my experiments:</p>\n\n<ol>\n<li>Thresholding to generate mask (threshold value 255-30=225) then perspective crop &amp; resize based on clustered aspect ratio (1:5 for all training data). This yields .83 CV but only .66 LB</li>\n<li>Replicate Iafoss's tile-pooling idea. .77CV as well as LB</li>\n</ol>\n\n<p>While thresholding with value 225 performed great locally, it was very bad on LB. It is possible that \n(1) test images have slightly different background colors (thus messing up the object-selection function), or\n (2) the test images have very different distribution of shape (thus messing up the aspect ratio).</p>\n\n<p>Please kindly share how your CV &amp; LB correlates with different approaches. With different preprocessing and data pipeline CV&amp;LB correlation is quite incomparable.</p>",
      "rawMarkdown": "The data tab claimed that \"slightly different procedures were in place for the images used in the test set than the training set\". Here are some of my experiments:\n\n1. Thresholding to generate mask (threshold value 255-30=225) then perspective crop &amp; resize based on clustered aspect ratio (1:5 for all training data). This yields .83 CV but only .66 LB\n2. Replicate Iafoss's tile-pooling idea. .77CV as well as LB\n\nWhile thresholding with value 225 performed great locally, it was very bad on LB. It is possible that \n(1) test images have slightly different background colors (thus messing up the object-selection function), or\n (2) the test images have very different distribution of shape (thus messing up the aspect ratio).\n\nPlease kindly share how your CV &amp; LB correlates with different approaches. With different preprocessing and data pipeline CV&amp;LB correlation is quite incomparable.",
      "votes": 3
    },
    {
      "id": 828877,
      "postDate": "2020-05-01T10:52:04.293Z",
      "content": "<p>There are no visual or technical differences between the test set and train set. All images were collected and then a selection was included in the test set. The slight difference referred to is that a few of the Radboud images contain pen marks, but since these pen marks can possibly be an indicator of cancer, we have made sure that there are no such images in the test set. </p>",
      "rawMarkdown": "There are no visual or technical differences between the test set and train set. All images were collected and then a selection was included in the test set. The slight difference referred to is that a few of the Radboud images contain pen marks, but since these pen marks can possibly be an indicator of cancer, we have made sure that there are no such images in the test set. ",
      "votes": 4
    }
  ],
  "comments": [
    {
      "id": 828877,
      "author_name": "PeterStröm",
      "author_url": "",
      "post_date": "2020-05-01T10:52:04.293000",
      "content": "<p>There are no visual or technical differences between the test set and train set. All images were collected and then a selection was included in the test set. The slight difference referred to is that a few of the Radboud images contain pen marks, but since these pen marks can possibly be an indicator of cancer, we have made sure that there are no such images in the test set. </p>",
      "votes": 4,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "828654": "The data tab claimed that \"slightly different procedures were in place for the images used in the test set than the training set\". Here are some of my experiments:\n\n1. Thresholding to generate mask (threshold value 255-30=225) then perspective crop &amp; resize based on clustered aspect ratio (1:5 for all training data). This yields .83 CV but only .66 LB\n2. Replicate Iafoss's tile-pooling idea. .77CV as well as LB\n\nWhile thresholding with value 225 performed great locally, it was very bad on LB. It is possible that \n(1) test images have slightly different background colors (thus messing up the object-selection function), or\n (2) the test images have very different distribution of shape (thus messing up the aspect ratio).\n\nPlease kindly share how your CV &amp; LB correlates with different approaches. With different preprocessing and data pipeline CV&amp;LB correlation is quite incomparable.",
    "828877": "There are no visual or technical differences between the test set and train set. All images were collected and then a selection was included in the test set. The slight difference referred to is that a few of the Radboud images contain pen marks, but since these pen marks can possibly be an indicator of cancer, we have made sure that there are no such images in the test set. "
  }
}