{
  "id": 347875,
  "title": "Using center_id as a feature?",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/347875",
  "author_name": "Martin Kovacevic Buvinic",
  "post_date": "2022-08-25T18:35:12.528000",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Does anyone know if the test set have the same center_ids as the train set? Maybe it can be used a feature in the case they are.</p>",
  "messages": [
    {
      "id": 1914113,
      "postDate": "2022-08-25T18:35:12.530Z",
      "content": "<p>Does anyone know if the test set have the same center_ids as the train set? Maybe it can be used a feature in the case they are.</p>",
      "rawMarkdown": "Does anyone know if the test set have the same center_ids as the train set? Maybe it can be used a feature in the case they are.",
      "votes": 4
    },
    {
      "id": 1961702,
      "postDate": "2022-09-29T10:18:19Z",
      "content": "<p>I tried to use <code>center_id</code> as a feature at one point. We know there are n=11 centers in the training set. When I run a notebook with this one-hot encoder on the hidden test set:</p>\n<pre><code>center_ids = np.arange(start=1.0, step=1.0, stop=12.0)\nonehot = OneHotEncoder(categories=center_ids.reshape(1, -1), sparse=False)\nfit_onehot = onehot.fit_transform(center_id)\n</code></pre>\n<p>The notebook returns an error. When I change one line to:</p>\n<pre><code>onehot = OneHotEncoder(categories=center_ids.reshape(1, -1), sparse=False, handle_unknown=\"ignore\")\n</code></pre>\n<p>The notebook returns correctly. I did notice that in <a href=\"https://jnis.bmj.com/content/neurintsurg/12/Suppl_1/A2.1.full.pdf\" target=\"_blank\">one abstract</a> they mention n=12 centers, not n=11 centers. So my conclusion is that there are likely centers in the hidden test set that are not in the training set, which makes using <code>center_id</code> challenging.</p>\n<p>Have you tried using <code>center_id</code> and if so what was your experience?</p>",
      "rawMarkdown": "I tried to use `center_id` as a feature at one point. We know there are n=11 centers in the training set. When I run a notebook with this one-hot encoder on the hidden test set:\n\n```\ncenter_ids = np.arange(start=1.0, step=1.0, stop=12.0)\nonehot = OneHotEncoder(categories=center_ids.reshape(1, -1), sparse=False)\nfit_onehot = onehot.fit_transform(center_id)\n```\n\nThe notebook returns an error. When I change one line to:\n\n```\nonehot = OneHotEncoder(categories=center_ids.reshape(1, -1), sparse=False, handle_unknown=\"ignore\")\n```\n\nThe notebook returns correctly. I did notice that in [one abstract](https://jnis.bmj.com/content/neurintsurg/12/Suppl_1/A2.1.full.pdf) they mention n=12 centers, not n=11 centers. So my conclusion is that there are likely centers in the hidden test set that are not in the training set, which makes using `center_id` challenging.\n\nHave you tried using `center_id` and if so what was your experience?\n\n",
      "replies": [
        {
          "id": 1965024,
          "postDate": "2022-10-01T04:41:31.537Z",
          "content": "<p>I believe there was some discussion on this. I think the center_id represents the variation that each center producing slides could have in terms of stain coloring. This could be used to stratify data while splitting for train-test. I have used it, but honestly cannot say it helped a lot, although I have changed so many parameters by now its hard to keep track :D</p>\n<p>But as the comment below mentions, all papers on stain normalization and color normalization/randomization say it makes predictions better. So probably that's one way to deal with stain variation.</p>",
          "rawMarkdown": "I believe there was some discussion on this. I think the center_id represents the variation that each center producing slides could have in terms of stain coloring. This could be used to stratify data while splitting for train-test. I have used it, but honestly cannot say it helped a lot, although I have changed so many parameters by now its hard to keep track :D\n\nBut as the comment below mentions, all papers on stain normalization and color normalization/randomization say it makes predictions better. So probably that's one way to deal with stain variation."
        }
      ]
    },
    {
      "id": 1952939,
      "postDate": "2022-09-24T05:06:39.727Z",
      "content": "<p>I think applying some color based augmentations and stain normalization is better choice.</p>",
      "rawMarkdown": "I think applying some color based augmentations and stain normalization is better choice."
    }
  ],
  "comments": [
    {
      "id": 1961702,
      "author_name": "Joe Marturano",
      "author_url": "",
      "post_date": "2022-09-29T10:18:19",
      "content": "<p>I tried to use <code>center_id</code> as a feature at one point. We know there are n=11 centers in the training set. When I run a notebook with this one-hot encoder on the hidden test set:</p>\n<pre><code>center_ids = np.arange(start=1.0, step=1.0, stop=12.0)\nonehot = OneHotEncoder(categories=center_ids.reshape(1, -1), sparse=False)\nfit_onehot = onehot.fit_transform(center_id)\n</code></pre>\n<p>The notebook returns an error. When I change one line to:</p>\n<pre><code>onehot = OneHotEncoder(categories=center_ids.reshape(1, -1), sparse=False, handle_unknown=\"ignore\")\n</code></pre>\n<p>The notebook returns correctly. I did notice that in <a href=\"https://jnis.bmj.com/content/neurintsurg/12/Suppl_1/A2.1.full.pdf\" target=\"_blank\">one abstract</a> they mention n=12 centers, not n=11 centers. So my conclusion is that there are likely centers in the hidden test set that are not in the training set, which makes using <code>center_id</code> challenging.</p>\n<p>Have you tried using <code>center_id</code> and if so what was your experience?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1965024,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-10-01T04:41:31.537000",
          "content": "<p>I believe there was some discussion on this. I think the center_id represents the variation that each center producing slides could have in terms of stain coloring. This could be used to stratify data while splitting for train-test. I have used it, but honestly cannot say it helped a lot, although I have changed so many parameters by now its hard to keep track :D</p>\n<p>But as the comment below mentions, all papers on stain normalization and color normalization/randomization say it makes predictions better. So probably that's one way to deal with stain variation.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1952939,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2022-09-24T05:06:39.727000",
      "content": "<p>I think applying some color based augmentations and stain normalization is better choice.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1914113": "Does anyone know if the test set have the same center_ids as the train set? Maybe it can be used a feature in the case they are.",
    "1961702": "I tried to use `center_id` as a feature at one point. We know there are n=11 centers in the training set. When I run a notebook with this one-hot encoder on the hidden test set:\n\n```\ncenter_ids = np.arange(start=1.0, step=1.0, stop=12.0)\nonehot = OneHotEncoder(categories=center_ids.reshape(1, -1), sparse=False)\nfit_onehot = onehot.fit_transform(center_id)\n```\n\nThe notebook returns an error. When I change one line to:\n\n```\nonehot = OneHotEncoder(categories=center_ids.reshape(1, -1), sparse=False, handle_unknown=\"ignore\")\n```\n\nThe notebook returns correctly. I did notice that in [one abstract](https://jnis.bmj.com/content/neurintsurg/12/Suppl_1/A2.1.full.pdf) they mention n=12 centers, not n=11 centers. So my conclusion is that there are likely centers in the hidden test set that are not in the training set, which makes using `center_id` challenging.\n\nHave you tried using `center_id` and if so what was your experience?\n\n",
    "1952939": "I think applying some color based augmentations and stain normalization is better choice."
  }
}