{
  "id": 376232,
  "title": "Patient Overlap and Data Leakage",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/376232",
  "author_name": "noura bentaher",
  "post_date": "2023-01-05T11:32:32.048000",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello everyone, I m trying to remove the overlap of patients by checking to see if a patient's ID appears in both the training set and the test set and should also verify that I don't have patient overlap in the training and validation sets. In this data set, we have 11913 unique ID. So I want to know how can I avoid this overlapping and split my data in train, test, and validation.</p>",
  "messages": [
    {
      "id": 2087378,
      "postDate": "2023-01-05T15:01:02.690Z",
      "content": "<p>You can check Stratified Group K Fold (<a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedGroupKFold.html)\" target=\"_blank\">https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedGroupKFold.html)</a>. This will allow you to group the data by patients and divide them into K - cross-validation subsets. Thanks to this procedure, in any fold the patients will not be repeated.</p>\n<p>In addition, if you train K - models, each validated on a different subgroup, and trained on the rest of the data, you can then select the best models and try to combine their results, this way you benefit from all the available data.</p>",
      "rawMarkdown": "You can check Stratified Group K Fold (https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedGroupKFold.html). This will allow you to group the data by patients and divide them into K - cross-validation subsets. Thanks to this procedure, in any fold the patients will not be repeated.\n\nIn addition, if you train K - models, each validated on a different subgroup, and trained on the rest of the data, you can then select the best models and try to combine their results, this way you benefit from all the available data.",
      "votes": 5,
      "replies": [
        {
          "id": 2090388,
          "postDate": "2023-01-07T09:20:30.603Z",
          "content": "<p>Thank you for this information!! good luck</p>",
          "rawMarkdown": "Thank you for this information!! good luck",
          "votes": 1
        },
        {
          "id": 2127554,
          "postDate": "2023-02-03T01:02:59.173Z",
          "content": "<p><a href=\"https://www.kaggle.com/dawidstachowiak\" target=\"_blank\">@dawidstachowiak</a> Do I remember correctly that this accepts also the <code>y</code> parameter in addition to <code>group</code>? Do you mind my asking what metadata column you have used there? That could be the <strong>density</strong> column (combined with <strong>cancer</strong> column or the <strong>birads</strong> column)? Or is it better to train a separate model for each density?</p>",
          "rawMarkdown": "@dawidstachowiak Do I remember correctly that this accepts also the `y` parameter in addition to `group`? Do you mind my asking what metadata column you have used there? That could be the **density** column (combined with **cancer** column or the **birads** column)? Or is it better to train a separate model for each density?"
        }
      ]
    },
    {
      "id": 2087163,
      "postDate": "2023-01-05T11:32:32.050Z",
      "content": "<p>Hello everyone, I m trying to remove the overlap of patients by checking to see if a patient's ID appears in both the training set and the test set and should also verify that I don't have patient overlap in the training and validation sets. In this data set, we have 11913 unique ID. So I want to know how can I avoid this overlapping and split my data in train, test, and validation.</p>",
      "rawMarkdown": "Hello everyone, I m trying to remove the overlap of patients by checking to see if a patient's ID appears in both the training set and the test set and should also verify that I don't have patient overlap in the training and validation sets. In this data set, we have 11913 unique ID. So I want to know how can I avoid this overlapping and split my data in train, test, and validation.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2087378,
      "author_name": "Pandas Warrior",
      "author_url": "",
      "post_date": "2023-01-05T15:01:02.690000",
      "content": "<p>You can check Stratified Group K Fold (<a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedGroupKFold.html)\" target=\"_blank\">https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedGroupKFold.html)</a>. This will allow you to group the data by patients and divide them into K - cross-validation subsets. Thanks to this procedure, in any fold the patients will not be repeated.</p>\n<p>In addition, if you train K - models, each validated on a different subgroup, and trained on the rest of the data, you can then select the best models and try to combine their results, this way you benefit from all the available data.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2090388,
          "author_name": "noura bentaher",
          "author_url": "",
          "post_date": "2023-01-07T09:20:30.603000",
          "content": "<p>Thank you for this information!! good luck</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2127554,
          "author_name": "Antti Isosalo",
          "author_url": "",
          "post_date": "2023-02-03T01:02:59.173000",
          "content": "<p><a href=\"https://www.kaggle.com/dawidstachowiak\" target=\"_blank\">@dawidstachowiak</a> Do I remember correctly that this accepts also the <code>y</code> parameter in addition to <code>group</code>? Do you mind my asking what metadata column you have used there? That could be the <strong>density</strong> column (combined with <strong>cancer</strong> column or the <strong>birads</strong> column)? Or is it better to train a separate model for each density?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2087378": "You can check Stratified Group K Fold (https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedGroupKFold.html). This will allow you to group the data by patients and divide them into K - cross-validation subsets. Thanks to this procedure, in any fold the patients will not be repeated.\n\nIn addition, if you train K - models, each validated on a different subgroup, and trained on the rest of the data, you can then select the best models and try to combine their results, this way you benefit from all the available data.",
    "2087163": "Hello everyone, I m trying to remove the overlap of patients by checking to see if a patient's ID appears in both the training set and the test set and should also verify that I don't have patient overlap in the training and validation sets. In this data set, we have 11913 unique ID. So I want to know how can I avoid this overlapping and split my data in train, test, and validation."
  }
}