{
  "id": 374203,
  "title": "CV Strategy",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/374203",
  "author_name": "happymentee",
  "post_date": "2022-12-26T01:34:44.140000",
  "votes": 0,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi! I'm a complete beginner at the Kaggle competition, and I'm confused about choosing the right CV strategy:<br>\n1, Should I use <code>StratifiedKFold</code> or <code>StratifiedGroupKFold</code> (in this competition )<br>\nI noticed many notebooks (such as <a href=\"https://www.kaggle.com/code/snnclsr/rsna-pytorch-baseline-training\" target=\"_blank\">1</a>, <a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train/notebook\" target=\"_blank\">2</a>, … using this:</p>\n<p><code>gkfold = StratifiedGroupKFold(n_splits=CFG.n_folds)\nfor fold_idx, (train_idx, val_idx) in enumerate(gkfold.split(df_all, y=df_all.cancer, groups=df_all.patient_id))</code></p>\n<p>However, when reading two solutions from <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372228#2065890\" target=\"_blank\">Remek Kinas</a> and <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371185\" target=\"_blank\">MARTIN KOVACEVIC BUVINIC</a>, I saw them using CV per <code>patient_id</code> and <code>laterality</code>. <strong>If they use this CV, whether <code>stratifiedKFold</code> or <code>StratifiedGroupKFold</code> been used?</strong></p>\n<p>**2, What features should be used to split fold (in general and in this comp)? **</p>",
  "messages": [
    {
      "id": 2076402,
      "postDate": "2022-12-26T13:20:19.343Z",
      "content": "<p>I use this strategy:</p>\n<pre><code>ds = pd.read_csv(\"train.csv\")\nds['split'] = ds['patient_id'].astype(str)+ '_' + ds['laterality'].astype(str)\n\nfolds = train.copy()\n\nif CFG.fold_split:\n    Fold = StratifiedGroupKFold(n_splits=CFG.n_fold)\n    for n, (train_index, val_index) in enumerate(Fold.split(folds, folds[CFG.target_col], groups=folds['split'])):\n        folds.loc[val_index, 'fold'] = int(n)\n    folds['fold'] = folds['fold'].astype(int)\n\nprint(folds.groupby(['fold', CFG.target_col]).size())\n</code></pre>",
      "rawMarkdown": "I use this strategy:\n\n```\nds = pd.read_csv(\"train.csv\")\nds['split'] = ds['patient_id'].astype(str)+ '_' + ds['laterality'].astype(str)\n\nfolds = train.copy()\n\nif CFG.fold_split:\n    Fold = StratifiedGroupKFold(n_splits=CFG.n_fold)\n    for n, (train_index, val_index) in enumerate(Fold.split(folds, folds[CFG.target_col], groups=folds['split'])):\n        folds.loc[val_index, 'fold'] = int(n)\n    folds['fold'] = folds['fold'].astype(int)\n    \nprint(folds.groupby(['fold', CFG.target_col]).size())\n```",
      "votes": 1,
      "replies": [
        {
          "id": 2077392,
          "postDate": "2022-12-27T14:35:09.503Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        },
        {
          "id": 2080706,
          "postDate": "2022-12-30T12:06:16.087Z",
          "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> I think you should not create your groups based on your 'split' column. You'll get different groups for a same patient. Even though it's probably not a big deal as right and left breasts are not exactly the same and probably do not contain both cancer, you still have a potentially leaky situation here. In order to mimic train vs public vs private (and real life setting), I think it's better to simply group by patient id so that you have a set of patients for your training set and a completely different for your validation set.</p>",
          "rawMarkdown": "@remekkinas I think you should not create your groups based on your 'split' column. You'll get different groups for a same patient. Even though it's probably not a big deal as right and left breasts are not exactly the same and probably do not contain both cancer, you still have a potentially leaky situation here. In order to mimic train vs public vs private (and real life setting), I think it's better to simply group by patient id so that you have a set of patients for your training set and a completely different for your validation set.",
          "votes": 3,
          "replies": [
            {
              "id": 2080792,
              "postDate": "2022-12-30T13:41:37.743Z",
              "content": "<p>Yes, you are right - safer is grouping by patient_id. I agree. Thank you for your comment.</p>",
              "rawMarkdown": "Yes, you are right - safer is grouping by patient_id. I agree. Thank you for your comment.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2075894,
      "postDate": "2022-12-26T01:34:44.140Z",
      "content": "<p>Hi! I'm a complete beginner at the Kaggle competition, and I'm confused about choosing the right CV strategy:<br>\n1, Should I use <code>StratifiedKFold</code> or <code>StratifiedGroupKFold</code> (in this competition )<br>\nI noticed many notebooks (such as <a href=\"https://www.kaggle.com/code/snnclsr/rsna-pytorch-baseline-training\" target=\"_blank\">1</a>, <a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train/notebook\" target=\"_blank\">2</a>, … using this:</p>\n<p><code>gkfold = StratifiedGroupKFold(n_splits=CFG.n_folds)\nfor fold_idx, (train_idx, val_idx) in enumerate(gkfold.split(df_all, y=df_all.cancer, groups=df_all.patient_id))</code></p>\n<p>However, when reading two solutions from <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372228#2065890\" target=\"_blank\">Remek Kinas</a> and <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371185\" target=\"_blank\">MARTIN KOVACEVIC BUVINIC</a>, I saw them using CV per <code>patient_id</code> and <code>laterality</code>. <strong>If they use this CV, whether <code>stratifiedKFold</code> or <code>StratifiedGroupKFold</code> been used?</strong></p>\n<p>**2, What features should be used to split fold (in general and in this comp)? **</p>",
      "rawMarkdown": "Hi! I'm a complete beginner at the Kaggle competition, and I'm confused about choosing the right CV strategy:\n1, Should I use `StratifiedKFold` or `StratifiedGroupKFold` (in this competition )\nI noticed many notebooks (such as [1](https://www.kaggle.com/code/snnclsr/rsna-pytorch-baseline-training), [2](https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train/notebook ), ... using this:\n\n`gkfold = StratifiedGroupKFold(n_splits=CFG.n_folds)\nfor fold_idx, (train_idx, val_idx) in enumerate(gkfold.split(df_all, y=df_all.cancer, groups=df_all.patient_id))`\n\nHowever, when reading two solutions from [Remek Kinas](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372228#2065890) and [MARTIN KOVACEVIC BUVINIC](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371185), I saw them using CV per `patient_id` and `laterality`. **If they use this CV, whether `stratifiedKFold` or `StratifiedGroupKFold` been used?**\n\n**2, What features should be used to split fold (in general and in this comp)? **"
    }
  ],
  "comments": [
    {
      "id": 2076402,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-12-26T13:20:19.343000",
      "content": "<p>I use this strategy:</p>\n<pre><code>ds = pd.read_csv(\"train.csv\")\nds['split'] = ds['patient_id'].astype(str)+ '_' + ds['laterality'].astype(str)\n\nfolds = train.copy()\n\nif CFG.fold_split:\n    Fold = StratifiedGroupKFold(n_splits=CFG.n_fold)\n    for n, (train_index, val_index) in enumerate(Fold.split(folds, folds[CFG.target_col], groups=folds['split'])):\n        folds.loc[val_index, 'fold'] = int(n)\n    folds['fold'] = folds['fold'].astype(int)\n\nprint(folds.groupby(['fold', CFG.target_col]).size())\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 2077392,
          "author_name": "happymentee",
          "author_url": "",
          "post_date": "2022-12-27T14:35:09.503000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2080706,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2022-12-30T12:06:16.087000",
          "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> I think you should not create your groups based on your 'split' column. You'll get different groups for a same patient. Even though it's probably not a big deal as right and left breasts are not exactly the same and probably do not contain both cancer, you still have a potentially leaky situation here. In order to mimic train vs public vs private (and real life setting), I think it's better to simply group by patient id so that you have a set of patients for your training set and a completely different for your validation set.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2080792,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2022-12-30T13:41:37.743000",
              "content": "<p>Yes, you are right - safer is grouping by patient_id. I agree. Thank you for your comment.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2076402": "I use this strategy:\n\n```\nds = pd.read_csv(\"train.csv\")\nds['split'] = ds['patient_id'].astype(str)+ '_' + ds['laterality'].astype(str)\n\nfolds = train.copy()\n\nif CFG.fold_split:\n    Fold = StratifiedGroupKFold(n_splits=CFG.n_fold)\n    for n, (train_index, val_index) in enumerate(Fold.split(folds, folds[CFG.target_col], groups=folds['split'])):\n        folds.loc[val_index, 'fold'] = int(n)\n    folds['fold'] = folds['fold'].astype(int)\n    \nprint(folds.groupby(['fold', CFG.target_col]).size())\n```",
    "2075894": "Hi! I'm a complete beginner at the Kaggle competition, and I'm confused about choosing the right CV strategy:\n1, Should I use `StratifiedKFold` or `StratifiedGroupKFold` (in this competition )\nI noticed many notebooks (such as [1](https://www.kaggle.com/code/snnclsr/rsna-pytorch-baseline-training), [2](https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train/notebook ), ... using this:\n\n`gkfold = StratifiedGroupKFold(n_splits=CFG.n_folds)\nfor fold_idx, (train_idx, val_idx) in enumerate(gkfold.split(df_all, y=df_all.cancer, groups=df_all.patient_id))`\n\nHowever, when reading two solutions from [Remek Kinas](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/372228#2065890) and [MARTIN KOVACEVIC BUVINIC](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371185), I saw them using CV per `patient_id` and `laterality`. **If they use this CV, whether `stratifiedKFold` or `StratifiedGroupKFold` been used?**\n\n**2, What features should be used to split fold (in general and in this comp)? **"
  }
}