{
  "id": 147958,
  "title": "Weighted Train Df Sampling Based on ISUP/Gleason",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/147958",
  "author_name": "Zac Dannelly",
  "post_date": "2020-05-02T14:54:13.721000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p><a href=\"https://stackoverflow.com/a/41528532\">Found this recipe</a> quite useful in testing and wanted to share. It has the ability to sample N amount from a dataframe with equal weighting based on a categorical/target column.</p>\n\n<p>Also have added some code to make it a percent if easier in your pipeline:</p>\n\n<p>```\ntrain_data_df = pd.read_csv(\"../input/prostate-cancer-grade-assessment/train.csv\")\nN_to_process = int(len(train_data_df) * pct_to_process)\nsample_df = categorical_sample(train_data_df, \"isup_grade\", N_to_process)</p>\n\n<p>def categorical_sample(df, cat_col, N):\n     cat_group = df.groupby(cat_col, group_keys=False)\n     sampled_df = cat_group.apply(lambda g: g.sample(int(N * len(g)/len(df))))\n     return sampled_df\n```</p>\n\n<p>Hope this is of some assistance with testing! </p>",
  "messages": [
    {
      "id": 830380,
      "postDate": "2020-05-02T14:54:13.723Z",
      "content": "<p><a href=\"https://stackoverflow.com/a/41528532\">Found this recipe</a> quite useful in testing and wanted to share. It has the ability to sample N amount from a dataframe with equal weighting based on a categorical/target column.</p>\n\n<p>Also have added some code to make it a percent if easier in your pipeline:</p>\n\n<p>```\ntrain_data_df = pd.read_csv(\"../input/prostate-cancer-grade-assessment/train.csv\")\nN_to_process = int(len(train_data_df) * pct_to_process)\nsample_df = categorical_sample(train_data_df, \"isup_grade\", N_to_process)</p>\n\n<p>def categorical_sample(df, cat_col, N):\n     cat_group = df.groupby(cat_col, group_keys=False)\n     sampled_df = cat_group.apply(lambda g: g.sample(int(N * len(g)/len(df))))\n     return sampled_df\n```</p>\n\n<p>Hope this is of some assistance with testing! </p>",
      "rawMarkdown": "[Found this recipe](https://stackoverflow.com/a/41528532) quite useful in testing and wanted to share. It has the ability to sample N amount from a dataframe with equal weighting based on a categorical/target column.\n\nAlso have added some code to make it a percent if easier in your pipeline:\n\n```\ntrain_data_df = pd.read_csv(\"../input/prostate-cancer-grade-assessment/train.csv\")\nN_to_process = int(len(train_data_df) * pct_to_process)\nsample_df = categorical_sample(train_data_df, \"isup_grade\", N_to_process)\n\ndef categorical_sample(df, cat_col, N):\n     cat_group = df.groupby(cat_col, group_keys=False)\n     sampled_df = cat_group.apply(lambda g: g.sample(int(N * len(g)/len(df))))\n     return sampled_df\n```\n\nHope this is of some assistance with testing! \n",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "830380": "[Found this recipe](https://stackoverflow.com/a/41528532) quite useful in testing and wanted to share. It has the ability to sample N amount from a dataframe with equal weighting based on a categorical/target column.\n\nAlso have added some code to make it a percent if easier in your pipeline:\n\n```\ntrain_data_df = pd.read_csv(\"../input/prostate-cancer-grade-assessment/train.csv\")\nN_to_process = int(len(train_data_df) * pct_to_process)\nsample_df = categorical_sample(train_data_df, \"isup_grade\", N_to_process)\n\ndef categorical_sample(df, cat_col, N):\n     cat_group = df.groupby(cat_col, group_keys=False)\n     sampled_df = cat_group.apply(lambda g: g.sample(int(N * len(g)/len(df))))\n     return sampled_df\n```\n\nHope this is of some assistance with testing! \n"
  }
}