{
  "id": 110578,
  "title": "Question on Stratified Split",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/110578",
  "author_name": "Ken Ho",
  "post_date": "2019-09-29T11:16:06.299000",
  "votes": 0,
  "comment_count": 6,
  "views": 0,
  "content": "<p>This is a multi-label problems. Therefore, I cannot simply pick 'any' for stratification. For now, I do not have any idea on this problem.</p>\n\n<p>What splitting method do you use?</p>",
  "messages": [
    {
      "id": 636397,
      "postDate": "2019-09-29T13:06:57.257Z",
      "content": "<p>I'm going to start with <code>GroupKFold</code> on the patient ID to prevent the same patient leaking between the training and validation sets. I haven't actually started yet, so I can't vouch for how well this works 😄 </p>",
      "rawMarkdown": "I'm going to start with `GroupKFold` on the patient ID to prevent the same patient leaking between the training and validation sets. I haven't actually started yet, so I can't vouch for how well this works 😄 ",
      "votes": 1,
      "replies": [
        {
          "id": 636440,
          "postDate": "2019-09-29T14:38:40.563Z",
          "content": "<p>Remember labelling is done at the slice level, not at the patient level, so any \"leakage\" you'd get would only occur if the experts were reviewing slices sequentially and \"carrying over\" knowledge between slices, which tbf could be the case.</p>",
          "rawMarkdown": "Remember labelling is done at the slice level, not at the patient level, so any \"leakage\" you'd get would only occur if the experts were reviewing slices sequentially and \"carrying over\" knowledge between slices, which tbf could be the case.",
          "votes": 4
        },
        {
          "id": 636491,
          "postDate": "2019-09-29T16:42:00.430Z",
          "content": "<p>There are also repeat scans that could cause train/val leakage too</p>",
          "rawMarkdown": "There are also repeat scans that could cause train/val leakage too",
          "votes": 1
        },
        {
          "id": 636495,
          "postDate": "2019-09-29T16:57:28.390Z",
          "content": "<p>Good point!</p>",
          "rawMarkdown": "Good point!",
          "votes": 1
        }
      ]
    },
    {
      "id": 636444,
      "postDate": "2019-09-29T14:46:18.433Z",
      "content": "<p>If you aren't going to do 3D reconstructions, you can use simple split, and because the data set is quit large you don't need to do few folds, you can just do a 1:10 split and be OK.\nIf you are intending to do some 3D post-processing, you don't want images from the same series to be on test and validate. </p>",
      "rawMarkdown": "If you aren't going to do 3D reconstructions, you can use simple split, and because the data set is quit large you don't need to do few folds, you can just do a 1:10 split and be OK.\nIf you are intending to do some 3D post-processing, you don't want images from the same series to be on test and validate. ",
      "votes": 2
    },
    {
      "id": 636398,
      "postDate": "2019-09-29T13:07:53.013Z",
      "content": "<p>Currently just doing random split so far i am getting close correlation between cv and lb..</p>",
      "rawMarkdown": "Currently just doing random split so far i am getting close correlation between cv and lb..",
      "votes": 2
    },
    {
      "id": 636361,
      "postDate": "2019-09-29T11:16:06.300Z",
      "content": "<p>This is a multi-label problems. Therefore, I cannot simply pick 'any' for stratification. For now, I do not have any idea on this problem.</p>\n\n<p>What splitting method do you use?</p>",
      "rawMarkdown": "This is a multi-label problems. Therefore, I cannot simply pick 'any' for stratification. For now, I do not have any idea on this problem.\n\nWhat splitting method do you use?"
    }
  ],
  "comments": [
    {
      "id": 636397,
      "author_name": "datasaurus",
      "author_url": "",
      "post_date": "2019-09-29T13:06:57.257000",
      "content": "<p>I'm going to start with <code>GroupKFold</code> on the patient ID to prevent the same patient leaking between the training and validation sets. I haven't actually started yet, so I can't vouch for how well this works 😄 </p>",
      "votes": 1,
      "replies": [
        {
          "id": 636440,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-09-29T14:38:40.563000",
          "content": "<p>Remember labelling is done at the slice level, not at the patient level, so any \"leakage\" you'd get would only occur if the experts were reviewing slices sequentially and \"carrying over\" knowledge between slices, which tbf could be the case.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 636491,
          "author_name": "datasaurus",
          "author_url": "",
          "post_date": "2019-09-29T16:42:00.430000",
          "content": "<p>There are also repeat scans that could cause train/val leakage too</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 636495,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-09-29T16:57:28.390000",
          "content": "<p>Good point!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 636444,
      "author_name": "yuval reina",
      "author_url": "",
      "post_date": "2019-09-29T14:46:18.433000",
      "content": "<p>If you aren't going to do 3D reconstructions, you can use simple split, and because the data set is quit large you don't need to do few folds, you can just do a 1:10 split and be OK.\nIf you are intending to do some 3D post-processing, you don't want images from the same series to be on test and validate. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 636398,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2019-09-29T13:07:53.013000",
      "content": "<p>Currently just doing random split so far i am getting close correlation between cv and lb..</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "636397": "I'm going to start with `GroupKFold` on the patient ID to prevent the same patient leaking between the training and validation sets. I haven't actually started yet, so I can't vouch for how well this works 😄 ",
    "636444": "If you aren't going to do 3D reconstructions, you can use simple split, and because the data set is quit large you don't need to do few folds, you can just do a 1:10 split and be OK.\nIf you are intending to do some 3D post-processing, you don't want images from the same series to be on test and validate. ",
    "636398": "Currently just doing random split so far i am getting close correlation between cv and lb..",
    "636361": "This is a multi-label problems. Therefore, I cannot simply pick 'any' for stratification. For now, I do not have any idea on this problem.\n\nWhat splitting method do you use?"
  }
}