{
  "id": 112426,
  "title": "Effect of imbalanced data on your metrics",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/112426",
  "author_name": "kambarakun",
  "post_date": "2019-10-12T19:29:23.935000",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Depending on the number of positive samples, not only accuracy but also logloss is affected.  </p>\n\n<p>I tried MCC (Matthews correlation coefficient) as a measure to correct apparent logloss deterioration when down-sampling. <br>\nMy MCC function is currently under verification. I would appreciate any advice on the code:)</p>\n\n<p>In the end, it's probably best to do GroupKFold validation with PatientID. Is there any other good idea?</p>\n\n<p>```</p>\n\n<h1>ypred and ytrue are fastai(equal PyTorch?) Tensor objects</h1>\n\n<p>def mcc(y_pred, y_true, thresh=0.5, eps=1e-15):\n    y_pred = (y_pred &gt; thresh).float()\n    y_true = y_true.float()\n    TP     = (y_pred     * y_true    ).sum(dim=0)\n    FP     = (y_pred     * (1-y_true)).sum(dim=0)\n    TN     = ((1-y_pred) * (1-y_true)).sum(dim=0)\n    FN     = ((1-y_pred) * y_true    ).sum(dim=0)\n    MCC    = (TP * TN - FP * FN) / ((TP + FP + eps) * (TP + FN + eps) * (TN + FP + eps) * (TN + FN + eps)).sqrt()\n    return MCC.mean()\n```</p>",
  "messages": [
    {
      "id": 649019,
      "postDate": "2019-10-14T21:34:30.927Z",
      "content": "<p>May'be some usefull tips. Use a large enough validation set. Use a test size of 0.20 or even 0.25. I started out with 0.15...but a little larger seems to give a better 'predictable' loss metric.</p>\n\n<p>Also use a good multi-class stratification module todo a good split. A library I use for this is: <a href=\"https://github.com/trent-b/iterative-stratification\">url</a>\nBut there are offcourse multiple good libraries that you can use.</p>\n\n<p>I have no experience with the MCC code you mention. For me the regular binary crossentropy works fine (at least until now..) even with modest up or downsampling.</p>\n\n<p>Good luck!</p>",
      "rawMarkdown": "May'be some usefull tips. Use a large enough validation set. Use a test size of 0.20 or even 0.25. I started out with 0.15...but a little larger seems to give a better 'predictable' loss metric.\n\nAlso use a good multi-class stratification module todo a good split. A library I use for this is: [url](https://github.com/trent-b/iterative-stratification)\nBut there are offcourse multiple good libraries that you can use.\n\nI have no experience with the MCC code you mention. For me the regular binary crossentropy works fine (at least until now..) even with modest up or downsampling.\n\nGood luck!",
      "votes": 3,
      "replies": [
        {
          "id": 649031,
          "postDate": "2019-10-14T21:56:18.330Z",
          "content": "<p>Hey Robin, I saw your link mentioned for the StratifiedKFold Multi Label, but how do you instantiate X and y without using too much memory? Sorry for my question, I am beginner</p>",
          "rawMarkdown": "Hey Robin, I saw your link mentioned for the StratifiedKFold Multi Label, but how do you instantiate X and y without using too much memory? Sorry for my question, I am beginner\n",
          "votes": 1
        },
        {
          "id": 649406,
          "postDate": "2019-10-15T09:58:31.917Z",
          "content": "<p>Hi Gabriel,</p>\n\n<p>No problem. So the train.csv is only about 100 Megabyte so doing the splitting into train and validation sets should not be a real issue for the Kaggle kernel with I believe 13GB's.</p>\n\n<p>See the code below for some example to get you started.\n```</p>\n\n<h1>Multi Label Stratified Split stuff...</h1>\n\n<p>msss = MultilabelStratifiedShuffleSplit(n_splits = 20, test_size = TEST_SIZE, random_state = SEED)\nX = train_df.index\nY = train_df.Label.values\nprint(X)\nprint(Y)    </p>\n\n<p>```\nYou first setup the MultiLabelStratifiedShuffleSplit...define a number of splits, test size and random state. \nNext you can can split. Note that msss is iterable so you can loop over your n_splits. Use the train and valid index again with your dataframe to get the correct split and pass it on to a data generator or something comparable.</p>\n\n<p>```</p>\n\n<h1>Get train and test index, shuffle train indexes.</h1>\n\n<p>msss_splits = next(msss.split(X, Y))\ntrain_idx = msss_splits[0]\nvalid_idx = msss_splits[1]\nprint(train_idx[:5]) <br>\nprint(valid_idx[:5])\n```</p>",
          "rawMarkdown": "Hi Gabriel,\n\nNo problem. So the train.csv is only about 100 Megabyte so doing the splitting into train and validation sets should not be a real issue for the Kaggle kernel with I believe 13GB's.\n\nSee the code below for some example to get you started.\n```\n# Multi Label Stratified Split stuff...\nmsss = MultilabelStratifiedShuffleSplit(n_splits = 20, test_size = TEST_SIZE, random_state = SEED)\nX = train_df.index\nY = train_df.Label.values\nprint(X)\nprint(Y)    \n\n```\nYou first setup the MultiLabelStratifiedShuffleSplit...define a number of splits, test size and random state. \nNext you can can split. Note that msss is iterable so you can loop over your n_splits. Use the train and valid index again with your dataframe to get the correct split and pass it on to a data generator or something comparable.\n\n```\n# Get train and test index, shuffle train indexes.\nmsss_splits = next(msss.split(X, Y))\ntrain_idx = msss_splits[0]\nvalid_idx = msss_splits[1]\nprint(train_idx[:5])    \nprint(valid_idx[:5])\n```",
          "votes": 1
        },
        {
          "id": 649419,
          "postDate": "2019-10-15T10:26:01.243Z",
          "content": "<p>Also, X is not actually required to generate the split indicies. You can just np.zeros(len(Y)) in place of X</p>",
          "rawMarkdown": "Also, X is not actually required to generate the split indicies. You can just np.zeros(len(Y)) in place of X",
          "votes": 1
        },
        {
          "id": 651220,
          "postDate": "2019-10-17T07:32:45.173Z",
          "content": "<p>I agree that stratification is another important thing to look out for. Thanks for the link to the library! I will try it out later but it seems a bit more elaborate than my current solution. So far I simply created a vector of all the different hemorrhage combinations / permutations in combination with scikit-learns stratification optins.</p>",
          "rawMarkdown": "I agree that stratification is another important thing to look out for. Thanks for the link to the library! I will try it out later but it seems a bit more elaborate than my current solution. So far I simply created a vector of all the different hemorrhage combinations / permutations in combination with scikit-learns stratification optins."
        }
      ]
    },
    {
      "id": 647530,
      "postDate": "2019-10-12T19:29:23.937Z",
      "content": "<p>Depending on the number of positive samples, not only accuracy but also logloss is affected.  </p>\n\n<p>I tried MCC (Matthews correlation coefficient) as a measure to correct apparent logloss deterioration when down-sampling. <br>\nMy MCC function is currently under verification. I would appreciate any advice on the code:)</p>\n\n<p>In the end, it's probably best to do GroupKFold validation with PatientID. Is there any other good idea?</p>\n\n<p>```</p>\n\n<h1>ypred and ytrue are fastai(equal PyTorch?) Tensor objects</h1>\n\n<p>def mcc(y_pred, y_true, thresh=0.5, eps=1e-15):\n    y_pred = (y_pred &gt; thresh).float()\n    y_true = y_true.float()\n    TP     = (y_pred     * y_true    ).sum(dim=0)\n    FP     = (y_pred     * (1-y_true)).sum(dim=0)\n    TN     = ((1-y_pred) * (1-y_true)).sum(dim=0)\n    FN     = ((1-y_pred) * y_true    ).sum(dim=0)\n    MCC    = (TP * TN - FP * FN) / ((TP + FP + eps) * (TP + FN + eps) * (TN + FP + eps) * (TN + FN + eps)).sqrt()\n    return MCC.mean()\n```</p>",
      "rawMarkdown": "Depending on the number of positive samples, not only accuracy but also logloss is affected.  \n\nI tried MCC (Matthews correlation coefficient) as a measure to correct apparent logloss deterioration when down-sampling.  \nMy MCC function is currently under verification. I would appreciate any advice on the code:)\n\nIn the end, it's probably best to do GroupKFold validation with PatientID. Is there any other good idea?\n\n```\n# ypred and ytrue are fastai(equal PyTorch?) Tensor objects\n\ndef mcc(y_pred, y_true, thresh=0.5, eps=1e-15):\n    y_pred = (y_pred &gt; thresh).float()\n    y_true = y_true.float()\n    TP     = (y_pred     * y_true    ).sum(dim=0)\n    FP     = (y_pred     * (1-y_true)).sum(dim=0)\n    TN     = ((1-y_pred) * (1-y_true)).sum(dim=0)\n    FN     = ((1-y_pred) * y_true    ).sum(dim=0)\n    MCC    = (TP * TN - FP * FN) / ((TP + FP + eps) * (TP + FN + eps) * (TN + FP + eps) * (TN + FN + eps)).sqrt()\n    return MCC.mean()\n```",
      "votes": 3
    },
    {
      "id": 651216,
      "postDate": "2019-10-17T07:25:00.087Z",
      "content": "<p>Have you also had a look at Balanced Accuracy in the context? I find it to be one of the easiest to understand metrics, but I am not sure if it's necessarily a good one. \nBAAC would be <code>((TP / P) + (TN / N) ) / 2</code> , only majority voting for example would lead in an BAAC of 0.5. In multi-label contexts, it might get a bit weird though. </p>",
      "rawMarkdown": "Have you also had a look at Balanced Accuracy in the context? I find it to be one of the easiest to understand metrics, but I am not sure if it's necessarily a good one. \nBAAC would be `((TP / P) + (TN / N) ) / 2` , only majority voting for example would lead in an BAAC of 0.5. In multi-label contexts, it might get a bit weird though. "
    }
  ],
  "comments": [
    {
      "id": 649019,
      "author_name": "Robin Smits",
      "author_url": "",
      "post_date": "2019-10-14T21:34:30.927000",
      "content": "<p>May'be some usefull tips. Use a large enough validation set. Use a test size of 0.20 or even 0.25. I started out with 0.15...but a little larger seems to give a better 'predictable' loss metric.</p>\n\n<p>Also use a good multi-class stratification module todo a good split. A library I use for this is: <a href=\"https://github.com/trent-b/iterative-stratification\">url</a>\nBut there are offcourse multiple good libraries that you can use.</p>\n\n<p>I have no experience with the MCC code you mention. For me the regular binary crossentropy works fine (at least until now..) even with modest up or downsampling.</p>\n\n<p>Good luck!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 649031,
          "author_name": "Gabriel",
          "author_url": "",
          "post_date": "2019-10-14T21:56:18.330000",
          "content": "<p>Hey Robin, I saw your link mentioned for the StratifiedKFold Multi Label, but how do you instantiate X and y without using too much memory? Sorry for my question, I am beginner</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 649406,
          "author_name": "Robin Smits",
          "author_url": "",
          "post_date": "2019-10-15T09:58:31.917000",
          "content": "<p>Hi Gabriel,</p>\n\n<p>No problem. So the train.csv is only about 100 Megabyte so doing the splitting into train and validation sets should not be a real issue for the Kaggle kernel with I believe 13GB's.</p>\n\n<p>See the code below for some example to get you started.\n```</p>\n\n<h1>Multi Label Stratified Split stuff...</h1>\n\n<p>msss = MultilabelStratifiedShuffleSplit(n_splits = 20, test_size = TEST_SIZE, random_state = SEED)\nX = train_df.index\nY = train_df.Label.values\nprint(X)\nprint(Y)    </p>\n\n<p>```\nYou first setup the MultiLabelStratifiedShuffleSplit...define a number of splits, test size and random state. \nNext you can can split. Note that msss is iterable so you can loop over your n_splits. Use the train and valid index again with your dataframe to get the correct split and pass it on to a data generator or something comparable.</p>\n\n<p>```</p>\n\n<h1>Get train and test index, shuffle train indexes.</h1>\n\n<p>msss_splits = next(msss.split(X, Y))\ntrain_idx = msss_splits[0]\nvalid_idx = msss_splits[1]\nprint(train_idx[:5]) <br>\nprint(valid_idx[:5])\n```</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 649419,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "2019-10-15T10:26:01.243000",
          "content": "<p>Also, X is not actually required to generate the split indicies. You can just np.zeros(len(Y)) in place of X</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 651220,
          "author_name": "srs",
          "author_url": "",
          "post_date": "2019-10-17T07:32:45.173000",
          "content": "<p>I agree that stratification is another important thing to look out for. Thanks for the link to the library! I will try it out later but it seems a bit more elaborate than my current solution. So far I simply created a vector of all the different hemorrhage combinations / permutations in combination with scikit-learns stratification optins.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 651216,
      "author_name": "srs",
      "author_url": "",
      "post_date": "2019-10-17T07:25:00.087000",
      "content": "<p>Have you also had a look at Balanced Accuracy in the context? I find it to be one of the easiest to understand metrics, but I am not sure if it's necessarily a good one. \nBAAC would be <code>((TP / P) + (TN / N) ) / 2</code> , only majority voting for example would lead in an BAAC of 0.5. In multi-label contexts, it might get a bit weird though. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "649019": "May'be some usefull tips. Use a large enough validation set. Use a test size of 0.20 or even 0.25. I started out with 0.15...but a little larger seems to give a better 'predictable' loss metric.\n\nAlso use a good multi-class stratification module todo a good split. A library I use for this is: [url](https://github.com/trent-b/iterative-stratification)\nBut there are offcourse multiple good libraries that you can use.\n\nI have no experience with the MCC code you mention. For me the regular binary crossentropy works fine (at least until now..) even with modest up or downsampling.\n\nGood luck!",
    "647530": "Depending on the number of positive samples, not only accuracy but also logloss is affected.  \n\nI tried MCC (Matthews correlation coefficient) as a measure to correct apparent logloss deterioration when down-sampling.  \nMy MCC function is currently under verification. I would appreciate any advice on the code:)\n\nIn the end, it's probably best to do GroupKFold validation with PatientID. Is there any other good idea?\n\n```\n# ypred and ytrue are fastai(equal PyTorch?) Tensor objects\n\ndef mcc(y_pred, y_true, thresh=0.5, eps=1e-15):\n    y_pred = (y_pred &gt; thresh).float()\n    y_true = y_true.float()\n    TP     = (y_pred     * y_true    ).sum(dim=0)\n    FP     = (y_pred     * (1-y_true)).sum(dim=0)\n    TN     = ((1-y_pred) * (1-y_true)).sum(dim=0)\n    FN     = ((1-y_pred) * y_true    ).sum(dim=0)\n    MCC    = (TP * TN - FP * FN) / ((TP + FP + eps) * (TP + FN + eps) * (TN + FP + eps) * (TN + FN + eps)).sqrt()\n    return MCC.mean()\n```",
    "651216": "Have you also had a look at Balanced Accuracy in the context? I find it to be one of the easiest to understand metrics, but I am not sure if it's necessarily a good one. \nBAAC would be `((TP / P) + (TN / N) ) / 2` , only majority voting for example would lead in an BAAC of 0.5. In multi-label contexts, it might get a bit weird though. "
  }
}