{
  "id": 371020,
  "title": "How to Handle Class Imbalance in Computer Vision",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/371020",
  "author_name": "moth",
  "post_date": "2022-12-07T15:17:52.594000",
  "votes": 14,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello guys, I was wondering which techniques beside:</p>\n<ul>\n<li>Undersampling majority class.</li>\n<li>Oversampling minority class.</li>\n<li>Focal Loss.</li>\n<li>Weighted Loss.</li>\n<li>Augmentations</li>\n</ul>\n<p>can be used in Computer Vision to handle class imbalance.</p>\n<p>Any other technique I am missing?</p>",
  "messages": [
    {
      "id": 2058043,
      "postDate": "2022-12-07T15:17:52.593Z",
      "content": "<p>Hello guys, I was wondering which techniques beside:</p>\n<ul>\n<li>Undersampling majority class.</li>\n<li>Oversampling minority class.</li>\n<li>Focal Loss.</li>\n<li>Weighted Loss.</li>\n<li>Augmentations</li>\n</ul>\n<p>can be used in Computer Vision to handle class imbalance.</p>\n<p>Any other technique I am missing?</p>",
      "rawMarkdown": "Hello guys, I was wondering which techniques beside:\n- Undersampling majority class.\n- Oversampling minority class.\n- Focal Loss.\n- Weighted Loss.\n- Augmentations\n\ncan be used in Computer Vision to handle class imbalance.\n\nAny other technique I am missing?",
      "votes": 13
    },
    {
      "id": 2059490,
      "postDate": "2022-12-08T22:45:16.887Z",
      "content": "<p>You were nice to quote a <a href=\"https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/172892\" target=\"_blank\">topic </a>I had on a related topic, thanks for that. In that topic I was advocating that fixing imbalance was not really useful. It was mostly true because the metric at the time was roc-auc, hence it did not depend on prediction calibration.</p>\n<p>Here the metric depends on calibration (or depends on the threshold you use for binary predictions). Fixing imbalance may therefore be more relevant here than it was when I wrote that topic.</p>",
      "rawMarkdown": "You were nice to quote a [topic ](https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/172892)I had on a related topic, thanks for that. In that topic I was advocating that fixing imbalance was not really useful. It was mostly true because the metric at the time was roc-auc, hence it did not depend on prediction calibration.\n\nHere the metric depends on calibration (or depends on the threshold you use for binary predictions). Fixing imbalance may therefore be more relevant here than it was when I wrote that topic.",
      "votes": 6,
      "replies": [
        {
          "id": 2059537,
          "postDate": "2022-12-09T01:58:50.097Z",
          "content": "<p>Thanks for clarifying on the matter.</p>",
          "rawMarkdown": "Thanks for clarifying on the matter."
        }
      ]
    },
    {
      "id": 2059531,
      "postDate": "2022-12-09T01:33:53.350Z",
      "content": "<p>Q: \"How to Handle Class Imbalance in Computer Vision?\" <br>\nA: Do experiments to verify</p>\n<p>suggest start training with at least some m pos samples in a batch size of n.</p>\n<p>in computer visions, the problems are multiple.e.g. if a positive sample is difficult to detect because of \"rare content\", then having balanced \"distribution\" may not help.</p>\n<p>there are a couple of issues in this competition:<br>\n[1] small lesion (maybe only 10x10 pixel in the whole of 1024x1024)<br>\n[2] \"uncommon lesion\" (may be common to human but uncommon in distribution), you cannot find \"similar neighbors\"</p>\n<p>i am not sure what is the problem contribution for now:<br>\npoor performance is due to<br>\ne.g. 20% due to class imbalance<br>\n       10% due to small object<br>\n       10% due to rare objects, etc ….</p>\n<hr>\n<p>another question you should ask is, how to show that the problem is class imbalance?<br>\ne.g. results improves when you use oversampling minority class<br>\ne.g. does failure of focal loss indicates that it is not a class imbalance problem?</p>",
      "rawMarkdown": "Q: \"How to Handle Class Imbalance in Computer Vision?\" \nA: Do experiments to verify\n\n suggest start training with at least some m pos samples in a batch size of n.\n\nin computer visions, the problems are multiple.e.g. if a positive sample is difficult to detect because of \"rare content\", then having balanced \"distribution\" may not help.\n\nthere are a couple of issues in this competition:\n[1] small lesion (maybe only 10x10 pixel in the whole of 1024x1024)\n[2] \"uncommon lesion\" (may be common to human but uncommon in distribution), you cannot find \"similar neighbors\"\n\n\ni am not sure what is the problem contribution for now:\npoor performance is due to\ne.g. 20% due to class imbalance\n       10% due to small object\n       10% due to rare objects, etc ....\n\n---\n\nanother question you should ask is, how to show that the problem is class imbalance?\ne.g. results improves when you use oversampling minority class\ne.g. does failure of focal loss indicates that it is not a class imbalance problem?\n",
      "votes": 3,
      "replies": [
        {
          "id": 2059541,
          "postDate": "2022-12-09T02:07:13.760Z",
          "content": "<p>I agree that one should always try and experiment if an idea boosts the performance but if we know in advance whether that boost is significant or not can save us a lot of time.</p>\n<p>I tried Focal Loss, it completely changed the model's prediction distribution from an \"exponential\" one to a more \"right-skewed\". It did not improve my performance though.</p>\n<p>It appears so far that increasing resolution is very beneficial. But with Kaggle's GPU it's barely impossible to train with 1024.</p>",
          "rawMarkdown": "I agree that one should always try and experiment if an idea boosts the performance but if we know in advance whether that boost is significant or not can save us a lot of time.\n\nI tried Focal Loss, it completely changed the model's prediction distribution from an \"exponential\" one to a more \"right-skewed\". It did not improve my performance though.\n\nIt appears so far that increasing resolution is very beneficial. But with Kaggle's GPU it's barely impossible to train with 1024."
        }
      ]
    },
    {
      "id": 2058762,
      "postDate": "2022-12-08T07:31:14.260Z",
      "content": "<p>It is also pretty cool to see the inference notebooks shared publicly either explicitly using multi task learning with auxiliary loss or having everything in place to use one 🙂</p>\n<p>Very suspect 😄 (but points to what people might be using)</p>",
      "rawMarkdown": "It is also pretty cool to see the inference notebooks shared publicly either explicitly using multi task learning with auxiliary loss or having everything in place to use one 🙂\n\nVery suspect 😄 (but points to what people might be using)",
      "votes": 1,
      "replies": [
        {
          "id": 2059375,
          "postDate": "2022-12-08T19:23:06.490Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>. I recently read this excellent <a href=\"https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/172892\" target=\"_blank\">topic</a> by <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> with an extensive discussion on the matter.</p>\n<p>TLDR: </p>\n<ul>\n<li>some people suggests doing nothing</li>\n<li>most people agree that oversampling leads to faster model convergence</li>\n<li>some say that what matters is the <strong>amount</strong> of minority samples not the <strong>ratio</strong> between positive/negative.</li>\n</ul>\n<p><strong>Important:</strong> I believe that that competition's metric was AUC-ROC, whereas here we have pF1. </p>\n<p>I would definitely like to hear more grandmasters opinions on this matter.</p>",
          "rawMarkdown": "Thanks @radek1. I recently read this excellent [topic](https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/172892) by @cpmpml with an extensive discussion on the matter.\n\nTLDR: \n- some people suggests doing nothing\n- most people agree that oversampling leads to faster model convergence\n- some say that what matters is the **amount** of minority samples not the **ratio** between positive/negative.\n\n**Important:** I believe that that competition's metric was AUC-ROC, whereas here we have pF1. \n\nI would definitely like to hear more grandmasters opinions on this matter.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2058760,
      "postDate": "2022-12-08T07:29:56.047Z",
      "content": "<ul>\n<li>semi-supervised pretraining</li>\n<li>pretraining on a bigger dataset</li>\n<li>(possibly) training with discriminative learning rates after the above, or parts of the model frozen (we probably don't need to alter the lower stacks of our CNN too much)</li>\n</ul>",
      "rawMarkdown": "* semi-supervised pretraining\n* pretraining on a bigger dataset\n* (possibly) training with discriminative learning rates after the above, or parts of the model frozen (we probably don't need to alter the lower stacks of our CNN too much)"
    }
  ],
  "comments": [
    {
      "id": 2059490,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2022-12-08T22:45:16.887000",
      "content": "<p>You were nice to quote a <a href=\"https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/172892\" target=\"_blank\">topic </a>I had on a related topic, thanks for that. In that topic I was advocating that fixing imbalance was not really useful. It was mostly true because the metric at the time was roc-auc, hence it did not depend on prediction calibration.</p>\n<p>Here the metric depends on calibration (or depends on the threshold you use for binary predictions). Fixing imbalance may therefore be more relevant here than it was when I wrote that topic.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2059537,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2022-12-09T01:58:50.097000",
          "content": "<p>Thanks for clarifying on the matter.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2059531,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-09T01:33:53.350000",
      "content": "<p>Q: \"How to Handle Class Imbalance in Computer Vision?\" <br>\nA: Do experiments to verify</p>\n<p>suggest start training with at least some m pos samples in a batch size of n.</p>\n<p>in computer visions, the problems are multiple.e.g. if a positive sample is difficult to detect because of \"rare content\", then having balanced \"distribution\" may not help.</p>\n<p>there are a couple of issues in this competition:<br>\n[1] small lesion (maybe only 10x10 pixel in the whole of 1024x1024)<br>\n[2] \"uncommon lesion\" (may be common to human but uncommon in distribution), you cannot find \"similar neighbors\"</p>\n<p>i am not sure what is the problem contribution for now:<br>\npoor performance is due to<br>\ne.g. 20% due to class imbalance<br>\n       10% due to small object<br>\n       10% due to rare objects, etc ….</p>\n<hr>\n<p>another question you should ask is, how to show that the problem is class imbalance?<br>\ne.g. results improves when you use oversampling minority class<br>\ne.g. does failure of focal loss indicates that it is not a class imbalance problem?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2059541,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2022-12-09T02:07:13.760000",
          "content": "<p>I agree that one should always try and experiment if an idea boosts the performance but if we know in advance whether that boost is significant or not can save us a lot of time.</p>\n<p>I tried Focal Loss, it completely changed the model's prediction distribution from an \"exponential\" one to a more \"right-skewed\". It did not improve my performance though.</p>\n<p>It appears so far that increasing resolution is very beneficial. But with Kaggle's GPU it's barely impossible to train with 1024.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2058762,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2022-12-08T07:31:14.260000",
      "content": "<p>It is also pretty cool to see the inference notebooks shared publicly either explicitly using multi task learning with auxiliary loss or having everything in place to use one 🙂</p>\n<p>Very suspect 😄 (but points to what people might be using)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2059375,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2022-12-08T19:23:06.490000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>. I recently read this excellent <a href=\"https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/172892\" target=\"_blank\">topic</a> by <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> with an extensive discussion on the matter.</p>\n<p>TLDR: </p>\n<ul>\n<li>some people suggests doing nothing</li>\n<li>most people agree that oversampling leads to faster model convergence</li>\n<li>some say that what matters is the <strong>amount</strong> of minority samples not the <strong>ratio</strong> between positive/negative.</li>\n</ul>\n<p><strong>Important:</strong> I believe that that competition's metric was AUC-ROC, whereas here we have pF1. </p>\n<p>I would definitely like to hear more grandmasters opinions on this matter.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2058760,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2022-12-08T07:29:56.047000",
      "content": "<ul>\n<li>semi-supervised pretraining</li>\n<li>pretraining on a bigger dataset</li>\n<li>(possibly) training with discriminative learning rates after the above, or parts of the model frozen (we probably don't need to alter the lower stacks of our CNN too much)</li>\n</ul>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2058043": "Hello guys, I was wondering which techniques beside:\n- Undersampling majority class.\n- Oversampling minority class.\n- Focal Loss.\n- Weighted Loss.\n- Augmentations\n\ncan be used in Computer Vision to handle class imbalance.\n\nAny other technique I am missing?",
    "2059490": "You were nice to quote a [topic ](https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/172892)I had on a related topic, thanks for that. In that topic I was advocating that fixing imbalance was not really useful. It was mostly true because the metric at the time was roc-auc, hence it did not depend on prediction calibration.\n\nHere the metric depends on calibration (or depends on the threshold you use for binary predictions). Fixing imbalance may therefore be more relevant here than it was when I wrote that topic.",
    "2059531": "Q: \"How to Handle Class Imbalance in Computer Vision?\" \nA: Do experiments to verify\n\n suggest start training with at least some m pos samples in a batch size of n.\n\nin computer visions, the problems are multiple.e.g. if a positive sample is difficult to detect because of \"rare content\", then having balanced \"distribution\" may not help.\n\nthere are a couple of issues in this competition:\n[1] small lesion (maybe only 10x10 pixel in the whole of 1024x1024)\n[2] \"uncommon lesion\" (may be common to human but uncommon in distribution), you cannot find \"similar neighbors\"\n\n\ni am not sure what is the problem contribution for now:\npoor performance is due to\ne.g. 20% due to class imbalance\n       10% due to small object\n       10% due to rare objects, etc ....\n\n---\n\nanother question you should ask is, how to show that the problem is class imbalance?\ne.g. results improves when you use oversampling minority class\ne.g. does failure of focal loss indicates that it is not a class imbalance problem?\n",
    "2058762": "It is also pretty cool to see the inference notebooks shared publicly either explicitly using multi task learning with auxiliary loss or having everything in place to use one 🙂\n\nVery suspect 😄 (but points to what people might be using)",
    "2058760": "* semi-supervised pretraining\n* pretraining on a bigger dataset\n* (possibly) training with discriminative learning rates after the above, or parts of the model frozen (we probably don't need to alter the lower stacks of our CNN too much)"
  }
}