{
  "id": 114627,
  "title": "tricks for kaggle log loss",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/114627",
  "author_name": "hengck23",
  "post_date": "2019-10-28T04:10:42.452000",
  "votes": 20,
  "comment_count": 7,
  "views": 0,
  "content": "<ol>\n<li><p>sampling/weighing is important in training</p>\n\n<ul><li>when the metric is accuracy, we usually use : predict_label = predict_prob &gt; threshold\nwe do not care about predict_prob, but rather thresholded prediction</li></ul></li>\n</ol>\n\n<p>we sometimes use loss weighing (e.g focal loss) or class balancing to speed up training or improve accuracy. but improving accuracy is \"not the same\" as reducing loss. the loss is modified to improve accuracy by artificially changing the train distribution and affecting the decision boundary. </p>\n\n<p><strong>in summary. we care only about the boundary, and care less about the decision value</strong></p>\n\n<p>for this challenge, we care about the actual log loss itself. so it is important to <strong>\"calibrate your probability\"</strong> to the actual one you expect in the test during training. you may have to use random sampling (instead of balanced sampling) at the final fine tuning. likewise you may need to disable any sample weighing.</p>\n\n<hr>\n\n<ol>\n<li>decimal is important</li>\n</ol>\n\n<p>there will be some different between probability expressed in float 32 bit and 64 bit</p>\n\n<hr>\n\n<ol>\n<li>probability shaping</li>\n</ol>\n\n<p>this is is very difficult (and very risky in scoring). e.g. a test sample is predicted to be negative with probability 0.0001. you may round it to zero.  you will gain a little if the truth is really negative, but you will lose big if it is truth is instead positive.</p>\n\n<p>if you can train with \"large margin\" and use some probability calibration trick such that  you are almost 100% sure that there cannot be positive truth below a certain low values, you can zero out the values. </p>",
  "messages": [
    {
      "id": 659671,
      "postDate": "2019-10-28T04:10:42.453Z",
      "content": "<ol>\n<li><p>sampling/weighing is important in training</p>\n\n<ul><li>when the metric is accuracy, we usually use : predict_label = predict_prob &gt; threshold\nwe do not care about predict_prob, but rather thresholded prediction</li></ul></li>\n</ol>\n\n<p>we sometimes use loss weighing (e.g focal loss) or class balancing to speed up training or improve accuracy. but improving accuracy is \"not the same\" as reducing loss. the loss is modified to improve accuracy by artificially changing the train distribution and affecting the decision boundary. </p>\n\n<p><strong>in summary. we care only about the boundary, and care less about the decision value</strong></p>\n\n<p>for this challenge, we care about the actual log loss itself. so it is important to <strong>\"calibrate your probability\"</strong> to the actual one you expect in the test during training. you may have to use random sampling (instead of balanced sampling) at the final fine tuning. likewise you may need to disable any sample weighing.</p>\n\n<hr>\n\n<ol>\n<li>decimal is important</li>\n</ol>\n\n<p>there will be some different between probability expressed in float 32 bit and 64 bit</p>\n\n<hr>\n\n<ol>\n<li>probability shaping</li>\n</ol>\n\n<p>this is is very difficult (and very risky in scoring). e.g. a test sample is predicted to be negative with probability 0.0001. you may round it to zero.  you will gain a little if the truth is really negative, but you will lose big if it is truth is instead positive.</p>\n\n<p>if you can train with \"large margin\" and use some probability calibration trick such that  you are almost 100% sure that there cannot be positive truth below a certain low values, you can zero out the values. </p>",
      "rawMarkdown": "1. sampling/weighing is important in training\n\n- when the metric is accuracy, we usually use : predict\\_label = predict\\_prob &gt; threshold\nwe do not care about predict\\_prob, but rather thresholded prediction\n\nwe sometimes use loss weighing (e.g focal loss) or class balancing to speed up training or improve accuracy. but improving accuracy is \"not the same\" as reducing loss. the loss is modified to improve accuracy by artificially changing the train distribution and affecting the decision boundary. \n\n**in summary. we care only about the boundary, and care less about the decision value**\n\n\nfor this challenge, we care about the actual log loss itself. so it is important to **\"calibrate your probability\"** to the actual one you expect in the test during training. you may have to use random sampling (instead of balanced sampling) at the final fine tuning. likewise you may need to disable any sample weighing.\n\n---\n\n2. decimal is important\n\nthere will be some different between probability expressed in float 32 bit and 64 bit\n\n\n---\n\n3. probability shaping\n\nthis is is very difficult (and very risky in scoring). e.g. a test sample is predicted to be negative with probability 0.0001. you may round it to zero.  you will gain a little if the truth is really negative, but you will lose big if it is truth is instead positive.\n\nif you can train with \"large margin\" and use some probability calibration trick such that  you are almost 100% sure that there cannot be positive truth below a certain low values, you can zero out the values. \n",
      "votes": 20
    },
    {
      "id": 663815,
      "postDate": "2019-11-02T17:41:25.107Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F1e6fb740684b6cdb77eef4f26d56f8e2%2FD62E03A7-9571-4737-B728-BD5D7DD57AE4.png?generation=1572716480280286&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F1e6fb740684b6cdb77eef4f26d56f8e2%2FD62E03A7-9571-4737-B728-BD5D7DD57AE4.png?generation=1572716480280286&amp;alt=media)\n",
      "votes": 1
    },
    {
      "id": 659680,
      "postDate": "2019-10-28T04:35:13.453Z",
      "content": "<p><img src=\"https://d3i71xaburhd42.cloudfront.net/68e8584f9a26ec2db828ea0cd13b8fc78a0e9a79/3-Figure2-1.png\" alt=\"\"></p>",
      "rawMarkdown": "![](https://d3i71xaburhd42.cloudfront.net/68e8584f9a26ec2db828ea0cd13b8fc78a0e9a79/3-Figure2-1.png)",
      "votes": 1
    },
    {
      "id": 3299349,
      "postDate": "2025-10-07T20:40:07.013Z",
      "content": "<p>First time working with log loss 🙀…Will try your ideas.Thanks for sharing </p>",
      "rawMarkdown": "First time working with log loss 🙀...Will try your ideas.Thanks for sharing "
    },
    {
      "id": 662138,
      "postDate": "2019-10-31T06:04:32.090Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F86746f8aa2a49bfbebe61d3ebf5370d9%2FSelection_060.png?generation=1572501869456145&amp;alt=media\" alt=\"\"></p>\n\n<p>@Jeremy Howard  has a notebook to detect such images. he used it to remove useless images without brain tissue to reduce training set</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F86746f8aa2a49bfbebe61d3ebf5370d9%2FSelection_060.png?generation=1572501869456145&amp;alt=media)\n\n\n@Jeremy Howard  has a notebook to detect such images. he used it to remove useless images without brain tissue to reduce training set"
    },
    {
      "id": 659893,
      "postDate": "2019-10-28T11:55:29.973Z",
      "content": "<p>one example of probability shaping:</p>\n\n<p>```</p>\n\n<p>NAME_TO_LABEL={\n'any' :              0,\n'epidural' :         1,\n'intraparenchymal' : 2,\n'intraventricular' : 3,\n'subarachnoid' :     4,\n'subdural' :         5,\n}\n    if 1: #reshape <br>\n        pos = probability_label[:,1:].max(-1)\n        neg = probability_label[:,1:].min(-1)\n        p0 = probability_label[:,0]</p>\n\n<pre><code>    probability_label[:,0][p0&amp;gt;0.99]=pos[p0&amp;gt;0.99]\n    probability_label[:,0][p0&amp;lt;0.01]=neg[p0&amp;lt;0.01]\n</code></pre>\n\n<p>```</p>\n\n<p>question ask:</p>\n\n<ul>\n<li>can p(any) be less or greater then p(epidural), p(intraparenchymal), etc </li>\n<li>if sum(p(epidural), p(intraparenchymal)) is large, can p(any) be small?</li>\n</ul>\n\n<p>a network to learning the shaping is better than heuristic  guess, e.g. lstm layer label refinement</p>\n\n<p>related: <a href=\"https://github.com/gpleiss/temperature_scaling\">https://github.com/gpleiss/temperature_scaling</a></p>",
      "rawMarkdown": "one example of probability shaping:\n\n```\n\nNAME_TO_LABEL={\n'any' :              0,\n'epidural' :         1,\n'intraparenchymal' : 2,\n'intraventricular' : 3,\n'subarachnoid' :     4,\n'subdural' :         5,\n}\n    if 1: #reshape  \n        pos = probability_label[:,1:].max(-1)\n        neg = probability_label[:,1:].min(-1)\n        p0 = probability_label[:,0]\n\n        probability_label[:,0][p0&gt;0.99]=pos[p0&gt;0.99]\n        probability_label[:,0][p0&lt;0.01]=neg[p0&lt;0.01]\n\n \n\n```\n\n\nquestion ask:\n\n-  can p(any) be less or greater then p(epidural), p(intraparenchymal), etc \n-  if sum(p(epidural), p(intraparenchymal)) is large, can p(any) be small?\n\na network to learning the shaping is better than heuristic  guess, e.g. lstm layer label refinement\n\nrelated: https://github.com/gpleiss/temperature_scaling"
    },
    {
      "id": 659758,
      "postDate": "2019-10-28T08:23:12.383Z",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> About the <em>probability shaping</em>, I'm asking with an open mind: Wouldn't it just \"even out\"? The 99 predictions you correctly round to 0 and the 1 prediction you incorrectly round to 0? So right now I'm thinking, it's better just to trust the Neural net :-D</p>\n\n<p>Another thought: because the any=0 is highly over-represented, the Neural net might tend to make overly pessimistic predictions for any=1 (i.e. give relatively low probabilities for them). Do you think this is True? If yes, would there be a way to solve this in a post-processing step? :-) </p>",
      "rawMarkdown": "@hengck23 About the *probability shaping*, I'm asking with an open mind: Wouldn't it just \"even out\"? The 99 predictions you correctly round to 0 and the 1 prediction you incorrectly round to 0? So right now I'm thinking, it's better just to trust the Neural net :-D\n\nAnother thought: because the any=0 is highly over-represented, the Neural net might tend to make overly pessimistic predictions for any=1 (i.e. give relatively low probabilities for them). Do you think this is True? If yes, would there be a way to solve this in a post-processing step? :-) ",
      "replies": [
        {
          "id": 659761,
          "postDate": "2019-10-28T08:27:15.620Z",
          "content": "<p>\"Wouldn't it just \"even out\"?\" </p>\n\n<p>the log loss is designed to even out. but if you have external post processing, then maybe you can make a gain.</p>\n\n<p>a modification of the log loss may help .... e.g if we know there are more negative samples than positive ...</p>\n\n<hr>\n\n<p>\" So right now I'm thinking, it's better just to trust the Neural net :-D\"</p>\n\n<p>so the sampling will be important for this</p>",
          "rawMarkdown": "\"Wouldn't it just \"even out\"?\" \n\nthe log loss is designed to even out. but if you have external post processing, then maybe you can make a gain.\n\na modification of the log loss may help .... e.g if we know there are more negative samples than positive ...\n\n---\n\n\" So right now I'm thinking, it's better just to trust the Neural net :-D\"\n\nso the sampling will be important for this",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 663815,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-02T17:41:25.107000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F1e6fb740684b6cdb77eef4f26d56f8e2%2FD62E03A7-9571-4737-B728-BD5D7DD57AE4.png?generation=1572716480280286&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 659680,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-10-28T04:35:13.453000",
      "content": "<p><img src=\"https://d3i71xaburhd42.cloudfront.net/68e8584f9a26ec2db828ea0cd13b8fc78a0e9a79/3-Figure2-1.png\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3299349,
      "author_name": "Sanjid Hasan",
      "author_url": "",
      "post_date": "2025-10-07T20:40:07.013000",
      "content": "<p>First time working with log loss 🙀…Will try your ideas.Thanks for sharing </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 662138,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-10-31T06:04:32.090000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F86746f8aa2a49bfbebe61d3ebf5370d9%2FSelection_060.png?generation=1572501869456145&amp;alt=media\" alt=\"\"></p>\n\n<p>@Jeremy Howard  has a notebook to detect such images. he used it to remove useless images without brain tissue to reduce training set</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 659893,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-10-28T11:55:29.973000",
      "content": "<p>one example of probability shaping:</p>\n\n<p>```</p>\n\n<p>NAME_TO_LABEL={\n'any' :              0,\n'epidural' :         1,\n'intraparenchymal' : 2,\n'intraventricular' : 3,\n'subarachnoid' :     4,\n'subdural' :         5,\n}\n    if 1: #reshape <br>\n        pos = probability_label[:,1:].max(-1)\n        neg = probability_label[:,1:].min(-1)\n        p0 = probability_label[:,0]</p>\n\n<pre><code>    probability_label[:,0][p0&amp;gt;0.99]=pos[p0&amp;gt;0.99]\n    probability_label[:,0][p0&amp;lt;0.01]=neg[p0&amp;lt;0.01]\n</code></pre>\n\n<p>```</p>\n\n<p>question ask:</p>\n\n<ul>\n<li>can p(any) be less or greater then p(epidural), p(intraparenchymal), etc </li>\n<li>if sum(p(epidural), p(intraparenchymal)) is large, can p(any) be small?</li>\n</ul>\n\n<p>a network to learning the shaping is better than heuristic  guess, e.g. lstm layer label refinement</p>\n\n<p>related: <a href=\"https://github.com/gpleiss/temperature_scaling\">https://github.com/gpleiss/temperature_scaling</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 659758,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2019-10-28T08:23:12.383000",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> About the <em>probability shaping</em>, I'm asking with an open mind: Wouldn't it just \"even out\"? The 99 predictions you correctly round to 0 and the 1 prediction you incorrectly round to 0? So right now I'm thinking, it's better just to trust the Neural net :-D</p>\n\n<p>Another thought: because the any=0 is highly over-represented, the Neural net might tend to make overly pessimistic predictions for any=1 (i.e. give relatively low probabilities for them). Do you think this is True? If yes, would there be a way to solve this in a post-processing step? :-) </p>",
      "votes": 0,
      "replies": [
        {
          "id": 659761,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-10-28T08:27:15.620000",
          "content": "<p>\"Wouldn't it just \"even out\"?\" </p>\n\n<p>the log loss is designed to even out. but if you have external post processing, then maybe you can make a gain.</p>\n\n<p>a modification of the log loss may help .... e.g if we know there are more negative samples than positive ...</p>\n\n<hr>\n\n<p>\" So right now I'm thinking, it's better just to trust the Neural net :-D\"</p>\n\n<p>so the sampling will be important for this</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "659671": "1. sampling/weighing is important in training\n\n- when the metric is accuracy, we usually use : predict\\_label = predict\\_prob &gt; threshold\nwe do not care about predict\\_prob, but rather thresholded prediction\n\nwe sometimes use loss weighing (e.g focal loss) or class balancing to speed up training or improve accuracy. but improving accuracy is \"not the same\" as reducing loss. the loss is modified to improve accuracy by artificially changing the train distribution and affecting the decision boundary. \n\n**in summary. we care only about the boundary, and care less about the decision value**\n\n\nfor this challenge, we care about the actual log loss itself. so it is important to **\"calibrate your probability\"** to the actual one you expect in the test during training. you may have to use random sampling (instead of balanced sampling) at the final fine tuning. likewise you may need to disable any sample weighing.\n\n---\n\n2. decimal is important\n\nthere will be some different between probability expressed in float 32 bit and 64 bit\n\n\n---\n\n3. probability shaping\n\nthis is is very difficult (and very risky in scoring). e.g. a test sample is predicted to be negative with probability 0.0001. you may round it to zero.  you will gain a little if the truth is really negative, but you will lose big if it is truth is instead positive.\n\nif you can train with \"large margin\" and use some probability calibration trick such that  you are almost 100% sure that there cannot be positive truth below a certain low values, you can zero out the values. \n",
    "663815": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F1e6fb740684b6cdb77eef4f26d56f8e2%2FD62E03A7-9571-4737-B728-BD5D7DD57AE4.png?generation=1572716480280286&amp;alt=media)\n",
    "659680": "![](https://d3i71xaburhd42.cloudfront.net/68e8584f9a26ec2db828ea0cd13b8fc78a0e9a79/3-Figure2-1.png)",
    "3299349": "First time working with log loss 🙀...Will try your ideas.Thanks for sharing ",
    "662138": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F86746f8aa2a49bfbebe61d3ebf5370d9%2FSelection_060.png?generation=1572501869456145&amp;alt=media)\n\n\n@Jeremy Howard  has a notebook to detect such images. he used it to remove useless images without brain tissue to reduce training set",
    "659893": "one example of probability shaping:\n\n```\n\nNAME_TO_LABEL={\n'any' :              0,\n'epidural' :         1,\n'intraparenchymal' : 2,\n'intraventricular' : 3,\n'subarachnoid' :     4,\n'subdural' :         5,\n}\n    if 1: #reshape  \n        pos = probability_label[:,1:].max(-1)\n        neg = probability_label[:,1:].min(-1)\n        p0 = probability_label[:,0]\n\n        probability_label[:,0][p0&gt;0.99]=pos[p0&gt;0.99]\n        probability_label[:,0][p0&lt;0.01]=neg[p0&lt;0.01]\n\n \n\n```\n\n\nquestion ask:\n\n-  can p(any) be less or greater then p(epidural), p(intraparenchymal), etc \n-  if sum(p(epidural), p(intraparenchymal)) is large, can p(any) be small?\n\na network to learning the shaping is better than heuristic  guess, e.g. lstm layer label refinement\n\nrelated: https://github.com/gpleiss/temperature_scaling",
    "659758": "@hengck23 About the *probability shaping*, I'm asking with an open mind: Wouldn't it just \"even out\"? The 99 predictions you correctly round to 0 and the 1 prediction you incorrectly round to 0? So right now I'm thinking, it's better just to trust the Neural net :-D\n\nAnother thought: because the any=0 is highly over-represented, the Neural net might tend to make overly pessimistic predictions for any=1 (i.e. give relatively low probabilities for them). Do you think this is True? If yes, would there be a way to solve this in a post-processing step? :-) "
  }
}