{
  "id": 447452,
  "title": "Any_injury calculation is faulty",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/447452",
  "author_name": "Raki",
  "post_date": "2023-10-16T00:22:00.195000",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I think the any_injury calculation was faulty, leading to a lot of problems, with scoring and model development.</p>\n<h1>The Problem</h1>\n<p>If you look in the <a href=\"https://www.kaggle.com/code/metric/rsna-trauma-metric/notebook\" target=\"_blank\">scoring notebook</a> it states: “2. Derive a new any_injury label by taking the max of 1 - p(healthy) for each label group.” and implements as: <code>any_injury_labels = (1 - solution[healthy_cols]).max(axis=1)</code>.<br>\nThis leads to extremely underestimating any_injury though!</p>\n<p>Let’s assume we have the case of independent probabilities for injury of <br>\n<code>healthy_pred = hp = [0.5, 0.5, 0.5, 0.5, 0.5]</code>  the true chance for any injury would be:<br>\n=&gt; <code>any_injury = 1 – hp[0] * hp[1]… = 31/32 = 0.97</code>.</p>\n<p>The max calculation screws this up. It just gives 0.5 for any_injury. Which is much too low.<br>\nIt would be best to deliberately overstate the injury probability for one injury type by a large amount, taking an increase in the error for this injury type, but vastly decreasing your loss in the any_injury category. </p>\n<h1>Demonstration</h1>\n<p>For this reason I also think that extravasation injury was not extremely over-represented in public test. I think it was only the injury type that the community converged on overestimating for this tradeoff, as it was empirically obvious that it helped decreasing loss. I think that it would have worked (almost) as well to use another category instead. </p>\n<p>I show an example here, you can see the implementation in <a href=\"https://www.kaggle.com/raki21/faulty-any-injury-in-metric-demonstration\" target=\"_blank\">this notebook</a>. </p>\n<p>You can see that predicting very high extravasation injury is detrimental to that category score, but helps by decreasing the overall loss.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F2891ab5c5f070d8fbefc2b70ae141df3%2FExtravasation%20and%20any_injury.png?generation=1697414956170079&amp;alt=media\" alt=\"\"></p>\n<p>You could also do this with other injury categories though!<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F6ec1f167993e6386ec894d46a89cfbd1%2FBowel%20and%20any_injury.png?generation=1697414962343909&amp;alt=media\" alt=\"\"></p>\n<h1>Consequences</h1>\n<p>I think a huge problem of people finding this out empirically was the following:<br>\nIf you take the <a href=\"https://www.kaggle.com/code/jasonheesanglee/rsna23-scale-h-implementation\" target=\"_blank\">mean approach</a> as baseline  and then see that you get a much better score/lower loss with high extravasation injury multiplier, you might wrongly conclude that there is much more extravasation injury and postprocess any model output accordingly.</p>\n<p>Let’s say you now trained a pretty good model that is especially strong at finding out when there is no extravasation injury present. If you now just correct extravasation injury upwards it probably only leads to worse score. You might now increase your expected extravasation group loss without getting the benefit of decreased any_injury loss as your extravasation is not the smallest healthy column anymore. So even though your predictions are better than the mean approach, the post processing is so much worse that you get a terrible score. </p>\n<p>A lot of people might have scrapped good models because empirically they didn’t perform well because of the bad any_injury implementation and them misidentifying the problem as overrepresentation of extravasation.<br>\nI think the correct approach for people to postprocess based on the faulty any_injury calculation would have been something like this:</p>\n<h1>Possible Post-Processing</h1>\n<ol>\n<li>Based on unweighted healthy columns do something like:</li>\n</ol>\n<pre><code>y_pred[] = (y_pred[] \n                              * y_pred[] \n                              * y_pred[] \n                              * y_pred[] \n                              * y_pred[])    \n\ny_pred[] =  – y_pred[]\ny_pred[] = (ln(y_pred[])*)**e \n</code></pre>\n<p>OR also train a model for any_injury prediction, which might be better as injuries are not independent in reality!</p>\n<ol>\n<li>Adjust other columns to weight and normalize healthy and injury columns to 1.</li>\n<li>Select column with lowest healthy prediction.</li>\n<li>Do this:</li>\n</ol>\n<pre><code>y_pred[highest_injury_type] = y_pred[]\ny_pred[highest_injury_type_healthy] =  – y_pred[highest_injury_type]\n</code></pre>\n<p>Keep in mind that this is still suboptimal with a much more complex calculation needed for perfect postprocessing. But this should be much better than just putting a factor x on extravasation injury.</p>\n<p>I didn’t share this before the competition end, because I only really joined the competition on Friday and thought it would be pretty detrimental to share the info at that point, maybe tossing up scores. I hope it is still illuminating.</p>",
  "messages": [
    {
      "id": 2483737,
      "postDate": "2023-10-16T00:22:00.197Z",
      "content": "<p>I think the any_injury calculation was faulty, leading to a lot of problems, with scoring and model development.</p>\n<h1>The Problem</h1>\n<p>If you look in the <a href=\"https://www.kaggle.com/code/metric/rsna-trauma-metric/notebook\" target=\"_blank\">scoring notebook</a> it states: “2. Derive a new any_injury label by taking the max of 1 - p(healthy) for each label group.” and implements as: <code>any_injury_labels = (1 - solution[healthy_cols]).max(axis=1)</code>.<br>\nThis leads to extremely underestimating any_injury though!</p>\n<p>Let’s assume we have the case of independent probabilities for injury of <br>\n<code>healthy_pred = hp = [0.5, 0.5, 0.5, 0.5, 0.5]</code>  the true chance for any injury would be:<br>\n=&gt; <code>any_injury = 1 – hp[0] * hp[1]… = 31/32 = 0.97</code>.</p>\n<p>The max calculation screws this up. It just gives 0.5 for any_injury. Which is much too low.<br>\nIt would be best to deliberately overstate the injury probability for one injury type by a large amount, taking an increase in the error for this injury type, but vastly decreasing your loss in the any_injury category. </p>\n<h1>Demonstration</h1>\n<p>For this reason I also think that extravasation injury was not extremely over-represented in public test. I think it was only the injury type that the community converged on overestimating for this tradeoff, as it was empirically obvious that it helped decreasing loss. I think that it would have worked (almost) as well to use another category instead. </p>\n<p>I show an example here, you can see the implementation in <a href=\"https://www.kaggle.com/raki21/faulty-any-injury-in-metric-demonstration\" target=\"_blank\">this notebook</a>. </p>\n<p>You can see that predicting very high extravasation injury is detrimental to that category score, but helps by decreasing the overall loss.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F2891ab5c5f070d8fbefc2b70ae141df3%2FExtravasation%20and%20any_injury.png?generation=1697414956170079&amp;alt=media\" alt=\"\"></p>\n<p>You could also do this with other injury categories though!<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F6ec1f167993e6386ec894d46a89cfbd1%2FBowel%20and%20any_injury.png?generation=1697414962343909&amp;alt=media\" alt=\"\"></p>\n<h1>Consequences</h1>\n<p>I think a huge problem of people finding this out empirically was the following:<br>\nIf you take the <a href=\"https://www.kaggle.com/code/jasonheesanglee/rsna23-scale-h-implementation\" target=\"_blank\">mean approach</a> as baseline  and then see that you get a much better score/lower loss with high extravasation injury multiplier, you might wrongly conclude that there is much more extravasation injury and postprocess any model output accordingly.</p>\n<p>Let’s say you now trained a pretty good model that is especially strong at finding out when there is no extravasation injury present. If you now just correct extravasation injury upwards it probably only leads to worse score. You might now increase your expected extravasation group loss without getting the benefit of decreased any_injury loss as your extravasation is not the smallest healthy column anymore. So even though your predictions are better than the mean approach, the post processing is so much worse that you get a terrible score. </p>\n<p>A lot of people might have scrapped good models because empirically they didn’t perform well because of the bad any_injury implementation and them misidentifying the problem as overrepresentation of extravasation.<br>\nI think the correct approach for people to postprocess based on the faulty any_injury calculation would have been something like this:</p>\n<h1>Possible Post-Processing</h1>\n<ol>\n<li>Based on unweighted healthy columns do something like:</li>\n</ol>\n<pre><code>y_pred[] = (y_pred[] \n                              * y_pred[] \n                              * y_pred[] \n                              * y_pred[] \n                              * y_pred[])    \n\ny_pred[] =  – y_pred[]\ny_pred[] = (ln(y_pred[])*)**e \n</code></pre>\n<p>OR also train a model for any_injury prediction, which might be better as injuries are not independent in reality!</p>\n<ol>\n<li>Adjust other columns to weight and normalize healthy and injury columns to 1.</li>\n<li>Select column with lowest healthy prediction.</li>\n<li>Do this:</li>\n</ol>\n<pre><code>y_pred[highest_injury_type] = y_pred[]\ny_pred[highest_injury_type_healthy] =  – y_pred[highest_injury_type]\n</code></pre>\n<p>Keep in mind that this is still suboptimal with a much more complex calculation needed for perfect postprocessing. But this should be much better than just putting a factor x on extravasation injury.</p>\n<p>I didn’t share this before the competition end, because I only really joined the competition on Friday and thought it would be pretty detrimental to share the info at that point, maybe tossing up scores. I hope it is still illuminating.</p>",
      "rawMarkdown": "I think the any_injury calculation was faulty, leading to a lot of problems, with scoring and model development.\n\n#The Problem\nIf you look in the [scoring notebook](https://www.kaggle.com/code/metric/rsna-trauma-metric/notebook) it states: “2. Derive a new any_injury label by taking the max of 1 - p(healthy) for each label group.” and implements as: `any_injury_labels = (1 - solution[healthy_cols]).max(axis=1)`.\nThis leads to extremely underestimating any_injury though!\n\nLet’s assume we have the case of independent probabilities for injury of \n`healthy_pred = hp = [0.5, 0.5, 0.5, 0.5, 0.5]`  the true chance for any injury would be:\n=> `any_injury = 1 – hp[0] * hp[1]… = 31/32 = 0.97`.\n\nThe max calculation screws this up. It just gives 0.5 for any_injury. Which is much too low.\nIt would be best to deliberately overstate the injury probability for one injury type by a large amount, taking an increase in the error for this injury type, but vastly decreasing your loss in the any_injury category. \n\n#Demonstration\nFor this reason I also think that extravasation injury was not extremely over-represented in public test. I think it was only the injury type that the community converged on overestimating for this tradeoff, as it was empirically obvious that it helped decreasing loss. I think that it would have worked (almost) as well to use another category instead. \n\nI show an example here, you can see the implementation in [this notebook](https://www.kaggle.com/raki21/faulty-any-injury-in-metric-demonstration). \n\nYou can see that predicting very high extravasation injury is detrimental to that category score, but helps by decreasing the overall loss.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F2891ab5c5f070d8fbefc2b70ae141df3%2FExtravasation%20and%20any_injury.png?generation=1697414956170079&alt=media)\n\nYou could also do this with other injury categories though!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F6ec1f167993e6386ec894d46a89cfbd1%2FBowel%20and%20any_injury.png?generation=1697414962343909&alt=media)\n\n#Consequences\nI think a huge problem of people finding this out empirically was the following:\nIf you take the [mean approach](https://www.kaggle.com/code/jasonheesanglee/rsna23-scale-h-implementation) as baseline  and then see that you get a much better score/lower loss with high extravasation injury multiplier, you might wrongly conclude that there is much more extravasation injury and postprocess any model output accordingly.\n\nLet’s say you now trained a pretty good model that is especially strong at finding out when there is no extravasation injury present. If you now just correct extravasation injury upwards it probably only leads to worse score. You might now increase your expected extravasation group loss without getting the benefit of decreased any_injury loss as your extravasation is not the smallest healthy column anymore. So even though your predictions are better than the mean approach, the post processing is so much worse that you get a terrible score. \n\nA lot of people might have scrapped good models because empirically they didn’t perform well because of the bad any_injury implementation and them misidentifying the problem as overrepresentation of extravasation.\nI think the correct approach for people to postprocess based on the faulty any_injury calculation would have been something like this:\n\n#Possible Post-Processing\n1. Based on unweighted healthy columns do something like:\n```python\ny_pred['real_any_healthy'] = (y_pred['bowel_healthy'] \n                              * y_pred['extravasation_healthy'] \n                              * y_pred['kidney_healthy'] \n                              * y_pred['liver_healthy'] \n                              * y_pred['spleen_healthy'])    \n\ny_pred['real_any_injury'] = 1 – y_pred['real_any_healthy']\ny_pred['weighted_any_injury'] = (ln(y_pred['real_any_healthy'])*6)**e \n```\nOR also train a model for any_injury prediction, which might be better as injuries are not independent in reality!\n\n2. Adjust other columns to weight and normalize healthy and injury columns to 1.\n4. Select column with lowest healthy prediction.\n5. Do this:\n```python\ny_pred[highest_injury_type] = y_pred['weighted_any_injury']\ny_pred[highest_injury_type_healthy] = 1 – y_pred[highest_injury_type]\n```\n\nKeep in mind that this is still suboptimal with a much more complex calculation needed for perfect postprocessing. But this should be much better than just putting a factor x on extravasation injury.\n\nI didn’t share this before the competition end, because I only really joined the competition on Friday and thought it would be pretty detrimental to share the info at that point, maybe tossing up scores. I hope it is still illuminating.\n",
      "votes": 5
    },
    {
      "id": 2483741,
      "postDate": "2023-10-16T00:31:40.173Z",
      "content": "<p>I believe someone raised this problem in the middle of the competition. In terms of results,  artificially increasing the probability of extravasation injury was indeed a trick in this competition. </p>\n<p>However, using the product of each health label's probability may also cause some problems, such as making the probability of any injury too high. I don't know which calculation method would be better.</p>",
      "rawMarkdown": "I believe someone raised this problem in the middle of the competition. In terms of results,  artificially increasing the probability of extravasation injury was indeed a trick in this competition. \n\nHowever, using the product of each health label's probability may also cause some problems, such as making the probability of any injury too high. I don't know which calculation method would be better.",
      "votes": 1,
      "replies": [
        {
          "id": 2483762,
          "postDate": "2023-10-16T01:06:23.310Z",
          "content": "<p>I expect it would be best to have a model make an any_injury probability prediction, scale it according to weight 6, then use that to adjust upwards the highest other injury prediction, so that the minimization of the any_injury loss does come at the lowest possible loss increase in another category. Even better to try to find an equilibrium, but that equilibrium will be heavily skewed towards any_injury.</p>",
          "rawMarkdown": "I expect it would be best to have a model make an any_injury probability prediction, scale it according to weight 6, then use that to adjust upwards the highest other injury prediction, so that the minimization of the any_injury loss does come at the lowest possible loss increase in another category. Even better to try to find an equilibrium, but that equilibrium will be heavily skewed towards any_injury.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2483903,
      "postDate": "2023-10-16T04:52:50.843Z",
      "content": "<p>Did question this <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/438020\" target=\"_blank\">here</a> and referred to posts on scaling up/post processing extravasation.  But little response.  Expect the metric was set up and would not change.  Pity you joined late. </p>",
      "rawMarkdown": "Did question this [here](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/438020) and referred to posts on scaling up/post processing extravasation.  But little response.  Expect the metric was set up and would not change.  Pity you joined late. \n"
    }
  ],
  "comments": [
    {
      "id": 2483741,
      "author_name": "NorthM344",
      "author_url": "",
      "post_date": "2023-10-16T00:31:40.173000",
      "content": "<p>I believe someone raised this problem in the middle of the competition. In terms of results,  artificially increasing the probability of extravasation injury was indeed a trick in this competition. </p>\n<p>However, using the product of each health label's probability may also cause some problems, such as making the probability of any injury too high. I don't know which calculation method would be better.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2483762,
          "author_name": "Raki",
          "author_url": "",
          "post_date": "2023-10-16T01:06:23.310000",
          "content": "<p>I expect it would be best to have a model make an any_injury probability prediction, scale it according to weight 6, then use that to adjust upwards the highest other injury prediction, so that the minimization of the any_injury loss does come at the lowest possible loss increase in another category. Even better to try to find an equilibrium, but that equilibrium will be heavily skewed towards any_injury.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2483903,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2023-10-16T04:52:50.843000",
      "content": "<p>Did question this <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/438020\" target=\"_blank\">here</a> and referred to posts on scaling up/post processing extravasation.  But little response.  Expect the metric was set up and would not change.  Pity you joined late. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2483737": "I think the any_injury calculation was faulty, leading to a lot of problems, with scoring and model development.\n\n#The Problem\nIf you look in the [scoring notebook](https://www.kaggle.com/code/metric/rsna-trauma-metric/notebook) it states: “2. Derive a new any_injury label by taking the max of 1 - p(healthy) for each label group.” and implements as: `any_injury_labels = (1 - solution[healthy_cols]).max(axis=1)`.\nThis leads to extremely underestimating any_injury though!\n\nLet’s assume we have the case of independent probabilities for injury of \n`healthy_pred = hp = [0.5, 0.5, 0.5, 0.5, 0.5]`  the true chance for any injury would be:\n=> `any_injury = 1 – hp[0] * hp[1]… = 31/32 = 0.97`.\n\nThe max calculation screws this up. It just gives 0.5 for any_injury. Which is much too low.\nIt would be best to deliberately overstate the injury probability for one injury type by a large amount, taking an increase in the error for this injury type, but vastly decreasing your loss in the any_injury category. \n\n#Demonstration\nFor this reason I also think that extravasation injury was not extremely over-represented in public test. I think it was only the injury type that the community converged on overestimating for this tradeoff, as it was empirically obvious that it helped decreasing loss. I think that it would have worked (almost) as well to use another category instead. \n\nI show an example here, you can see the implementation in [this notebook](https://www.kaggle.com/raki21/faulty-any-injury-in-metric-demonstration). \n\nYou can see that predicting very high extravasation injury is detrimental to that category score, but helps by decreasing the overall loss.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F2891ab5c5f070d8fbefc2b70ae141df3%2FExtravasation%20and%20any_injury.png?generation=1697414956170079&alt=media)\n\nYou could also do this with other injury categories though!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3747152%2F6ec1f167993e6386ec894d46a89cfbd1%2FBowel%20and%20any_injury.png?generation=1697414962343909&alt=media)\n\n#Consequences\nI think a huge problem of people finding this out empirically was the following:\nIf you take the [mean approach](https://www.kaggle.com/code/jasonheesanglee/rsna23-scale-h-implementation) as baseline  and then see that you get a much better score/lower loss with high extravasation injury multiplier, you might wrongly conclude that there is much more extravasation injury and postprocess any model output accordingly.\n\nLet’s say you now trained a pretty good model that is especially strong at finding out when there is no extravasation injury present. If you now just correct extravasation injury upwards it probably only leads to worse score. You might now increase your expected extravasation group loss without getting the benefit of decreased any_injury loss as your extravasation is not the smallest healthy column anymore. So even though your predictions are better than the mean approach, the post processing is so much worse that you get a terrible score. \n\nA lot of people might have scrapped good models because empirically they didn’t perform well because of the bad any_injury implementation and them misidentifying the problem as overrepresentation of extravasation.\nI think the correct approach for people to postprocess based on the faulty any_injury calculation would have been something like this:\n\n#Possible Post-Processing\n1. Based on unweighted healthy columns do something like:\n```python\ny_pred['real_any_healthy'] = (y_pred['bowel_healthy'] \n                              * y_pred['extravasation_healthy'] \n                              * y_pred['kidney_healthy'] \n                              * y_pred['liver_healthy'] \n                              * y_pred['spleen_healthy'])    \n\ny_pred['real_any_injury'] = 1 – y_pred['real_any_healthy']\ny_pred['weighted_any_injury'] = (ln(y_pred['real_any_healthy'])*6)**e \n```\nOR also train a model for any_injury prediction, which might be better as injuries are not independent in reality!\n\n2. Adjust other columns to weight and normalize healthy and injury columns to 1.\n4. Select column with lowest healthy prediction.\n5. Do this:\n```python\ny_pred[highest_injury_type] = y_pred['weighted_any_injury']\ny_pred[highest_injury_type_healthy] = 1 – y_pred[highest_injury_type]\n```\n\nKeep in mind that this is still suboptimal with a much more complex calculation needed for perfect postprocessing. But this should be much better than just putting a factor x on extravasation injury.\n\nI didn’t share this before the competition end, because I only really joined the competition on Friday and thought it would be pretty detrimental to share the info at that point, maybe tossing up scores. I hope it is still illuminating.\n",
    "2483741": "I believe someone raised this problem in the middle of the competition. In terms of results,  artificially increasing the probability of extravasation injury was indeed a trick in this competition. \n\nHowever, using the product of each health label's probability may also cause some problems, such as making the probability of any injury too high. I don't know which calculation method would be better.",
    "2483903": "Did question this [here](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/438020) and referred to posts on scaling up/post processing extravasation.  But little response.  Expect the metric was set up and would not change.  Pity you joined late. \n"
  }
}