{
  "id": 190155,
  "title": "Question about consistency rules",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/190155",
  "author_name": "Alexander Soare",
  "post_date": "2020-10-10T12:21:23.775000",
  "votes": 6,
  "comment_count": 5,
  "views": 0,
  "content": "<p>From the <a href=\"https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check\" target=\"_blank\">HOST CONFIRMED - Label Consistency Check</a> notebook I see that:</p>\n<blockquote>\n  <p>If there is at least one image per <code>StudyInstanceUID</code> with <code>pe_present_on_image</code> &gt; 0.5, then…</p>\n</blockquote>\n<p>refers to situations in which we believe the study indicates positive for PE.</p>\n<p>I'm unsure about this though. What if all N images in the study have <code>pe_present_on_image</code> = 0.49. Then, assuming independence of images, the probability that the study indicates negative for PE is (1-0.49)^N, which with just N=2 becomes around 0.25. And I assume you now see where I'm going with this.</p>\n<p>So I don't think it's necessary that at least one image to have <code>pe_present_on_image</code> &gt; 0.5 for us to conclude that the study has a more than 0.5 chance of indicating PE. Therefore I think it would make sense to be able to specify things like RV/LV ratio, the location of the PE, or the nature of the condition (chronic/acute).</p>",
  "messages": [
    {
      "id": 1045184,
      "postDate": "2020-10-10T12:21:23.777Z",
      "content": "<p>From the <a href=\"https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check\" target=\"_blank\">HOST CONFIRMED - Label Consistency Check</a> notebook I see that:</p>\n<blockquote>\n  <p>If there is at least one image per <code>StudyInstanceUID</code> with <code>pe_present_on_image</code> &gt; 0.5, then…</p>\n</blockquote>\n<p>refers to situations in which we believe the study indicates positive for PE.</p>\n<p>I'm unsure about this though. What if all N images in the study have <code>pe_present_on_image</code> = 0.49. Then, assuming independence of images, the probability that the study indicates negative for PE is (1-0.49)^N, which with just N=2 becomes around 0.25. And I assume you now see where I'm going with this.</p>\n<p>So I don't think it's necessary that at least one image to have <code>pe_present_on_image</code> &gt; 0.5 for us to conclude that the study has a more than 0.5 chance of indicating PE. Therefore I think it would make sense to be able to specify things like RV/LV ratio, the location of the PE, or the nature of the condition (chronic/acute).</p>",
      "rawMarkdown": "From the [HOST CONFIRMED - Label Consistency Check](https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check) notebook I see that:\n\n> If there is at least one image per `StudyInstanceUID` with `pe_present_on_image` > 0.5, then...\n\nrefers to situations in which we believe the study indicates positive for PE.\n\nI'm unsure about this though. What if all N images in the study have `pe_present_on_image` = 0.49. Then, assuming independence of images, the probability that the study indicates negative for PE is (1-0.49)^N, which with just N=2 becomes around 0.25. And I assume you now see where I'm going with this.\n\nSo I don't think it's necessary that at least one image to have `pe_present_on_image` > 0.5 for us to conclude that the study has a more than 0.5 chance of indicating PE. Therefore I think it would make sense to be able to specify things like RV/LV ratio, the location of the PE, or the nature of the condition (chronic/acute).",
      "votes": 6
    },
    {
      "id": 1045249,
      "postDate": "2020-10-10T13:05:02.353Z",
      "content": "<p>I agree with <a href=\"https://www.kaggle.com/alexandersoare\" target=\"_blank\">@alexandersoare</a>. These consistency rules add unnecessary complexity to the problem. I don't find it contradictory that a model can identify a suspicious image to have &gt;50% chance of a PE, but when reviewing the entire stack of images as a whole conclude that the study is negative for the abnormality. What this inadvertently will cause is that people will manually write codes on top of model outputs, which will paradoxically make AI outputs even more difficult to explain. </p>",
      "rawMarkdown": "I agree with @alexandersoare. These consistency rules add unnecessary complexity to the problem. I don't find it contradictory that a model can identify a suspicious image to have >50% chance of a PE, but when reviewing the entire stack of images as a whole conclude that the study is negative for the abnormality. What this inadvertently will cause is that people will manually write codes on top of model outputs, which will paradoxically make AI outputs even more difficult to explain. ",
      "votes": 4,
      "replies": [
        {
          "id": 1050729,
          "postDate": "2020-10-15T17:06:24.543Z",
          "content": "<p><a href=\"https://www.kaggle.com/yeeseng\" target=\"_blank\">@yeeseng</a> when trying to met label consistency on above of the AI model than getting higher loss. Did you get your current leaderboard loss with or without meeting label consistency requirements?</p>",
          "rawMarkdown": "@yeeseng when trying to met label consistency on above of the AI model than getting higher loss. Did you get your current leaderboard loss with or without meeting label consistency requirements?"
        }
      ]
    },
    {
      "id": 1057689,
      "postDate": "2020-10-22T21:49:38.867Z",
      "content": "<p>These issues were previously raised and <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183473\" target=\"_blank\">discussed</a>.  <a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> nicely <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183473#1014525\" target=\"_blank\">summarized the rationale</a> for the rules as currently written.</p>",
      "rawMarkdown": "These issues were previously raised and [discussed](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183473).  @richardepstein nicely [summarized the rationale](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183473#1014525) for the rules as currently written.",
      "replies": [
        {
          "id": 1058035,
          "postDate": "2020-10-23T09:11:03.080Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> thanks for the response. I've read that summary and I agree. What I'm referring to is how we define \"medically consistent\". I think that if all images in a study have 49% predicted probability of having PE present, then it is medically <strong>inconsistent</strong> to then have to say that the overall study has more than 50% chance of being negative for PE.</p>",
          "rawMarkdown": "Hi @anthracene thanks for the response. I've read that summary and I agree. What I'm referring to is how we define \"medically consistent\". I think that if all images in a study have 49% predicted probability of having PE present, then it is medically **inconsistent** to then have to say that the overall study has more than 50% chance of being negative for PE."
        }
      ]
    },
    {
      "id": 1047006,
      "postDate": "2020-10-12T07:13:30.780Z",
      "content": "<p><a href=\"https://www.kaggle.com/alexandersoare\" target=\"_blank\">@alexandersoare</a> may be mean could make sense as all images belong to same CT ..</p>",
      "rawMarkdown": "@alexandersoare may be mean could make sense as all images belong to same CT .."
    }
  ],
  "comments": [
    {
      "id": 1045249,
      "author_name": "Yee Ng",
      "author_url": "",
      "post_date": "2020-10-10T13:05:02.353000",
      "content": "<p>I agree with <a href=\"https://www.kaggle.com/alexandersoare\" target=\"_blank\">@alexandersoare</a>. These consistency rules add unnecessary complexity to the problem. I don't find it contradictory that a model can identify a suspicious image to have &gt;50% chance of a PE, but when reviewing the entire stack of images as a whole conclude that the study is negative for the abnormality. What this inadvertently will cause is that people will manually write codes on top of model outputs, which will paradoxically make AI outputs even more difficult to explain. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 1050729,
          "author_name": "SumanSudhir",
          "author_url": "",
          "post_date": "2020-10-15T17:06:24.543000",
          "content": "<p><a href=\"https://www.kaggle.com/yeeseng\" target=\"_blank\">@yeeseng</a> when trying to met label consistency on above of the AI model than getting higher loss. Did you get your current leaderboard loss with or without meeting label consistency requirements?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1057689,
      "author_name": "John Mongan",
      "author_url": "",
      "post_date": "2020-10-22T21:49:38.867000",
      "content": "<p>These issues were previously raised and <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183473\" target=\"_blank\">discussed</a>.  <a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> nicely <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183473#1014525\" target=\"_blank\">summarized the rationale</a> for the rules as currently written.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1058035,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2020-10-23T09:11:03.080000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> thanks for the response. I've read that summary and I agree. What I'm referring to is how we define \"medically consistent\". I think that if all images in a study have 49% predicted probability of having PE present, then it is medically <strong>inconsistent</strong> to then have to say that the overall study has more than 50% chance of being negative for PE.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1047006,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2020-10-12T07:13:30.780000",
      "content": "<p><a href=\"https://www.kaggle.com/alexandersoare\" target=\"_blank\">@alexandersoare</a> may be mean could make sense as all images belong to same CT ..</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1045184": "From the [HOST CONFIRMED - Label Consistency Check](https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check) notebook I see that:\n\n> If there is at least one image per `StudyInstanceUID` with `pe_present_on_image` > 0.5, then...\n\nrefers to situations in which we believe the study indicates positive for PE.\n\nI'm unsure about this though. What if all N images in the study have `pe_present_on_image` = 0.49. Then, assuming independence of images, the probability that the study indicates negative for PE is (1-0.49)^N, which with just N=2 becomes around 0.25. And I assume you now see where I'm going with this.\n\nSo I don't think it's necessary that at least one image to have `pe_present_on_image` > 0.5 for us to conclude that the study has a more than 0.5 chance of indicating PE. Therefore I think it would make sense to be able to specify things like RV/LV ratio, the location of the PE, or the nature of the condition (chronic/acute).",
    "1045249": "I agree with @alexandersoare. These consistency rules add unnecessary complexity to the problem. I don't find it contradictory that a model can identify a suspicious image to have >50% chance of a PE, but when reviewing the entire stack of images as a whole conclude that the study is negative for the abnormality. What this inadvertently will cause is that people will manually write codes on top of model outputs, which will paradoxically make AI outputs even more difficult to explain. ",
    "1057689": "These issues were previously raised and [discussed](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183473).  @richardepstein nicely [summarized the rationale](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183473#1014525) for the rules as currently written.",
    "1047006": "@alexandersoare may be mean could make sense as all images belong to same CT .."
  }
}