{
  "id": 145596,
  "title": "How is ground truth decided for measuring pathologist accuracy?",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/145596",
  "author_name": "Matt",
  "post_date": "2020-04-23T19:50:11.158000",
  "votes": 0,
  "comment_count": 6,
  "views": 0,
  "content": "<p>My question does not pertain specifically to the competition dataset but in regards to the literature in general.</p>\n\n<p>Many of the published articles that show a deep learning system being used to predict Gleason scores compares their accuracy against the accuracy of pathologists. For exmaple, in <a href=\"https://www.nature.com/articles/s41746-019-0112-2.pdf\">this Nature paper by Google AI</a>, they say the accuracy of 29 pathologists was 0.61 while their model achieved 0.7. But how do they actually decide what the ground truth is? It must have been assigned by a pathologist, but who's to say that label is necessarily \"correct\" compared to the other pathologists? Is it decided by consensus like the dataset in this competition?</p>\n\n<p>Thanks for responses in advance, I hope I explained that clearly enough.</p>",
  "messages": [
    {
      "id": 818814,
      "postDate": "2020-04-24T06:17:15.587Z",
      "content": "<p>Good question - and indeed a hot topic in the field. In most cases the ground truth is constructed in similar fashion to the test data of this challenge - based on the consensus of several pathologists. In some cases the assessment performed by a pathologist, who is specialized in the disease/tissue in question and/or has many years of experience is treated as ground truth, and the assessments by general pathologists are compared against the expert. This of course has pitfalls as even the most knowledgeable experts may simply have a bad day from time to time and there is variation between even highly experienced pathologists.</p>\n\n<p>In some cancers, there are proteins or other molecules whose presence in the tissue is recognized as a strong indication of cancer, and they can be picked up by applying a specific chemical stain (immunohistochemistry, IHC) before scanning the image. In most cases IHC is not done for all samples due to costs and the extra work, but in a research setting it's been sometimes performed to obtain a potentially more reliable ground truth. The pathologists' performance on the usual H&amp;E stained tissue (i.e. the type used in this challenge) can then be compared against the \"biochemical ground truth\". Training models on the IHC-based labels is sometimes called \"antibody-supervised learning\".</p>",
      "rawMarkdown": "Good question - and indeed a hot topic in the field. In most cases the ground truth is constructed in similar fashion to the test data of this challenge - based on the consensus of several pathologists. In some cases the assessment performed by a pathologist, who is specialized in the disease/tissue in question and/or has many years of experience is treated as ground truth, and the assessments by general pathologists are compared against the expert. This of course has pitfalls as even the most knowledgeable experts may simply have a bad day from time to time and there is variation between even highly experienced pathologists.\n\nIn some cancers, there are proteins or other molecules whose presence in the tissue is recognized as a strong indication of cancer, and they can be picked up by applying a specific chemical stain (immunohistochemistry, IHC) before scanning the image. In most cases IHC is not done for all samples due to costs and the extra work, but in a research setting it's been sometimes performed to obtain a potentially more reliable ground truth. The pathologists' performance on the usual H&amp;E stained tissue (i.e. the type used in this challenge) can then be compared against the \"biochemical ground truth\". Training models on the IHC-based labels is sometimes called \"antibody-supervised learning\".",
      "votes": 4,
      "replies": [
        {
          "id": 818985,
          "postDate": "2020-04-24T09:08:22.970Z",
          "content": "<p>Wow, you answered both my questions today! Thank you very much for this thorough answer 😃 </p>",
          "rawMarkdown": "Wow, you answered both my questions today! Thank you very much for this thorough answer 😃 "
        }
      ]
    },
    {
      "id": 818414,
      "postDate": "2020-04-23T21:27:44.983Z",
      "content": "<p>I believe the ground truth is a consensus prediction from multiple pathologists. </p>\n\n<p>Edit: In the particular Nature paper, I found this:</p>\n\n<blockquote>\n  <p>For each slide, the reference standard was provided by one genitourinary specialist pathologist. To improve accuracy, the specialist reviewing each slide also had access to initial Gleason pattern percentage estimates and free-text comments from prior reviews of at least three general pathologists.</p>\n</blockquote>\n\n<p>So it seems like in that case it was just a single sub-specialist pathologist who had access to additional information.</p>",
      "rawMarkdown": "I believe the ground truth is a consensus prediction from multiple pathologists. \n\nEdit: In the particular Nature paper, I found this:\n&gt;For each slide, the reference standard was provided by one genitourinary specialist pathologist. To improve accuracy, the specialist reviewing each slide also had access to initial Gleason pattern percentage estimates and free-text comments from prior reviews of at least three general pathologists.\n\nSo it seems like in that case it was just a single sub-specialist pathologist who had access to additional information.",
      "votes": 1
    },
    {
      "id": 818422,
      "postDate": "2020-04-23T21:39:59.483Z",
      "content": "<p>see <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145619\">https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145619</a></p>",
      "rawMarkdown": "see https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145619",
      "replies": [
        {
          "id": 818425,
          "postDate": "2020-04-23T21:44:03.587Z",
          "content": "<p>I've already read that document, it doesn't really answer my question.</p>",
          "rawMarkdown": "I've already read that document, it doesn't really answer my question."
        },
        {
          "id": 818431,
          "postDate": "2020-04-23T21:51:51.853Z",
          "content": "<p>Got you, I read your question incorrectly</p>",
          "rawMarkdown": "Got you, I read your question incorrectly"
        }
      ]
    },
    {
      "id": 818340,
      "postDate": "2020-04-23T19:50:11.157Z",
      "content": "<p>My question does not pertain specifically to the competition dataset but in regards to the literature in general.</p>\n\n<p>Many of the published articles that show a deep learning system being used to predict Gleason scores compares their accuracy against the accuracy of pathologists. For exmaple, in <a href=\"https://www.nature.com/articles/s41746-019-0112-2.pdf\">this Nature paper by Google AI</a>, they say the accuracy of 29 pathologists was 0.61 while their model achieved 0.7. But how do they actually decide what the ground truth is? It must have been assigned by a pathologist, but who's to say that label is necessarily \"correct\" compared to the other pathologists? Is it decided by consensus like the dataset in this competition?</p>\n\n<p>Thanks for responses in advance, I hope I explained that clearly enough.</p>",
      "rawMarkdown": "My question does not pertain specifically to the competition dataset but in regards to the literature in general.\n\nMany of the published articles that show a deep learning system being used to predict Gleason scores compares their accuracy against the accuracy of pathologists. For exmaple, in [this Nature paper by Google AI](https://www.nature.com/articles/s41746-019-0112-2.pdf), they say the accuracy of 29 pathologists was 0.61 while their model achieved 0.7. But how do they actually decide what the ground truth is? It must have been assigned by a pathologist, but who's to say that label is necessarily \"correct\" compared to the other pathologists? Is it decided by consensus like the dataset in this competition?\n\nThanks for responses in advance, I hope I explained that clearly enough."
    }
  ],
  "comments": [
    {
      "id": 818814,
      "author_name": "Kimmo Kartasalo",
      "author_url": "",
      "post_date": "2020-04-24T06:17:15.587000",
      "content": "<p>Good question - and indeed a hot topic in the field. In most cases the ground truth is constructed in similar fashion to the test data of this challenge - based on the consensus of several pathologists. In some cases the assessment performed by a pathologist, who is specialized in the disease/tissue in question and/or has many years of experience is treated as ground truth, and the assessments by general pathologists are compared against the expert. This of course has pitfalls as even the most knowledgeable experts may simply have a bad day from time to time and there is variation between even highly experienced pathologists.</p>\n\n<p>In some cancers, there are proteins or other molecules whose presence in the tissue is recognized as a strong indication of cancer, and they can be picked up by applying a specific chemical stain (immunohistochemistry, IHC) before scanning the image. In most cases IHC is not done for all samples due to costs and the extra work, but in a research setting it's been sometimes performed to obtain a potentially more reliable ground truth. The pathologists' performance on the usual H&amp;E stained tissue (i.e. the type used in this challenge) can then be compared against the \"biochemical ground truth\". Training models on the IHC-based labels is sometimes called \"antibody-supervised learning\".</p>",
      "votes": 4,
      "replies": [
        {
          "id": 818985,
          "author_name": "Matt",
          "author_url": "",
          "post_date": "2020-04-24T09:08:22.970000",
          "content": "<p>Wow, you answered both my questions today! Thank you very much for this thorough answer 😃 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 818414,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2020-04-23T21:27:44.983000",
      "content": "<p>I believe the ground truth is a consensus prediction from multiple pathologists. </p>\n\n<p>Edit: In the particular Nature paper, I found this:</p>\n\n<blockquote>\n  <p>For each slide, the reference standard was provided by one genitourinary specialist pathologist. To improve accuracy, the specialist reviewing each slide also had access to initial Gleason pattern percentage estimates and free-text comments from prior reviews of at least three general pathologists.</p>\n</blockquote>\n\n<p>So it seems like in that case it was just a single sub-specialist pathologist who had access to additional information.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 818422,
      "author_name": "Rafi Hai",
      "author_url": "",
      "post_date": "2020-04-23T21:39:59.483000",
      "content": "<p>see <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145619\">https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145619</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 818425,
          "author_name": "Matt",
          "author_url": "",
          "post_date": "2020-04-23T21:44:03.587000",
          "content": "<p>I've already read that document, it doesn't really answer my question.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 818431,
          "author_name": "Rafi Hai",
          "author_url": "",
          "post_date": "2020-04-23T21:51:51.853000",
          "content": "<p>Got you, I read your question incorrectly</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "818814": "Good question - and indeed a hot topic in the field. In most cases the ground truth is constructed in similar fashion to the test data of this challenge - based on the consensus of several pathologists. In some cases the assessment performed by a pathologist, who is specialized in the disease/tissue in question and/or has many years of experience is treated as ground truth, and the assessments by general pathologists are compared against the expert. This of course has pitfalls as even the most knowledgeable experts may simply have a bad day from time to time and there is variation between even highly experienced pathologists.\n\nIn some cancers, there are proteins or other molecules whose presence in the tissue is recognized as a strong indication of cancer, and they can be picked up by applying a specific chemical stain (immunohistochemistry, IHC) before scanning the image. In most cases IHC is not done for all samples due to costs and the extra work, but in a research setting it's been sometimes performed to obtain a potentially more reliable ground truth. The pathologists' performance on the usual H&amp;E stained tissue (i.e. the type used in this challenge) can then be compared against the \"biochemical ground truth\". Training models on the IHC-based labels is sometimes called \"antibody-supervised learning\".",
    "818414": "I believe the ground truth is a consensus prediction from multiple pathologists. \n\nEdit: In the particular Nature paper, I found this:\n&gt;For each slide, the reference standard was provided by one genitourinary specialist pathologist. To improve accuracy, the specialist reviewing each slide also had access to initial Gleason pattern percentage estimates and free-text comments from prior reviews of at least three general pathologists.\n\nSo it seems like in that case it was just a single sub-specialist pathologist who had access to additional information.",
    "818422": "see https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145619",
    "818340": "My question does not pertain specifically to the competition dataset but in regards to the literature in general.\n\nMany of the published articles that show a deep learning system being used to predict Gleason scores compares their accuracy against the accuracy of pathologists. For exmaple, in [this Nature paper by Google AI](https://www.nature.com/articles/s41746-019-0112-2.pdf), they say the accuracy of 29 pathologists was 0.61 while their model achieved 0.7. But how do they actually decide what the ground truth is? It must have been assigned by a pathologist, but who's to say that label is necessarily \"correct\" compared to the other pathologists? Is it decided by consensus like the dataset in this competition?\n\nThanks for responses in advance, I hope I explained that clearly enough."
  }
}