{
  "id": 146297,
  "title": "Question about mask unique values and gleason score",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/146297",
  "author_name": "shihyung",
  "post_date": "2020-04-26T15:14:20.001000",
  "votes": 10,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi all,\n   I have a question about the mask unique values and gleason score:</p>\n\n<p>According to discussion, like:\n<a href=\"https://www.kaggle.com/rohitsingh9990/panda-eda-better-visualization-simple-baseline/comments\">https://www.kaggle.com/rohitsingh9990/panda-eda-better-visualization-simple-baseline/comments</a>\n'''\nLoading label masks\nRadboudumc: Prostate glands are individually labelled. Valid values are:\n0: background (non tissue) or unknown\n1: stroma (connective tissue, non-epithelium tissue)\n2: healthy (benign) epithelium\"\n3: cancerous epithelium (Gleason 3)\n4: cancerous epithelium (Gleason 4)\n5: cancerous epithelium (Gleason 5)\n'''\n, which listed how Radbound masks are labeled.</p>\n\n<p>But in some figures, it seems that mask values can't be related to gleason score.</p>\n\n<p>example 1:</p>\n\n<p>train.csv:   0068d4c7529e34fd4c9da863ce01a161, radboud, 3, 4+3</p>\n\n<p>mask file: 0068d4c7529e34fd4c9da863ce01a161_mask.tiff</p>\n\n<p>image shape:  (10496, 6912, 3)</p>\n\n<p>unique_value:  [0 1 2 3]</p>\n\n<p>unique_count:  [61748560   10325036   474116   640]</p>\n\n<p>in this example, gleason score has '4', but mask value doesn't </p>\n\n<p>example 2:</p>\n\n<p>train.csv:  00d8a8c04886379e266406fdeff81c45, radboud, 5, 4+5</p>\n\n<p>mask file:  00d8a8c04886379e266406fdeff81c45_mask.tiff</p>\n\n<p>image shape:  (24064, 12288, 3)</p>\n\n<p>unique_value:  [0 1 2]</p>\n\n<p>unique_count:  [274690448  19930008   1077976]</p>\n\n<p>in this example, gleason score has '4' and '5', but mask value doesn't </p>\n\n<p>Some Radbound masks have this weird result, but some don't.\nIs it mask erroneously labeled or just my misunderstanding?</p>\n\n<p>Thanks for your help.</p>",
  "messages": [
    {
      "id": 822003,
      "postDate": "2020-04-26T15:14:20.003Z",
      "content": "<p>Hi all,\n   I have a question about the mask unique values and gleason score:</p>\n\n<p>According to discussion, like:\n<a href=\"https://www.kaggle.com/rohitsingh9990/panda-eda-better-visualization-simple-baseline/comments\">https://www.kaggle.com/rohitsingh9990/panda-eda-better-visualization-simple-baseline/comments</a>\n'''\nLoading label masks\nRadboudumc: Prostate glands are individually labelled. Valid values are:\n0: background (non tissue) or unknown\n1: stroma (connective tissue, non-epithelium tissue)\n2: healthy (benign) epithelium\"\n3: cancerous epithelium (Gleason 3)\n4: cancerous epithelium (Gleason 4)\n5: cancerous epithelium (Gleason 5)\n'''\n, which listed how Radbound masks are labeled.</p>\n\n<p>But in some figures, it seems that mask values can't be related to gleason score.</p>\n\n<p>example 1:</p>\n\n<p>train.csv:   0068d4c7529e34fd4c9da863ce01a161, radboud, 3, 4+3</p>\n\n<p>mask file: 0068d4c7529e34fd4c9da863ce01a161_mask.tiff</p>\n\n<p>image shape:  (10496, 6912, 3)</p>\n\n<p>unique_value:  [0 1 2 3]</p>\n\n<p>unique_count:  [61748560   10325036   474116   640]</p>\n\n<p>in this example, gleason score has '4', but mask value doesn't </p>\n\n<p>example 2:</p>\n\n<p>train.csv:  00d8a8c04886379e266406fdeff81c45, radboud, 5, 4+5</p>\n\n<p>mask file:  00d8a8c04886379e266406fdeff81c45_mask.tiff</p>\n\n<p>image shape:  (24064, 12288, 3)</p>\n\n<p>unique_value:  [0 1 2]</p>\n\n<p>unique_count:  [274690448  19930008   1077976]</p>\n\n<p>in this example, gleason score has '4' and '5', but mask value doesn't </p>\n\n<p>Some Radbound masks have this weird result, but some don't.\nIs it mask erroneously labeled or just my misunderstanding?</p>\n\n<p>Thanks for your help.</p>",
      "rawMarkdown": "Hi all,\n   I have a question about the mask unique values and gleason score:\n\nAccording to discussion, like:\nhttps://www.kaggle.com/rohitsingh9990/panda-eda-better-visualization-simple-baseline/comments\n'''\nLoading label masks\nRadboudumc: Prostate glands are individually labelled. Valid values are:\n0: background (non tissue) or unknown\n1: stroma (connective tissue, non-epithelium tissue)\n2: healthy (benign) epithelium\"\n3: cancerous epithelium (Gleason 3)\n4: cancerous epithelium (Gleason 4)\n5: cancerous epithelium (Gleason 5)\n'''\n, which listed how Radbound masks are labeled.\n\n\nBut in some figures, it seems that mask values can't be related to gleason score.\n\nexample 1:\n\ntrain.csv:   0068d4c7529e34fd4c9da863ce01a161, radboud, 3, 4+3\n\nmask file: 0068d4c7529e34fd4c9da863ce01a161_mask.tiff\n\nimage shape:  (10496, 6912, 3)\n\nunique_value:  [0 1 2 3]\n\nunique_count:  [61748560   10325036   474116   640]\n\nin this example, gleason score has '4', but mask value doesn't \n\n\nexample 2:\n\ntrain.csv:  00d8a8c04886379e266406fdeff81c45, radboud, 5, 4+5\n\nmask file:  00d8a8c04886379e266406fdeff81c45_mask.tiff\n\nimage shape:  (24064, 12288, 3)\n\nunique_value:  [0 1 2]\n\nunique_count:  [274690448  19930008   1077976]\n\nin this example, gleason score has '4' and '5', but mask value doesn't \n\n\nSome Radbound masks have this weird result, but some don't.\nIs it mask erroneously labeled or just my misunderstanding?\n\nThanks for your help.",
      "votes": 10
    },
    {
      "id": 822016,
      "postDate": "2020-04-26T15:25:30.767Z",
      "content": "<p>Here are some other examples:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F1ff90338738903468cac86890ab56d49%2F1.png?generation=1587977630564316&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F81ecf583bda3f291077ff5caf98ef779%2F2.png?generation=1587977647758045&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2Fd328fef84aeacec3294785719248f551%2F3.png?generation=1587977671542038&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2Fbb352a12d54fd98ca4c7b41c8f4ba631%2F5.png?generation=1587977735495649&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F20c3bb7381c60a43ce9d3aafb52eed61%2F4.png?generation=1587977695343047&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2Fca2574890d1cb02d9de3c0db3fd7f975%2F6.png?generation=1587977760361512&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F091c4925ab2215783ccc6167d855211e%2F7.png?generation=1587977776229239&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Here are some other examples:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F1ff90338738903468cac86890ab56d49%2F1.png?generation=1587977630564316&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F81ecf583bda3f291077ff5caf98ef779%2F2.png?generation=1587977647758045&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2Fd328fef84aeacec3294785719248f551%2F3.png?generation=1587977671542038&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2Fbb352a12d54fd98ca4c7b41c8f4ba631%2F5.png?generation=1587977735495649&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F20c3bb7381c60a43ce9d3aafb52eed61%2F4.png?generation=1587977695343047&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2Fca2574890d1cb02d9de3c0db3fd7f975%2F6.png?generation=1587977760361512&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F091c4925ab2215783ccc6167d855211e%2F7.png?generation=1587977776229239&amp;alt=media)\n",
      "votes": 5,
      "replies": [
        {
          "id": 887232,
          "postDate": "2020-06-15T15:02:13.270Z",
          "content": "<p>I wish I had seen that post earlier. Thanks for the hard work on these tables !</p>",
          "rawMarkdown": "I wish I had seen that post earlier. Thanks for the hard work on these tables !"
        }
      ]
    },
    {
      "id": 824165,
      "postDate": "2020-04-28T07:45:43.363Z",
      "content": "<p>Good analysis! The labels of the biopsies in the training set were determined based on the patient's report. The masks were generated by a set of (deep learning) algorithms. In both labels there can be noise due to this semi-automatic method. </p>\n\n<p>The biopsy level label could for example be wrong due to an error by a pathologist, by an error in our data collection or because the patient report was incompletely described. For example, the first case in your list b9402571270ab62c1e0a690fd33fc9e6 does not seem to contain cancer so the mask is correct. (Note: I'm not a pathologist so this my best guess.) This slide does contain colon tissue (foreign tissue). It could be the case that the pathologist indicated 3+3 for the whole glass slide, while in fact some tissue parts did not contain cancer. This would result in an incorrect label in our set.</p>\n\n<p>The masks can also be wrong, just because our automated method is not perfect. Especially, Gleason 5 can be hard to detect. Sometimes there is also overgrading (benign or foreign tissue marked as cancer).</p>\n\n<p>If you want to know more, I suggest to read <a href=\"https://arxiv.org/abs/1907.07980\">our paper</a> or the <a href=\"https://www.wouterbulten.nl/blog/tech/automated-gleason-grading-deep-learning/#labeling\">summary on my blog</a> on the labeling method. </p>\n\n<p>Good luck!</p>",
      "rawMarkdown": "Good analysis! The labels of the biopsies in the training set were determined based on the patient's report. The masks were generated by a set of (deep learning) algorithms. In both labels there can be noise due to this semi-automatic method. \n\nThe biopsy level label could for example be wrong due to an error by a pathologist, by an error in our data collection or because the patient report was incompletely described. For example, the first case in your list b9402571270ab62c1e0a690fd33fc9e6 does not seem to contain cancer so the mask is correct. (Note: I'm not a pathologist so this my best guess.) This slide does contain colon tissue (foreign tissue). It could be the case that the pathologist indicated 3+3 for the whole glass slide, while in fact some tissue parts did not contain cancer. This would result in an incorrect label in our set.\n\nThe masks can also be wrong, just because our automated method is not perfect. Especially, Gleason 5 can be hard to detect. Sometimes there is also overgrading (benign or foreign tissue marked as cancer).\n\nIf you want to know more, I suggest to read [our paper](https://arxiv.org/abs/1907.07980) or the [summary on my blog](https://www.wouterbulten.nl/blog/tech/automated-gleason-grading-deep-learning/#labeling) on the labeling method. \n\nGood luck!",
      "votes": 2,
      "replies": [
        {
          "id": 826979,
          "postDate": "2020-04-30T02:23:48.627Z",
          "content": "<p>Thanks.\nSo it means that the masks are automatically generated by other AI/ML programs.\nBut how about the mask differences between two data_provider?</p>\n\n<p>For example, in the following figure, \nmasks by \"Karolinska\" are labeled continuously at the whole region,\nbut masks by \"Radbound\" are labeled discretely at each pixel.\nAre they both generated by the same AI/ML program?\nthx.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F3a4d13ed14da1c17843d8cfcd413fc46%2F2020-04-30%2010.14.05.png?generation=1588213162601914&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Thanks.\nSo it means that the masks are automatically generated by other AI/ML programs.\nBut how about the mask differences between two data_provider?\n\nFor example, in the following figure, \nmasks by \"Karolinska\" are labeled continuously at the whole region,\nbut masks by \"Radbound\" are labeled discretely at each pixel.\nAre they both generated by the same AI/ML program?\nthx.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F3a4d13ed14da1c17843d8cfcd413fc46%2F2020-04-30%2010.14.05.png?generation=1588213162601914&amp;alt=media)\n"
        },
        {
          "id": 827218,
          "postDate": "2020-04-30T06:25:43.760Z",
          "content": "<p>Hi, Wouter meant the Radboud masks in his reply. The Karolinska masks are not generated based on AI predictions but on rough annotations by the pathologist. In that sense, they may be more reliable, but as you can see they only highlight approximate regions with cancer (with no info on the grade).</p>",
          "rawMarkdown": "Hi, Wouter meant the Radboud masks in his reply. The Karolinska masks are not generated based on AI predictions but on rough annotations by the pathologist. In that sense, they may be more reliable, but as you can see they only highlight approximate regions with cancer (with no info on the grade).",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 822016,
      "author_name": "shihyung",
      "author_url": "",
      "post_date": "2020-04-26T15:25:30.767000",
      "content": "<p>Here are some other examples:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F1ff90338738903468cac86890ab56d49%2F1.png?generation=1587977630564316&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F81ecf583bda3f291077ff5caf98ef779%2F2.png?generation=1587977647758045&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2Fd328fef84aeacec3294785719248f551%2F3.png?generation=1587977671542038&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2Fbb352a12d54fd98ca4c7b41c8f4ba631%2F5.png?generation=1587977735495649&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F20c3bb7381c60a43ce9d3aafb52eed61%2F4.png?generation=1587977695343047&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2Fca2574890d1cb02d9de3c0db3fd7f975%2F6.png?generation=1587977760361512&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F091c4925ab2215783ccc6167d855211e%2F7.png?generation=1587977776229239&amp;alt=media\" alt=\"\"></p>",
      "votes": 5,
      "replies": [
        {
          "id": 887232,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-06-15T15:02:13.270000",
          "content": "<p>I wish I had seen that post earlier. Thanks for the hard work on these tables !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 824165,
      "author_name": "Wouter Bulten",
      "author_url": "",
      "post_date": "2020-04-28T07:45:43.363000",
      "content": "<p>Good analysis! The labels of the biopsies in the training set were determined based on the patient's report. The masks were generated by a set of (deep learning) algorithms. In both labels there can be noise due to this semi-automatic method. </p>\n\n<p>The biopsy level label could for example be wrong due to an error by a pathologist, by an error in our data collection or because the patient report was incompletely described. For example, the first case in your list b9402571270ab62c1e0a690fd33fc9e6 does not seem to contain cancer so the mask is correct. (Note: I'm not a pathologist so this my best guess.) This slide does contain colon tissue (foreign tissue). It could be the case that the pathologist indicated 3+3 for the whole glass slide, while in fact some tissue parts did not contain cancer. This would result in an incorrect label in our set.</p>\n\n<p>The masks can also be wrong, just because our automated method is not perfect. Especially, Gleason 5 can be hard to detect. Sometimes there is also overgrading (benign or foreign tissue marked as cancer).</p>\n\n<p>If you want to know more, I suggest to read <a href=\"https://arxiv.org/abs/1907.07980\">our paper</a> or the <a href=\"https://www.wouterbulten.nl/blog/tech/automated-gleason-grading-deep-learning/#labeling\">summary on my blog</a> on the labeling method. </p>\n\n<p>Good luck!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 826979,
          "author_name": "shihyung",
          "author_url": "",
          "post_date": "2020-04-30T02:23:48.627000",
          "content": "<p>Thanks.\nSo it means that the masks are automatically generated by other AI/ML programs.\nBut how about the mask differences between two data_provider?</p>\n\n<p>For example, in the following figure, \nmasks by \"Karolinska\" are labeled continuously at the whole region,\nbut masks by \"Radbound\" are labeled discretely at each pixel.\nAre they both generated by the same AI/ML program?\nthx.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F3a4d13ed14da1c17843d8cfcd413fc46%2F2020-04-30%2010.14.05.png?generation=1588213162601914&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 827218,
          "author_name": "Kimmo Kartasalo",
          "author_url": "",
          "post_date": "2020-04-30T06:25:43.760000",
          "content": "<p>Hi, Wouter meant the Radboud masks in his reply. The Karolinska masks are not generated based on AI predictions but on rough annotations by the pathologist. In that sense, they may be more reliable, but as you can see they only highlight approximate regions with cancer (with no info on the grade).</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "822003": "Hi all,\n   I have a question about the mask unique values and gleason score:\n\nAccording to discussion, like:\nhttps://www.kaggle.com/rohitsingh9990/panda-eda-better-visualization-simple-baseline/comments\n'''\nLoading label masks\nRadboudumc: Prostate glands are individually labelled. Valid values are:\n0: background (non tissue) or unknown\n1: stroma (connective tissue, non-epithelium tissue)\n2: healthy (benign) epithelium\"\n3: cancerous epithelium (Gleason 3)\n4: cancerous epithelium (Gleason 4)\n5: cancerous epithelium (Gleason 5)\n'''\n, which listed how Radbound masks are labeled.\n\n\nBut in some figures, it seems that mask values can't be related to gleason score.\n\nexample 1:\n\ntrain.csv:   0068d4c7529e34fd4c9da863ce01a161, radboud, 3, 4+3\n\nmask file: 0068d4c7529e34fd4c9da863ce01a161_mask.tiff\n\nimage shape:  (10496, 6912, 3)\n\nunique_value:  [0 1 2 3]\n\nunique_count:  [61748560   10325036   474116   640]\n\nin this example, gleason score has '4', but mask value doesn't \n\n\nexample 2:\n\ntrain.csv:  00d8a8c04886379e266406fdeff81c45, radboud, 5, 4+5\n\nmask file:  00d8a8c04886379e266406fdeff81c45_mask.tiff\n\nimage shape:  (24064, 12288, 3)\n\nunique_value:  [0 1 2]\n\nunique_count:  [274690448  19930008   1077976]\n\nin this example, gleason score has '4' and '5', but mask value doesn't \n\n\nSome Radbound masks have this weird result, but some don't.\nIs it mask erroneously labeled or just my misunderstanding?\n\nThanks for your help.",
    "822016": "Here are some other examples:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F1ff90338738903468cac86890ab56d49%2F1.png?generation=1587977630564316&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F81ecf583bda3f291077ff5caf98ef779%2F2.png?generation=1587977647758045&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2Fd328fef84aeacec3294785719248f551%2F3.png?generation=1587977671542038&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2Fbb352a12d54fd98ca4c7b41c8f4ba631%2F5.png?generation=1587977735495649&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F20c3bb7381c60a43ce9d3aafb52eed61%2F4.png?generation=1587977695343047&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2Fca2574890d1cb02d9de3c0db3fd7f975%2F6.png?generation=1587977760361512&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1935020%2F091c4925ab2215783ccc6167d855211e%2F7.png?generation=1587977776229239&amp;alt=media)\n",
    "824165": "Good analysis! The labels of the biopsies in the training set were determined based on the patient's report. The masks were generated by a set of (deep learning) algorithms. In both labels there can be noise due to this semi-automatic method. \n\nThe biopsy level label could for example be wrong due to an error by a pathologist, by an error in our data collection or because the patient report was incompletely described. For example, the first case in your list b9402571270ab62c1e0a690fd33fc9e6 does not seem to contain cancer so the mask is correct. (Note: I'm not a pathologist so this my best guess.) This slide does contain colon tissue (foreign tissue). It could be the case that the pathologist indicated 3+3 for the whole glass slide, while in fact some tissue parts did not contain cancer. This would result in an incorrect label in our set.\n\nThe masks can also be wrong, just because our automated method is not perfect. Especially, Gleason 5 can be hard to detect. Sometimes there is also overgrading (benign or foreign tissue marked as cancer).\n\nIf you want to know more, I suggest to read [our paper](https://arxiv.org/abs/1907.07980) or the [summary on my blog](https://www.wouterbulten.nl/blog/tech/automated-gleason-grading-deep-learning/#labeling) on the labeling method. \n\nGood luck!"
  }
}