{
  "id": 147322,
  "title": "1.5M patch's coordinates with targets",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/147322",
  "author_name": "SM",
  "post_date": "2020-04-30T08:19:30.907000",
  "votes": 15,
  "comment_count": 7,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/sermakarevich/15m-patchs-coordinates-with-targets/download\">Here</a> you can find 1.5M coordinates of (hopefully) non-empty patches of size (512, 512) with their mask labels. \n<code>\n(10752, 22016)  [0 1 2] 43111f9c3a9aca9f3b1daf09f392f3c8.tiff\n</code>\n- If there is no mask for a WSI - <code>target</code> col is empty. \n- If there are no/few patches for a WSI - there is something wrong with either WSI itself (some are empty like 3790f55cad63053e956fb73027179707) or with mask (we cant read_region from first level of some masks or some masks are empty). \n- an example of code to generate the file is <a href=\"https://www.kaggle.com/sermakarevich/prepare-patches-coordinates-and-targets\">here</a>. Some modifications were made for handling errors, checking if a WSI was already processed and saving results for every patch instead of collecting all results. </p>\n\n<p>Hope this helps somebody. Good luck. </p>",
  "messages": [
    {
      "id": 827329,
      "postDate": "2020-04-30T08:19:30.907Z",
      "content": "<p><a href=\"https://www.kaggle.com/sermakarevich/15m-patchs-coordinates-with-targets/download\">Here</a> you can find 1.5M coordinates of (hopefully) non-empty patches of size (512, 512) with their mask labels. \n<code>\n(10752, 22016)  [0 1 2] 43111f9c3a9aca9f3b1daf09f392f3c8.tiff\n</code>\n- If there is no mask for a WSI - <code>target</code> col is empty. \n- If there are no/few patches for a WSI - there is something wrong with either WSI itself (some are empty like 3790f55cad63053e956fb73027179707) or with mask (we cant read_region from first level of some masks or some masks are empty). \n- an example of code to generate the file is <a href=\"https://www.kaggle.com/sermakarevich/prepare-patches-coordinates-and-targets\">here</a>. Some modifications were made for handling errors, checking if a WSI was already processed and saving results for every patch instead of collecting all results. </p>\n\n<p>Hope this helps somebody. Good luck. </p>",
      "rawMarkdown": "[Here](https://www.kaggle.com/sermakarevich/15m-patchs-coordinates-with-targets/download) you can find 1.5M coordinates of (hopefully) non-empty patches of size (512, 512) with their mask labels. \n```\n(10752, 22016)\t[0 1 2]\t43111f9c3a9aca9f3b1daf09f392f3c8.tiff\n```\n- If there is no mask for a WSI - `target` col is empty. \n- If there are no/few patches for a WSI - there is something wrong with either WSI itself (some are empty like 3790f55cad63053e956fb73027179707) or with mask (we cant read_region from first level of some masks or some masks are empty). \n- an example of code to generate the file is [here](https://www.kaggle.com/sermakarevich/prepare-patches-coordinates-and-targets). Some modifications were made for handling errors, checking if a WSI was already processed and saving results for every patch instead of collecting all results. \n\nHope this helps somebody. Good luck. ",
      "votes": 15
    },
    {
      "id": 829104,
      "postDate": "2020-05-01T13:59:05.777Z",
      "content": "<p>Nice work done there <a href=\"/sermakarevich\">@sermakarevich</a> .\nHow did you encode the square position in that coordinate ? The coordinates (for example (10752, 22016)  represents the TL corner of the 512, 512 image ?\nAlso, can you give more information about how did you attached the label to a region ?</p>\n\n<p>Thank you and keep up the good work</p>",
      "rawMarkdown": "Nice work done there @sermakarevich .\nHow did you encode the square position in that coordinate ? The coordinates (for example (10752, 22016)  represents the TL corner of the 512, 512 image ?\nAlso, can you give more information about how did you attached the label to a region ?\n\nThank you and keep up the good work\n",
      "replies": [
        {
          "id": 829141,
          "postDate": "2020-05-01T14:38:35.180Z",
          "content": "<p>coordinates work like this:\n<code>\nwsi.read_region((10752, 22016), 0, (512, 512))\n</code>\ntargets for the same area:\n<code>\nunique_mask_values = np.unique(mask_array)\n</code></p>",
          "rawMarkdown": "coordinates work like this:\n```\nwsi.read_region((10752, 22016), 0, (512, 512))\n```\ntargets for the same area:\n```\nunique_mask_values = np.unique(mask_array)\n```",
          "votes": 2
        },
        {
          "id": 829374,
          "postDate": "2020-05-01T18:32:27.807Z",
          "content": "<p>Thank you for the quick reply <a href=\"/sermakarevich\">@sermakarevich</a> .\nI haven't yet used the segmentation mask but from what I see in the description the labels are different depending on the provider (karolinska vs radboud)</p>\n\n<p>Radboud: Prostate glands are individually labelled. Valid values are:\n0: background (non tissue) or unknown\n1: stroma (connective tissue, non-epithelium tissue)\n2: healthy (benign) epithelium\n3: cancerous epithelium (Gleason 3)\n4: cancerous epithelium (Gleason 4)\n5: cancerous epithelium (Gleason 5)</p>\n\n<p>Karolinska: Regions are labelled. Valid values are:\n0: background (non tissue) or unknown\n1: benign tissue (stroma and epithelium combined)\n2: cancerous tissue (stroma and epithelium combined)</p>\n\n<p>So, when trained, you have to check out the picture to what provider belongs and treat the labels accordantly.</p>\n\n<p>Another aspect, did you consider to eliminate as a label a category that is having too fewer pixels in that crop area ? let's consider a 512 x 512 area which is having category 5 just 2 pixels and category 3 the rest of them, it will be a little confusing for the model to extrapolate from that 2 pixels that the \ncorrect labels are 3 and 5 also</p>\n\n<p>Thank you</p>",
          "rawMarkdown": "Thank you for the quick reply @sermakarevich .\nI haven't yet used the segmentation mask but from what I see in the description the labels are different depending on the provider (karolinska vs radboud)\n\nRadboud: Prostate glands are individually labelled. Valid values are:\n0: background (non tissue) or unknown\n1: stroma (connective tissue, non-epithelium tissue)\n2: healthy (benign) epithelium\n3: cancerous epithelium (Gleason 3)\n4: cancerous epithelium (Gleason 4)\n5: cancerous epithelium (Gleason 5)\n\nKarolinska: Regions are labelled. Valid values are:\n0: background (non tissue) or unknown\n1: benign tissue (stroma and epithelium combined)\n2: cancerous tissue (stroma and epithelium combined)\n\nSo, when trained, you have to check out the picture to what provider belongs and treat the labels accordantly.\n\nAnother aspect, did you consider to eliminate as a label a category that is having too fewer pixels in that crop area ? let's consider a 512 x 512 area which is having category 5 just 2 pixels and category 3 the rest of them, it will be a little confusing for the model to extrapolate from that 2 pixels that the \ncorrect labels are 3 and 5 also\n\nThank you"
        },
        {
          "id": 829410,
          "postDate": "2020-05-01T19:00:14.700Z",
          "content": "<p>Correct, the targets have to be further processed before training. Also we need to deal with patches without targets. </p>\n\n<p>Maybe we can consider joining scores 1 and 2 of <code>radboud</code> masks into score 1. Also we can try to assign both gleason_scores to every patch with score 2 for karolinska masks. When there is only one gleason pattern, like 3+3, we would be mostly correct. If there are multiple patterns - we would be partially correct. Having combo of correct + mostly correct + partially correct should be enough to distinguish between patches. However as it was mentioned by organizers: these masks are weakly supervised and they are not that perfect. For example this yellow area is labeled as 0 in mask image:  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F265850%2Fc83501740aac17bf1b9da3fb63b391d9%2Fmask_example.png?generation=1588359610832291&amp;alt=media\" alt=\"\"></p>\n\n<p>If you take a look into code example you can see there are 2 thresholds: <code>white_area_score</code> and <code>max_white_area_mean</code> which I used to drop mostly empty/grey patches. </p>",
          "rawMarkdown": "Correct, the targets have to be further processed before training. Also we need to deal with patches without targets. \n\nMaybe we can consider joining scores 1 and 2 of `radboud` masks into score 1. Also we can try to assign both gleason_scores to every patch with score 2 for karolinska masks. When there is only one gleason pattern, like 3+3, we would be mostly correct. If there are multiple patterns - we would be partially correct. Having combo of correct + mostly correct + partially correct should be enough to distinguish between patches. However as it was mentioned by organizers: these masks are weakly supervised and they are not that perfect. For example this yellow area is labeled as 0 in mask image:  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F265850%2Fc83501740aac17bf1b9da3fb63b391d9%2Fmask_example.png?generation=1588359610832291&amp;alt=media)\n\n\nIf you take a look into code example you can see there are 2 thresholds: `white_area_score` and `max_white_area_mean ` which I used to drop mostly empty/grey patches. \n\n",
          "votes": 3
        }
      ]
    },
    {
      "id": 827352,
      "postDate": "2020-04-30T08:35:09.337Z",
      "content": "<p>Thanks for sharing! I created something similar but learning from such a large dataset can be difficult with limited resources.</p>",
      "rawMarkdown": "Thanks for sharing! I created something similar but learning from such a large dataset can be difficult with limited resources."
    },
    {
      "id": 828126,
      "postDate": "2020-04-30T19:21:48.173Z",
      "content": "<p>Thanks for sharing :)</p>",
      "rawMarkdown": "Thanks for sharing :)"
    },
    {
      "id": 827340,
      "postDate": "2020-04-30T08:27:49.273Z",
      "content": "<p>Thanks for sharing :)</p>",
      "rawMarkdown": "Thanks for sharing :)"
    }
  ],
  "comments": [
    {
      "id": 829104,
      "author_name": "Vlad Vaduva",
      "author_url": "",
      "post_date": "2020-05-01T13:59:05.777000",
      "content": "<p>Nice work done there <a href=\"/sermakarevich\">@sermakarevich</a> .\nHow did you encode the square position in that coordinate ? The coordinates (for example (10752, 22016)  represents the TL corner of the 512, 512 image ?\nAlso, can you give more information about how did you attached the label to a region ?</p>\n\n<p>Thank you and keep up the good work</p>",
      "votes": 0,
      "replies": [
        {
          "id": 829141,
          "author_name": "SM",
          "author_url": "",
          "post_date": "2020-05-01T14:38:35.180000",
          "content": "<p>coordinates work like this:\n<code>\nwsi.read_region((10752, 22016), 0, (512, 512))\n</code>\ntargets for the same area:\n<code>\nunique_mask_values = np.unique(mask_array)\n</code></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 829374,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-05-01T18:32:27.807000",
          "content": "<p>Thank you for the quick reply <a href=\"/sermakarevich\">@sermakarevich</a> .\nI haven't yet used the segmentation mask but from what I see in the description the labels are different depending on the provider (karolinska vs radboud)</p>\n\n<p>Radboud: Prostate glands are individually labelled. Valid values are:\n0: background (non tissue) or unknown\n1: stroma (connective tissue, non-epithelium tissue)\n2: healthy (benign) epithelium\n3: cancerous epithelium (Gleason 3)\n4: cancerous epithelium (Gleason 4)\n5: cancerous epithelium (Gleason 5)</p>\n\n<p>Karolinska: Regions are labelled. Valid values are:\n0: background (non tissue) or unknown\n1: benign tissue (stroma and epithelium combined)\n2: cancerous tissue (stroma and epithelium combined)</p>\n\n<p>So, when trained, you have to check out the picture to what provider belongs and treat the labels accordantly.</p>\n\n<p>Another aspect, did you consider to eliminate as a label a category that is having too fewer pixels in that crop area ? let's consider a 512 x 512 area which is having category 5 just 2 pixels and category 3 the rest of them, it will be a little confusing for the model to extrapolate from that 2 pixels that the \ncorrect labels are 3 and 5 also</p>\n\n<p>Thank you</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 829410,
          "author_name": "SM",
          "author_url": "",
          "post_date": "2020-05-01T19:00:14.700000",
          "content": "<p>Correct, the targets have to be further processed before training. Also we need to deal with patches without targets. </p>\n\n<p>Maybe we can consider joining scores 1 and 2 of <code>radboud</code> masks into score 1. Also we can try to assign both gleason_scores to every patch with score 2 for karolinska masks. When there is only one gleason pattern, like 3+3, we would be mostly correct. If there are multiple patterns - we would be partially correct. Having combo of correct + mostly correct + partially correct should be enough to distinguish between patches. However as it was mentioned by organizers: these masks are weakly supervised and they are not that perfect. For example this yellow area is labeled as 0 in mask image:  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F265850%2Fc83501740aac17bf1b9da3fb63b391d9%2Fmask_example.png?generation=1588359610832291&amp;alt=media\" alt=\"\"></p>\n\n<p>If you take a look into code example you can see there are 2 thresholds: <code>white_area_score</code> and <code>max_white_area_mean</code> which I used to drop mostly empty/grey patches. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 827352,
      "author_name": "Matt",
      "author_url": "",
      "post_date": "2020-04-30T08:35:09.337000",
      "content": "<p>Thanks for sharing! I created something similar but learning from such a large dataset can be difficult with limited resources.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 828126,
      "author_name": "Rasheeq Zaman",
      "author_url": "",
      "post_date": "2020-04-30T19:21:48.173000",
      "content": "<p>Thanks for sharing :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 827340,
      "author_name": "Alberto Maria Falletta",
      "author_url": "",
      "post_date": "2020-04-30T08:27:49.273000",
      "content": "<p>Thanks for sharing :)</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "827329": "[Here](https://www.kaggle.com/sermakarevich/15m-patchs-coordinates-with-targets/download) you can find 1.5M coordinates of (hopefully) non-empty patches of size (512, 512) with their mask labels. \n```\n(10752, 22016)\t[0 1 2]\t43111f9c3a9aca9f3b1daf09f392f3c8.tiff\n```\n- If there is no mask for a WSI - `target` col is empty. \n- If there are no/few patches for a WSI - there is something wrong with either WSI itself (some are empty like 3790f55cad63053e956fb73027179707) or with mask (we cant read_region from first level of some masks or some masks are empty). \n- an example of code to generate the file is [here](https://www.kaggle.com/sermakarevich/prepare-patches-coordinates-and-targets). Some modifications were made for handling errors, checking if a WSI was already processed and saving results for every patch instead of collecting all results. \n\nHope this helps somebody. Good luck. ",
    "829104": "Nice work done there @sermakarevich .\nHow did you encode the square position in that coordinate ? The coordinates (for example (10752, 22016)  represents the TL corner of the 512, 512 image ?\nAlso, can you give more information about how did you attached the label to a region ?\n\nThank you and keep up the good work\n",
    "827352": "Thanks for sharing! I created something similar but learning from such a large dataset can be difficult with limited resources.",
    "828126": "Thanks for sharing :)",
    "827340": "Thanks for sharing :)"
  }
}