{
  "id": 163576,
  "title": "Clinical-grade computational pathology for Prostate Cancer using MIL (Multi-instance Learning)",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/163576",
  "author_name": "SumitJha",
  "post_date": "2020-07-02T15:26:35.260000",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I like to share one of the best approaches to solve this problem. I have started with this approach but due to my engagement with other work, I couldn't complete it :(</p>\n\n<p>I think this approch  might help others as this method is published in <strong>Nature  by Paige AI</strong></p>\n\n<p><strong>Note</strong> I am an engineer, not a pathologist but whatever I share here is based on my interaction with the senior pathologists.</p>\n\n<p>In clinical practices, Gleason Score is a must-have solution so any algorithm which gives ISUP w/o Gleason Score is not good for pathology.</p>\n\n<p><strong>The issues at hand</strong>-\n- The annotation is very noisy\n- GT Mask of the WSI doesn't have detailed Gleason score information</p>\n\n<p>What  MIL Nature paper says-</p>\n\n<p>&gt; The development of decision support systems for pathology and their deployment in clinical practice have been hindered by the need for large manually annotated datasets. To overcome this problem, <strong>we present a multiple instance learning-based deep learning system that uses only the reported diagnoses as labels for training</strong>, thereby avoiding expensive and time-consuming\npixel-wise manual annotations. We evaluated this framework at scale on a dataset of 44,732 whole slide images from 15,187 patients without any form of data curation. Tests on prostate cancer, basal cell carcinoma and breast cancer metastases to axillary lymph nodes resulted in areas under the curve above 0.98 for all cancer types. Its clinical application would allow\npathologists to exclude 65–75% of slides while retaining 100% sensitivity. Our results show that this system has the ability to train accurate classification models at unprecedented scale, laying the foundation for the deployment of computational decision support systems in clinical practice.</p>\n\n<p><strong>My Proposed Approach</strong></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F505538%2F9db2256fdd1ccda30e8537469f3ceb29%2Fmil.PNG?generation=1593702600561260&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>MIL Algorithm</strong>\nThis works well with binary classifiers but not sure how it works for multi-class classification.  I have started with multi-class but realized the algorithm limits. So started the following paper and developed an approach for tumor/normal classification first.  After that, all normal slides will be Gealson Score -0 and for others w we need another classifier/CNN model to get Gleason-3/4/5.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F505538%2F9382debff24b94e2cd47655ee0bd92fb%2Fmilflow.PNG?generation=1593702948273717&amp;alt=media\" alt=\"\"></p>\n\n<p>For detailed -\nRead paper-\nNature paper - <a href=\"https://www.nature.com/articles/s41591-019-0508-1\">https://www.nature.com/articles/s41591-019-0508-1</a></p>\n\n<p>Github code-\n<a href=\"https://github.com/MSKCC-Computational-Pathology/MIL-nature-medicine-2019\">https://github.com/MSKCC-Computational-Pathology/MIL-nature-medicine-2019</a></p>\n\n<p>MyInference Kernel\n<a href=\"https://www.kaggle.com/sumitjha19/gleason-to-isup-score-resnet50\">My Inference Kernel</a>\n<a href=\"https://www.kaggle.com/sumitjha19/gleason-to-isup-score-resnet50\">https://www.kaggle.com/sumitjha19/gleason-to-isup-score-resnet50</a></p>\n\n<p><strong>My Experience</strong>\n- This approach needs Huge RAM and GPU capacity to run. \n- Kaggle kernel doesn't meet RAM requirements.</p>\n\n<p>I don't think we have so much time to train a model from scratch. But, I am open for collaboration on this approach.</p>",
  "messages": [
    {
      "id": 912576,
      "postDate": "2020-07-02T15:26:35.260Z",
      "content": "<p>I like to share one of the best approaches to solve this problem. I have started with this approach but due to my engagement with other work, I couldn't complete it :(</p>\n\n<p>I think this approch  might help others as this method is published in <strong>Nature  by Paige AI</strong></p>\n\n<p><strong>Note</strong> I am an engineer, not a pathologist but whatever I share here is based on my interaction with the senior pathologists.</p>\n\n<p>In clinical practices, Gleason Score is a must-have solution so any algorithm which gives ISUP w/o Gleason Score is not good for pathology.</p>\n\n<p><strong>The issues at hand</strong>-\n- The annotation is very noisy\n- GT Mask of the WSI doesn't have detailed Gleason score information</p>\n\n<p>What  MIL Nature paper says-</p>\n\n<p>&gt; The development of decision support systems for pathology and their deployment in clinical practice have been hindered by the need for large manually annotated datasets. To overcome this problem, <strong>we present a multiple instance learning-based deep learning system that uses only the reported diagnoses as labels for training</strong>, thereby avoiding expensive and time-consuming\npixel-wise manual annotations. We evaluated this framework at scale on a dataset of 44,732 whole slide images from 15,187 patients without any form of data curation. Tests on prostate cancer, basal cell carcinoma and breast cancer metastases to axillary lymph nodes resulted in areas under the curve above 0.98 for all cancer types. Its clinical application would allow\npathologists to exclude 65–75% of slides while retaining 100% sensitivity. Our results show that this system has the ability to train accurate classification models at unprecedented scale, laying the foundation for the deployment of computational decision support systems in clinical practice.</p>\n\n<p><strong>My Proposed Approach</strong></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F505538%2F9db2256fdd1ccda30e8537469f3ceb29%2Fmil.PNG?generation=1593702600561260&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>MIL Algorithm</strong>\nThis works well with binary classifiers but not sure how it works for multi-class classification.  I have started with multi-class but realized the algorithm limits. So started the following paper and developed an approach for tumor/normal classification first.  After that, all normal slides will be Gealson Score -0 and for others w we need another classifier/CNN model to get Gleason-3/4/5.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F505538%2F9382debff24b94e2cd47655ee0bd92fb%2Fmilflow.PNG?generation=1593702948273717&amp;alt=media\" alt=\"\"></p>\n\n<p>For detailed -\nRead paper-\nNature paper - <a href=\"https://www.nature.com/articles/s41591-019-0508-1\">https://www.nature.com/articles/s41591-019-0508-1</a></p>\n\n<p>Github code-\n<a href=\"https://github.com/MSKCC-Computational-Pathology/MIL-nature-medicine-2019\">https://github.com/MSKCC-Computational-Pathology/MIL-nature-medicine-2019</a></p>\n\n<p>MyInference Kernel\n<a href=\"https://www.kaggle.com/sumitjha19/gleason-to-isup-score-resnet50\">My Inference Kernel</a>\n<a href=\"https://www.kaggle.com/sumitjha19/gleason-to-isup-score-resnet50\">https://www.kaggle.com/sumitjha19/gleason-to-isup-score-resnet50</a></p>\n\n<p><strong>My Experience</strong>\n- This approach needs Huge RAM and GPU capacity to run. \n- Kaggle kernel doesn't meet RAM requirements.</p>\n\n<p>I don't think we have so much time to train a model from scratch. But, I am open for collaboration on this approach.</p>",
      "rawMarkdown": "I like to share one of the best approaches to solve this problem. I have started with this approach but due to my engagement with other work, I couldn't complete it :(\n\nI think this approch  might help others as this method is published in **Nature  by Paige AI**\n\n**Note** I am an engineer, not a pathologist but whatever I share here is based on my interaction with the senior pathologists.\n\nIn clinical practices, Gleason Score is a must-have solution so any algorithm which gives ISUP w/o Gleason Score is not good for pathology.\n\n**The issues at hand**-\n- The annotation is very noisy\n- GT Mask of the WSI doesn't have detailed Gleason score information\n\n\nWhat  MIL Nature paper says-\n\n&gt; The development of decision support systems for pathology and their deployment in clinical practice have been hindered by the need for large manually annotated datasets. To overcome this problem, **we present a multiple instance learning-based deep learning system that uses only the reported diagnoses as labels for training**, thereby avoiding expensive and time-consuming\npixel-wise manual annotations. We evaluated this framework at scale on a dataset of 44,732 whole slide images from 15,187 patients without any form of data curation. Tests on prostate cancer, basal cell carcinoma and breast cancer metastases to axillary lymph nodes resulted in areas under the curve above 0.98 for all cancer types. Its clinical application would allow\npathologists to exclude 65–75% of slides while retaining 100% sensitivity. Our results show that this system has the ability to train accurate classification models at unprecedented scale, laying the foundation for the deployment of computational decision support systems in clinical practice.\n\n **My Proposed Approach**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F505538%2F9db2256fdd1ccda30e8537469f3ceb29%2Fmil.PNG?generation=1593702600561260&amp;alt=media)\n\n\n\n**MIL Algorithm**\nThis works well with binary classifiers but not sure how it works for multi-class classification.  I have started with multi-class but realized the algorithm limits. So started the following paper and developed an approach for tumor/normal classification first.  After that, all normal slides will be Gealson Score -0 and for others w we need another classifier/CNN model to get Gleason-3/4/5.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F505538%2F9382debff24b94e2cd47655ee0bd92fb%2Fmilflow.PNG?generation=1593702948273717&amp;alt=media)\n\n\nFor detailed -\nRead paper-\nNature paper - https://www.nature.com/articles/s41591-019-0508-1\n\nGithub code-\nhttps://github.com/MSKCC-Computational-Pathology/MIL-nature-medicine-2019\n\nMyInference Kernel\n[My Inference Kernel](https://www.kaggle.com/sumitjha19/gleason-to-isup-score-resnet50)\nhttps://www.kaggle.com/sumitjha19/gleason-to-isup-score-resnet50\n\n**My Experience**\n- This approach needs Huge RAM and GPU capacity to run. \n- Kaggle kernel doesn't meet RAM requirements.\n\nI don't think we have so much time to train a model from scratch. But, I am open for collaboration on this approach.\n\n\n\n\n\n\n\n\n\n\n\n",
      "votes": 5
    },
    {
      "id": 912678,
      "postDate": "2020-07-02T16:35:53.203Z",
      "content": "<p>Very interesting but, as you say, this is a multi-class classification. And you must consider that Gleason score is just a part of the work, the submission it's about ISUP grade. I wish you succesful with your interesting approach.</p>",
      "rawMarkdown": "Very interesting but, as you say, this is a multi-class classification. And you must consider that Gleason score is just a part of the work, the submission it's about ISUP grade. I wish you succesful with your interesting approach.",
      "replies": [
        {
          "id": 913463,
          "postDate": "2020-07-03T08:24:26.243Z",
          "content": "<p>Yes, this is a multi-class classification.  I  am not thinking of developing a solution to win this challenge rather my thought process was how to help pathologists to use such a solution in day to day work.  That will be a greater value to the patient/pathologist/oncologist. There I came to know that just giving ISUP might not work as they need Gleason score too.</p>\n\n<p>The nature paper talks bout just tumor/normal c classification with the data set without pixel-level annotations. So, I am extending this idea to solve this problem with an extra model for classifying tumor into Gleason Score - _3/4/5. And from the Gleason score to ISUP is a simple rule.</p>\n\n<p>I don't have time to complete this so posted to the community if anyone can take it further.</p>",
          "rawMarkdown": "Yes, this is a multi-class classification.  I  am not thinking of developing a solution to win this challenge rather my thought process was how to help pathologists to use such a solution in day to day work.  That will be a greater value to the patient/pathologist/oncologist. There I came to know that just giving ISUP might not work as they need Gleason score too.\n\nThe nature paper talks bout just tumor/normal c classification with the data set without pixel-level annotations. So, I am extending this idea to solve this problem with an extra model for classifying tumor into Gleason Score - _3/4/5. And from the Gleason score to ISUP is a simple rule.\n\nI don't have time to complete this so posted to the community if anyone can take it further.\n",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 912678,
      "author_name": "Hiram Coria 🧬",
      "author_url": "",
      "post_date": "2020-07-02T16:35:53.203000",
      "content": "<p>Very interesting but, as you say, this is a multi-class classification. And you must consider that Gleason score is just a part of the work, the submission it's about ISUP grade. I wish you succesful with your interesting approach.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 913463,
          "author_name": "SumitJha",
          "author_url": "",
          "post_date": "2020-07-03T08:24:26.243000",
          "content": "<p>Yes, this is a multi-class classification.  I  am not thinking of developing a solution to win this challenge rather my thought process was how to help pathologists to use such a solution in day to day work.  That will be a greater value to the patient/pathologist/oncologist. There I came to know that just giving ISUP might not work as they need Gleason score too.</p>\n\n<p>The nature paper talks bout just tumor/normal c classification with the data set without pixel-level annotations. So, I am extending this idea to solve this problem with an extra model for classifying tumor into Gleason Score - _3/4/5. And from the Gleason score to ISUP is a simple rule.</p>\n\n<p>I don't have time to complete this so posted to the community if anyone can take it further.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "912576": "I like to share one of the best approaches to solve this problem. I have started with this approach but due to my engagement with other work, I couldn't complete it :(\n\nI think this approch  might help others as this method is published in **Nature  by Paige AI**\n\n**Note** I am an engineer, not a pathologist but whatever I share here is based on my interaction with the senior pathologists.\n\nIn clinical practices, Gleason Score is a must-have solution so any algorithm which gives ISUP w/o Gleason Score is not good for pathology.\n\n**The issues at hand**-\n- The annotation is very noisy\n- GT Mask of the WSI doesn't have detailed Gleason score information\n\n\nWhat  MIL Nature paper says-\n\n&gt; The development of decision support systems for pathology and their deployment in clinical practice have been hindered by the need for large manually annotated datasets. To overcome this problem, **we present a multiple instance learning-based deep learning system that uses only the reported diagnoses as labels for training**, thereby avoiding expensive and time-consuming\npixel-wise manual annotations. We evaluated this framework at scale on a dataset of 44,732 whole slide images from 15,187 patients without any form of data curation. Tests on prostate cancer, basal cell carcinoma and breast cancer metastases to axillary lymph nodes resulted in areas under the curve above 0.98 for all cancer types. Its clinical application would allow\npathologists to exclude 65–75% of slides while retaining 100% sensitivity. Our results show that this system has the ability to train accurate classification models at unprecedented scale, laying the foundation for the deployment of computational decision support systems in clinical practice.\n\n **My Proposed Approach**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F505538%2F9db2256fdd1ccda30e8537469f3ceb29%2Fmil.PNG?generation=1593702600561260&amp;alt=media)\n\n\n\n**MIL Algorithm**\nThis works well with binary classifiers but not sure how it works for multi-class classification.  I have started with multi-class but realized the algorithm limits. So started the following paper and developed an approach for tumor/normal classification first.  After that, all normal slides will be Gealson Score -0 and for others w we need another classifier/CNN model to get Gleason-3/4/5.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F505538%2F9382debff24b94e2cd47655ee0bd92fb%2Fmilflow.PNG?generation=1593702948273717&amp;alt=media)\n\n\nFor detailed -\nRead paper-\nNature paper - https://www.nature.com/articles/s41591-019-0508-1\n\nGithub code-\nhttps://github.com/MSKCC-Computational-Pathology/MIL-nature-medicine-2019\n\nMyInference Kernel\n[My Inference Kernel](https://www.kaggle.com/sumitjha19/gleason-to-isup-score-resnet50)\nhttps://www.kaggle.com/sumitjha19/gleason-to-isup-score-resnet50\n\n**My Experience**\n- This approach needs Huge RAM and GPU capacity to run. \n- Kaggle kernel doesn't meet RAM requirements.\n\nI don't think we have so much time to train a model from scratch. But, I am open for collaboration on this approach.\n\n\n\n\n\n\n\n\n\n\n\n",
    "912678": "Very interesting but, as you say, this is a multi-class classification. And you must consider that Gleason score is just a part of the work, the submission it's about ISUP grade. I wish you succesful with your interesting approach."
  }
}