{
  "id": 357452,
  "title": "Resources for Multiple Instance Learning (MIL)",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/357452",
  "author_name": "Gunes Evitan",
  "post_date": "2022-10-04T11:35:50.314000",
  "votes": 12,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I know the competition is ending very soon but I got interested in this topic and I'm trying some of the papers. These are the resources I found. Unfortunately, there aren't many papers but the good thing is most of them are tested in this particular task. I hope this will be useful for people researching MIL.</p>\n<p>I started with this survey paper to get a general idea about MIL.<br>\nMultiple Instance Learning for Digital Pathology: A Review on the State-of-the-Art, Limitations &amp; Future Potential<br>\n<a href=\"https://arxiv.org/pdf/2206.04425.pdf\" target=\"_blank\">https://arxiv.org/pdf/2206.04425.pdf</a></p>\n<p>This paper proposes a stochastic tile sampling method but I don't think it's practical here. <br>\nMonte-Carlo Sampling applied to Multiple Instance Learning for Histological Image Classification<br>\n<a href=\"https://arxiv.org/abs/1812.11560\" target=\"_blank\">https://arxiv.org/abs/1812.11560</a></p>\n<p>This is what I'm trying to do right now. An attention mechanism for multiple instance feature maps could be useful.<br>\nAttention-based Deep Multiple Instance Learning<br>\n<a href=\"https://arxiv.org/pdf/1802.04712v4.pdf\" target=\"_blank\">https://arxiv.org/pdf/1802.04712v4.pdf</a></p>\n<p>TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image Classification<br>\n<a href=\"https://arxiv.org/abs/2106.00908\" target=\"_blank\">https://arxiv.org/abs/2106.00908</a></p>",
  "messages": [
    {
      "id": 1970957,
      "postDate": "2022-10-04T11:35:50.313Z",
      "content": "<p>I know the competition is ending very soon but I got interested in this topic and I'm trying some of the papers. These are the resources I found. Unfortunately, there aren't many papers but the good thing is most of them are tested in this particular task. I hope this will be useful for people researching MIL.</p>\n<p>I started with this survey paper to get a general idea about MIL.<br>\nMultiple Instance Learning for Digital Pathology: A Review on the State-of-the-Art, Limitations &amp; Future Potential<br>\n<a href=\"https://arxiv.org/pdf/2206.04425.pdf\" target=\"_blank\">https://arxiv.org/pdf/2206.04425.pdf</a></p>\n<p>This paper proposes a stochastic tile sampling method but I don't think it's practical here. <br>\nMonte-Carlo Sampling applied to Multiple Instance Learning for Histological Image Classification<br>\n<a href=\"https://arxiv.org/abs/1812.11560\" target=\"_blank\">https://arxiv.org/abs/1812.11560</a></p>\n<p>This is what I'm trying to do right now. An attention mechanism for multiple instance feature maps could be useful.<br>\nAttention-based Deep Multiple Instance Learning<br>\n<a href=\"https://arxiv.org/pdf/1802.04712v4.pdf\" target=\"_blank\">https://arxiv.org/pdf/1802.04712v4.pdf</a></p>\n<p>TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image Classification<br>\n<a href=\"https://arxiv.org/abs/2106.00908\" target=\"_blank\">https://arxiv.org/abs/2106.00908</a></p>",
      "rawMarkdown": "I know the competition is ending very soon but I got interested in this topic and I'm trying some of the papers. These are the resources I found. Unfortunately, there aren't many papers but the good thing is most of them are tested in this particular task. I hope this will be useful for people researching MIL.\n\nI started with this survey paper to get a general idea about MIL.\nMultiple Instance Learning for Digital Pathology: A Review on the State-of-the-Art, Limitations & Future Potential\nhttps://arxiv.org/pdf/2206.04425.pdf\n\nThis paper proposes a stochastic tile sampling method but I don't think it's practical here. \nMonte-Carlo Sampling applied to Multiple Instance Learning for Histological Image Classification\nhttps://arxiv.org/abs/1812.11560\n\nThis is what I'm trying to do right now. An attention mechanism for multiple instance feature maps could be useful.\nAttention-based Deep Multiple Instance Learning\nhttps://arxiv.org/pdf/1802.04712v4.pdf\n\nTransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image Classification\nhttps://arxiv.org/abs/2106.00908",
      "votes": 12
    },
    {
      "id": 1973219,
      "postDate": "2022-10-05T14:28:17.663Z",
      "content": "<p>As I mentioned in the previous thread - I tried MIL with TensorFlow/Keras, and it failed at this task miserably, because of the weak signal.<br>\nI'll probably write a guide on doing binary class MIL and publish it with practical code as there aren't many resources covering the concept.</p>",
      "rawMarkdown": "As I mentioned in the previous thread - I tried MIL with TensorFlow/Keras, and it failed at this task miserably, because of the weak signal.\nI'll probably write a guide on doing binary class MIL and publish it with practical code as there aren't many resources covering the concept.",
      "votes": 4,
      "replies": [
        {
          "id": 1973495,
          "postDate": "2022-10-05T16:25:58.177Z",
          "content": "<p>I have expressly tried MIL as well, and I also feel I could not get it work very well. Would love to see what all methodologies were used for training and slide level aggregation. Thanks <a href=\"https://www.kaggle.com/davidlandup\" target=\"_blank\">@davidlandup</a> </p>",
          "rawMarkdown": "I have expressly tried MIL as well, and I also feel I could not get it work very well. Would love to see what all methodologies were used for training and slide level aggregation. Thanks @davidlandup ",
          "votes": 1
        },
        {
          "id": 1973506,
          "postDate": "2022-10-05T16:32:55.980Z",
          "content": "<p>Thanks for sharing your experiences. What exactly didn't work for you?</p>",
          "rawMarkdown": "Thanks for sharing your experiences. What exactly didn't work for you?"
        },
        {
          "id": 1973660,
          "postDate": "2022-10-05T18:46:32.317Z",
          "content": "<p>To be honest, when I say did not work, I mean more that I could not make it work rather than blaming the technique. Now how much of that is my lack of experience and how much is lack of signal remains to be seen. </p>\n<p>In short, <a href=\"https://www.kaggle.com/code/icemantd/feature-cluster-tile-psuedo-label-train-mayo\" target=\"_blank\">this notebook</a> summarizes my training methodology. I did a shabby job writing this, and it probably has an error in computing log_loss for categorical target variables, but the method worked to some extent inspired by the research article mentioned. With this portion I trained a somewhat discriminatory tile feature extractor. </p>\n<p>For the next part I aggregate these features using attention to maximum logit of a tile from a slide, and using aggregation methodology inspired by <a href=\"https://arxiv.org/pdf/2011.08939.pdf\" target=\"_blank\">this</a> paper. Finally, I used a random forest classifier on aggregated features for classifications.</p>\n<p>The main issues I found in my implementation, that could explain a lot of blame of not working well are:</p>\n<p>1) My tile creation and selection was very naive, based on simple threshold methods and scoring by a pre-trained network. I also found that at one point my model was suffering from bias stemming from more tiles from larger images and very few tiles from smaller images (model was doing better on larger size images).<br>\n2) Since this a complicated 2-step method, I could not be rigorous in checking CV performance. I also found it difficult to reliably evaluate feature extractor training, for which I used maxpooling scores - i.e. taking maximum predicted logits for the two classes for any tile in the slide image and taking log_loss based on that as validation score on holdout data to stop training.</p>\n<p>Much to learn still, but this itself was exciting (and at times frustrating). So I am very interested to see more direct implementations of these papers that generally do well on such problems - assuming data signal issues were not drastic in nature :)</p>",
          "rawMarkdown": "To be honest, when I say did not work, I mean more that I could not make it work rather than blaming the technique. Now how much of that is my lack of experience and how much is lack of signal remains to be seen. \n\nIn short, [this notebook](https://www.kaggle.com/code/icemantd/feature-cluster-tile-psuedo-label-train-mayo) summarizes my training methodology. I did a shabby job writing this, and it probably has an error in computing log_loss for categorical target variables, but the method worked to some extent inspired by the research article mentioned. With this portion I trained a somewhat discriminatory tile feature extractor. \n\nFor the next part I aggregate these features using attention to maximum logit of a tile from a slide, and using aggregation methodology inspired by [this](https://arxiv.org/pdf/2011.08939.pdf) paper. Finally, I used a random forest classifier on aggregated features for classifications.\n\nThe main issues I found in my implementation, that could explain a lot of blame of not working well are:\n\n1) My tile creation and selection was very naive, based on simple threshold methods and scoring by a pre-trained network. I also found that at one point my model was suffering from bias stemming from more tiles from larger images and very few tiles from smaller images (model was doing better on larger size images).\n2) Since this a complicated 2-step method, I could not be rigorous in checking CV performance. I also found it difficult to reliably evaluate feature extractor training, for which I used maxpooling scores - i.e. taking maximum predicted logits for the two classes for any tile in the slide image and taking log_loss based on that as validation score on holdout data to stop training.\n\nMuch to learn still, but this itself was exciting (and at times frustrating). So I am very interested to see more direct implementations of these papers that generally do well on such problems - assuming data signal issues were not drastic in nature :)",
          "votes": 1
        },
        {
          "id": 1973773,
          "postDate": "2022-10-05T20:30:10.360Z",
          "content": "<p>I'm a simple person, I didn't use any mil, but I achieved 0.9976 accuracy with all sorts of tricks))))) my log_loss was at a minimum 0.004)))) It was fun and the result surprised me (I mean the metric of the competition, or its public version), it seems that none of us understood the metric of the organizers.</p>",
          "rawMarkdown": "I'm a simple person, I didn't use any mil, but I achieved 0.9976 accuracy with all sorts of tricks))))) my log_loss was at a minimum 0.004)))) It was fun and the result surprised me (I mean the metric of the competition, or its public version), it seems that none of us understood the metric of the organizers."
        }
      ]
    },
    {
      "id": 1971239,
      "postDate": "2022-10-04T13:50:26.087Z",
      "content": "<p>This is a famous paper on MIL for WSI: <a href=\"https://www.nature.com/articles/s41591-019-0508-1\" target=\"_blank\">https://www.nature.com/articles/s41591-019-0508-1</a></p>",
      "rawMarkdown": "This is a famous paper on MIL for WSI: https://www.nature.com/articles/s41591-019-0508-1"
    }
  ],
  "comments": [
    {
      "id": 1973219,
      "author_name": "David Landup",
      "author_url": "",
      "post_date": "2022-10-05T14:28:17.663000",
      "content": "<p>As I mentioned in the previous thread - I tried MIL with TensorFlow/Keras, and it failed at this task miserably, because of the weak signal.<br>\nI'll probably write a guide on doing binary class MIL and publish it with practical code as there aren't many resources covering the concept.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1973495,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-10-05T16:25:58.177000",
          "content": "<p>I have expressly tried MIL as well, and I also feel I could not get it work very well. Would love to see what all methodologies were used for training and slide level aggregation. Thanks <a href=\"https://www.kaggle.com/davidlandup\" target=\"_blank\">@davidlandup</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1973506,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-10-05T16:32:55.980000",
          "content": "<p>Thanks for sharing your experiences. What exactly didn't work for you?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1973660,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-10-05T18:46:32.317000",
          "content": "<p>To be honest, when I say did not work, I mean more that I could not make it work rather than blaming the technique. Now how much of that is my lack of experience and how much is lack of signal remains to be seen. </p>\n<p>In short, <a href=\"https://www.kaggle.com/code/icemantd/feature-cluster-tile-psuedo-label-train-mayo\" target=\"_blank\">this notebook</a> summarizes my training methodology. I did a shabby job writing this, and it probably has an error in computing log_loss for categorical target variables, but the method worked to some extent inspired by the research article mentioned. With this portion I trained a somewhat discriminatory tile feature extractor. </p>\n<p>For the next part I aggregate these features using attention to maximum logit of a tile from a slide, and using aggregation methodology inspired by <a href=\"https://arxiv.org/pdf/2011.08939.pdf\" target=\"_blank\">this</a> paper. Finally, I used a random forest classifier on aggregated features for classifications.</p>\n<p>The main issues I found in my implementation, that could explain a lot of blame of not working well are:</p>\n<p>1) My tile creation and selection was very naive, based on simple threshold methods and scoring by a pre-trained network. I also found that at one point my model was suffering from bias stemming from more tiles from larger images and very few tiles from smaller images (model was doing better on larger size images).<br>\n2) Since this a complicated 2-step method, I could not be rigorous in checking CV performance. I also found it difficult to reliably evaluate feature extractor training, for which I used maxpooling scores - i.e. taking maximum predicted logits for the two classes for any tile in the slide image and taking log_loss based on that as validation score on holdout data to stop training.</p>\n<p>Much to learn still, but this itself was exciting (and at times frustrating). So I am very interested to see more direct implementations of these papers that generally do well on such problems - assuming data signal issues were not drastic in nature :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1973773,
          "author_name": "Zaakcii Ru",
          "author_url": "",
          "post_date": "2022-10-05T20:30:10.360000",
          "content": "<p>I'm a simple person, I didn't use any mil, but I achieved 0.9976 accuracy with all sorts of tricks))))) my log_loss was at a minimum 0.004)))) It was fun and the result surprised me (I mean the metric of the competition, or its public version), it seems that none of us understood the metric of the organizers.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1971239,
      "author_name": "MITEL-UNIUD",
      "author_url": "",
      "post_date": "2022-10-04T13:50:26.087000",
      "content": "<p>This is a famous paper on MIL for WSI: <a href=\"https://www.nature.com/articles/s41591-019-0508-1\" target=\"_blank\">https://www.nature.com/articles/s41591-019-0508-1</a></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1970957": "I know the competition is ending very soon but I got interested in this topic and I'm trying some of the papers. These are the resources I found. Unfortunately, there aren't many papers but the good thing is most of them are tested in this particular task. I hope this will be useful for people researching MIL.\n\nI started with this survey paper to get a general idea about MIL.\nMultiple Instance Learning for Digital Pathology: A Review on the State-of-the-Art, Limitations & Future Potential\nhttps://arxiv.org/pdf/2206.04425.pdf\n\nThis paper proposes a stochastic tile sampling method but I don't think it's practical here. \nMonte-Carlo Sampling applied to Multiple Instance Learning for Histological Image Classification\nhttps://arxiv.org/abs/1812.11560\n\nThis is what I'm trying to do right now. An attention mechanism for multiple instance feature maps could be useful.\nAttention-based Deep Multiple Instance Learning\nhttps://arxiv.org/pdf/1802.04712v4.pdf\n\nTransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image Classification\nhttps://arxiv.org/abs/2106.00908",
    "1973219": "As I mentioned in the previous thread - I tried MIL with TensorFlow/Keras, and it failed at this task miserably, because of the weak signal.\nI'll probably write a guide on doing binary class MIL and publish it with practical code as there aren't many resources covering the concept.",
    "1971239": "This is a famous paper on MIL for WSI: https://www.nature.com/articles/s41591-019-0508-1"
  }
}