{
  "id": 348987,
  "title": "Tiling did not work for me",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/348987",
  "author_name": "moth",
  "post_date": "2022-08-30T20:30:07.773000",
  "votes": 2,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hello guys,</p>\n<p>I was wondering if anyone has had good results so far with tiling. My approach has been classifying on tiles and then fusing the scores from all the tiles from each image to get the final prediction for each image. </p>\n<p>This process is far more compute intensive and complex than resizing WSIs and training classifiers on resized images. What has been your approach so far? </p>",
  "messages": [
    {
      "id": 1923407,
      "postDate": "2022-09-02T07:49:49.677Z",
      "content": "<p>I extensively investigated tiling as well. Different sizes, shrinkage, different augmentations, different transforms, normalization methods, etc. It never seems to yield a good score. Very depressing since many scientific papers rely on the method to get good results. </p>",
      "rawMarkdown": "I extensively investigated tiling as well. Different sizes, shrinkage, different augmentations, different transforms, normalization methods, etc. It never seems to yield a good score. Very depressing since many scientific papers rely on the method to get good results. ",
      "votes": 3,
      "replies": [
        {
          "id": 1923723,
          "postDate": "2022-09-02T13:32:43.950Z",
          "content": "<p>As you mentioned there are multiple papers on this topic, plus I think the intuition goes along with this idea since tiling means you have much more data and you dont lose information by resizing</p>",
          "rawMarkdown": "As you mentioned there are multiple papers on this topic, plus I think the intuition goes along with this idea since tiling means you have much more data and you dont lose information by resizing"
        }
      ]
    },
    {
      "id": 1922525,
      "postDate": "2022-09-01T14:50:25.393Z",
      "content": "<p>It requires more compute but you also get the benefit of high resolution which can be a deal breaker (Not sure it is the case here..) </p>",
      "rawMarkdown": "It requires more compute but you also get the benefit of high resolution which can be a deal breaker (Not sure it is the case here..) ",
      "votes": 1,
      "replies": [
        {
          "id": 1922777,
          "postDate": "2022-09-01T17:43:00.373Z",
          "content": "<p>That was my intuition but my results dont correlate with this.</p>",
          "rawMarkdown": "That was my intuition but my results dont correlate with this."
        }
      ]
    },
    {
      "id": 1921736,
      "postDate": "2022-09-01T02:51:42.517Z",
      "content": "<p>Can I ask how you classified on tiles? We don't have instance level labels right?</p>",
      "rawMarkdown": "Can I ask how you classified on tiles? We don't have instance level labels right?",
      "votes": 1,
      "replies": [
        {
          "id": 1922566,
          "postDate": "2022-09-01T15:12:45.123Z",
          "content": "<p>I split the original images into lower dimensional tiles, removed background and assigned the image class to the tile. Then predicted the class probability and averaged.</p>",
          "rawMarkdown": "I split the original images into lower dimensional tiles, removed background and assigned the image class to the tile. Then predicted the class probability and averaged.",
          "votes": 1
        },
        {
          "id": 1922604,
          "postDate": "2022-09-01T15:33:37.597Z",
          "content": "<p>but every tile may not contain the signal right? or is it the case?</p>",
          "rawMarkdown": "but every tile may not contain the signal right? or is it the case?"
        }
      ]
    },
    {
      "id": 1921561,
      "postDate": "2022-09-01T00:03:32.570Z",
      "content": "<p>Hi, I think that using tiles reduce the possibility of applying complex models and everything is focused on a simple model but maybe with a lot of customization, I agree with what you mentioned about tiles required more processing , memory, etc.</p>\n<p>Thanks</p>",
      "rawMarkdown": "Hi, I think that using tiles reduce the possibility of applying complex models and everything is focused on a simple model but maybe with a lot of customization, I agree with what you mentioned about tiles required more processing , memory, etc.\n\nThanks",
      "votes": 1,
      "replies": [
        {
          "id": 1922535,
          "postDate": "2022-09-01T14:53:04.147Z",
          "content": "<p>Definitely more computation and memory is required. Results dont match my intuition since through tiling you get images with higher resolution and more amount compared to resizing, maybe its just my pipeline who knows </p>",
          "rawMarkdown": "Definitely more computation and memory is required. Results dont match my intuition since through tiling you get images with higher resolution and more amount compared to resizing, maybe its just my pipeline who knows "
        }
      ]
    },
    {
      "id": 1923696,
      "postDate": "2022-09-02T13:13:10.040Z",
      "content": "<p>you need a trained a patch feature extraction \"first\" (self or weak supervised) if you want to use tiles.</p>\n<p>i think it is not good enough to subsample a few tiles per image. you are likely to miss the signal.<br>\nAlternatively, if you can detect ROI in the whole slide, then extract tiles from the ROI.</p>\n<p>it is also not good enough to label each tile by the class. only a few tile contains the signal</p>\n<hr>\n<p>on a side note the heatmap from full image classification model could reveal this?</p>",
      "rawMarkdown": "you need a trained a patch feature extraction \"first\" (self or weak supervised) if you want to use tiles.\n\ni think it is not good enough to subsample a few tiles per image. you are likely to miss the signal.\nAlternatively, if you can detect ROI in the whole slide, then extract tiles from the ROI.\n\nit is also not good enough to label each tile by the class. only a few tile contains the signal\n\n---\n\non a side note the heatmap from full image classification model could reveal this?",
      "votes": 2,
      "replies": [
        {
          "id": 1923719,
          "postDate": "2022-09-02T13:30:54.583Z",
          "content": "<p>I performed clot detection vs background and then trained a model on only-clots images, still this didnt work that well on test data. My ROC at CV was around 0.75, which is not great but definitely not bad either 🤷🏼</p>",
          "rawMarkdown": "I performed clot detection vs background and then trained a model on only-clots images, still this didnt work that well on test data. My ROC at CV was around 0.75, which is not great but definitely not bad either 🤷🏼"
        }
      ]
    },
    {
      "id": 1920052,
      "postDate": "2022-08-30T20:30:07.773Z",
      "content": "<p>Hello guys,</p>\n<p>I was wondering if anyone has had good results so far with tiling. My approach has been classifying on tiles and then fusing the scores from all the tiles from each image to get the final prediction for each image. </p>\n<p>This process is far more compute intensive and complex than resizing WSIs and training classifiers on resized images. What has been your approach so far? </p>",
      "rawMarkdown": "Hello guys,\n\nI was wondering if anyone has had good results so far with tiling. My approach has been classifying on tiles and then fusing the scores from all the tiles from each image to get the final prediction for each image. \n\nThis process is far more compute intensive and complex than resizing WSIs and training classifiers on resized images. What has been your approach so far? ",
      "votes": 2
    },
    {
      "id": 1939184,
      "postDate": "2022-09-14T15:12:01.897Z",
      "content": "<p>In my experience tile selection from a whole slide is a key component. If you use tiles with some threshold for background avoidance, there's a lot of noisy signal in this dataset, difficult to train any meaningful network. I have tried using a pretrained network to rank tiles and use top x% tiles for training - this was a bit better, but still there's lots of way to overfit the model. Plus there is the tile prediction aggregation method that adds more complexity to classify slide level image. Here also, in my very limited experience, I found that taking top x% tiles for aggregating works decently well and taking many tiles from images did not help that much but increased computation load a lot.</p>\n<p>Of course these are something I tried, I'm pretty sure some of my implementations have been poor :)</p>",
      "rawMarkdown": "In my experience tile selection from a whole slide is a key component. If you use tiles with some threshold for background avoidance, there's a lot of noisy signal in this dataset, difficult to train any meaningful network. I have tried using a pretrained network to rank tiles and use top x% tiles for training - this was a bit better, but still there's lots of way to overfit the model. Plus there is the tile prediction aggregation method that adds more complexity to classify slide level image. Here also, in my very limited experience, I found that taking top x% tiles for aggregating works decently well and taking many tiles from images did not help that much but increased computation load a lot.\n\nOf course these are something I tried, I'm pretty sure some of my implementations have been poor :)"
    }
  ],
  "comments": [
    {
      "id": 1923407,
      "author_name": "Arthur",
      "author_url": "",
      "post_date": "2022-09-02T07:49:49.677000",
      "content": "<p>I extensively investigated tiling as well. Different sizes, shrinkage, different augmentations, different transforms, normalization methods, etc. It never seems to yield a good score. Very depressing since many scientific papers rely on the method to get good results. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1923723,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2022-09-02T13:32:43.950000",
          "content": "<p>As you mentioned there are multiple papers on this topic, plus I think the intuition goes along with this idea since tiling means you have much more data and you dont lose information by resizing</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1922525,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-09-01T14:50:25.393000",
      "content": "<p>It requires more compute but you also get the benefit of high resolution which can be a deal breaker (Not sure it is the case here..) </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1922777,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2022-09-01T17:43:00.373000",
          "content": "<p>That was my intuition but my results dont correlate with this.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1921736,
      "author_name": "Rajat Chaudhari",
      "author_url": "",
      "post_date": "2022-09-01T02:51:42.517000",
      "content": "<p>Can I ask how you classified on tiles? We don't have instance level labels right?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1922566,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2022-09-01T15:12:45.123000",
          "content": "<p>I split the original images into lower dimensional tiles, removed background and assigned the image class to the tile. Then predicted the class probability and averaged.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1922604,
          "author_name": "Rajat Chaudhari",
          "author_url": "",
          "post_date": "2022-09-01T15:33:37.597000",
          "content": "<p>but every tile may not contain the signal right? or is it the case?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1921561,
      "author_name": "Pablo Larrosa",
      "author_url": "",
      "post_date": "2022-09-01T00:03:32.570000",
      "content": "<p>Hi, I think that using tiles reduce the possibility of applying complex models and everything is focused on a simple model but maybe with a lot of customization, I agree with what you mentioned about tiles required more processing , memory, etc.</p>\n<p>Thanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1922535,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2022-09-01T14:53:04.147000",
          "content": "<p>Definitely more computation and memory is required. Results dont match my intuition since through tiling you get images with higher resolution and more amount compared to resizing, maybe its just my pipeline who knows </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1923696,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-09-02T13:13:10.040000",
      "content": "<p>you need a trained a patch feature extraction \"first\" (self or weak supervised) if you want to use tiles.</p>\n<p>i think it is not good enough to subsample a few tiles per image. you are likely to miss the signal.<br>\nAlternatively, if you can detect ROI in the whole slide, then extract tiles from the ROI.</p>\n<p>it is also not good enough to label each tile by the class. only a few tile contains the signal</p>\n<hr>\n<p>on a side note the heatmap from full image classification model could reveal this?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1923719,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2022-09-02T13:30:54.583000",
          "content": "<p>I performed clot detection vs background and then trained a model on only-clots images, still this didnt work that well on test data. My ROC at CV was around 0.75, which is not great but definitely not bad either 🤷🏼</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1939184,
      "author_name": "tdiceman",
      "author_url": "",
      "post_date": "2022-09-14T15:12:01.897000",
      "content": "<p>In my experience tile selection from a whole slide is a key component. If you use tiles with some threshold for background avoidance, there's a lot of noisy signal in this dataset, difficult to train any meaningful network. I have tried using a pretrained network to rank tiles and use top x% tiles for training - this was a bit better, but still there's lots of way to overfit the model. Plus there is the tile prediction aggregation method that adds more complexity to classify slide level image. Here also, in my very limited experience, I found that taking top x% tiles for aggregating works decently well and taking many tiles from images did not help that much but increased computation load a lot.</p>\n<p>Of course these are something I tried, I'm pretty sure some of my implementations have been poor :)</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1923407": "I extensively investigated tiling as well. Different sizes, shrinkage, different augmentations, different transforms, normalization methods, etc. It never seems to yield a good score. Very depressing since many scientific papers rely on the method to get good results. ",
    "1922525": "It requires more compute but you also get the benefit of high resolution which can be a deal breaker (Not sure it is the case here..) ",
    "1921736": "Can I ask how you classified on tiles? We don't have instance level labels right?",
    "1921561": "Hi, I think that using tiles reduce the possibility of applying complex models and everything is focused on a simple model but maybe with a lot of customization, I agree with what you mentioned about tiles required more processing , memory, etc.\n\nThanks",
    "1923696": "you need a trained a patch feature extraction \"first\" (self or weak supervised) if you want to use tiles.\n\ni think it is not good enough to subsample a few tiles per image. you are likely to miss the signal.\nAlternatively, if you can detect ROI in the whole slide, then extract tiles from the ROI.\n\nit is also not good enough to label each tile by the class. only a few tile contains the signal\n\n---\n\non a side note the heatmap from full image classification model could reveal this?",
    "1920052": "Hello guys,\n\nI was wondering if anyone has had good results so far with tiling. My approach has been classifying on tiles and then fusing the scores from all the tiles from each image to get the final prediction for each image. \n\nThis process is far more compute intensive and complex than resizing WSIs and training classifiers on resized images. What has been your approach so far? ",
    "1939184": "In my experience tile selection from a whole slide is a key component. If you use tiles with some threshold for background avoidance, there's a lot of noisy signal in this dataset, difficult to train any meaningful network. I have tried using a pretrained network to rank tiles and use top x% tiles for training - this was a bit better, but still there's lots of way to overfit the model. Plus there is the tile prediction aggregation method that adds more complexity to classify slide level image. Here also, in my very limited experience, I found that taking top x% tiles for aggregating works decently well and taking many tiles from images did not help that much but increased computation load a lot.\n\nOf course these are something I tried, I'm pretty sure some of my implementations have been poor :)"
  }
}