{
  "id": 463824,
  "title": "Segmentation Model Value",
  "url": "/competitions/UBC-OCEAN/discussion/463824",
  "author_name": "chemdatafarmer",
  "post_date": "2023-12-27T13:21:26.055000",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi Everyone,</p>\n<p>Has anyone had any luck with segmentation models to extract more tile level annotations for the data that was not masked for us?</p>\n<p>I've used the tiles that <a href=\"https://www.kaggle.com/jirkaborovec\" target=\"_blank\">@jirkaborovec</a> processed from the ~152 images with masks to train an instance classifier that has ~98% accuracy on a validation set taken from the same set (i.e. the tiles were from slides that have tiles in the training set, but the validation tiles are not in the training set). However, when I transfer this instance based classifier to a slide level MIL like context on slides it has seen vs slides it hasn't (i.e. the 152 WSIs with masks vs the ones without) I see a pretty steep drop off in accuracy even when the model has some pretty heavy augmentation and regularization techniques used during training. This shows me there's a lot of information left on the table if I could train with closer to the full image set, and I'm curious what others experiences have been.</p>",
  "messages": [
    {
      "id": 2576620,
      "postDate": "2023-12-28T00:06:09.437Z",
      "content": "<p>I had a similar approach as you, with similar results (above 99% of accuracy on training), but the problem is that one type of cancer contains tissues that belong to others kinds of cancer. For example, the image could be labeled as LGSC and contain only LGSC tissue, but other could be labeled as EC and contain tissues like LGSC and CC. Sadly, the data is not labeled for a segmentation task, but in my opinion, it would be the easiest way to solve the lack of data.</p>",
      "rawMarkdown": "I had a similar approach as you, with similar results (above 99% of accuracy on training), but the problem is that one type of cancer contains tissues that belong to others kinds of cancer. For example, the image could be labeled as LGSC and contain only LGSC tissue, but other could be labeled as EC and contain tissues like LGSC and CC. Sadly, the data is not labeled for a segmentation task, but in my opinion, it would be the easiest way to solve the lack of data.",
      "votes": 3,
      "replies": [
        {
          "id": 2576679,
          "postDate": "2023-12-28T03:03:11.600Z",
          "content": "<p>I was trying the very same approach, i used u net and it would score, 99.6% accuracy on validation, but it only had 0.33 IoU. </p>\n<p>I wonder what other approach could be used for this challenge with an incorrect scoring metric</p>",
          "rawMarkdown": "I was trying the very same approach, i used u net and it would score, 99.6% accuracy on validation, but it only had 0.33 IoU. \n\nI wonder what other approach could be used for this challenge with an incorrect scoring metric",
          "votes": 2,
          "replies": [
            {
              "id": 2577370,
              "postDate": "2023-12-28T14:44:11.657Z",
              "content": "<p>I suspect that when we build classifiers on tiles from the 152 segmented images, the model learns something about the image as well as the cancer phenotype and our validation accuracy is a bit misleading (even if the training set is heavily augmented). If I take a mosaic of 64 tiles from the WSIs and count votes where the model is &gt;90% confident the tile is of class y, I get ~70% balanced accuracy, but that only translates to 38% on the leaderboard. If I predict on WSIs that have tiles as part of the training data, I get near perfect balanced accuracy. If I predict on the 25 TMAs (not used for training) I get ~65% balanced accuracy. If I were to get 60% correct on the test set without predicting \"other\", I think the LB score should be around 0.44.</p>\n<p>I wish we had info on where the slides were from. It would be awesome to compare splits where we hold different clinics out vs when we compare splits that hold samples out from the same clinics.</p>",
              "rawMarkdown": "I suspect that when we build classifiers on tiles from the 152 segmented images, the model learns something about the image as well as the cancer phenotype and our validation accuracy is a bit misleading (even if the training set is heavily augmented). If I take a mosaic of 64 tiles from the WSIs and count votes where the model is >90% confident the tile is of class y, I get ~70% balanced accuracy, but that only translates to 38% on the leaderboard. If I predict on WSIs that have tiles as part of the training data, I get near perfect balanced accuracy. If I predict on the 25 TMAs (not used for training) I get ~65% balanced accuracy. If I were to get 60% correct on the test set without predicting \"other\", I think the LB score should be around 0.44.\n\nI wish we had info on where the slides were from. It would be awesome to compare splits where we hold different clinics out vs when we compare splits that hold samples out from the same clinics.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2576080,
      "postDate": "2023-12-27T13:21:26.057Z",
      "content": "<p>Hi Everyone,</p>\n<p>Has anyone had any luck with segmentation models to extract more tile level annotations for the data that was not masked for us?</p>\n<p>I've used the tiles that <a href=\"https://www.kaggle.com/jirkaborovec\" target=\"_blank\">@jirkaborovec</a> processed from the ~152 images with masks to train an instance classifier that has ~98% accuracy on a validation set taken from the same set (i.e. the tiles were from slides that have tiles in the training set, but the validation tiles are not in the training set). However, when I transfer this instance based classifier to a slide level MIL like context on slides it has seen vs slides it hasn't (i.e. the 152 WSIs with masks vs the ones without) I see a pretty steep drop off in accuracy even when the model has some pretty heavy augmentation and regularization techniques used during training. This shows me there's a lot of information left on the table if I could train with closer to the full image set, and I'm curious what others experiences have been.</p>",
      "rawMarkdown": "Hi Everyone,\n\nHas anyone had any luck with segmentation models to extract more tile level annotations for the data that was not masked for us?\n\nI've used the tiles that @jirkaborovec processed from the ~152 images with masks to train an instance classifier that has ~98% accuracy on a validation set taken from the same set (i.e. the tiles were from slides that have tiles in the training set, but the validation tiles are not in the training set). However, when I transfer this instance based classifier to a slide level MIL like context on slides it has seen vs slides it hasn't (i.e. the 152 WSIs with masks vs the ones without) I see a pretty steep drop off in accuracy even when the model has some pretty heavy augmentation and regularization techniques used during training. This shows me there's a lot of information left on the table if I could train with closer to the full image set, and I'm curious what others experiences have been.",
      "votes": 3
    },
    {
      "id": 2576091,
      "postDate": "2023-12-27T13:32:56.270Z",
      "content": "<p>It's worth noting that when I say \"transfer this instance based classifier to a slide level MIL like context\", what I mean is using a simple aggregation scoring method (max logits, sum logits, etc) to take the tile level predictions and come up with a slide level prediction. I haven't had time yet to try training a MIL classifier with the instance based classifier as an embedding layer just yet. If anyone wants to chime in on how that's gone for them, I'd be happy to hear about it!</p>",
      "rawMarkdown": "It's worth noting that when I say \"transfer this instance based classifier to a slide level MIL like context\", what I mean is using a simple aggregation scoring method (max logits, sum logits, etc) to take the tile level predictions and come up with a slide level prediction. I haven't had time yet to try training a MIL classifier with the instance based classifier as an embedding layer just yet. If anyone wants to chime in on how that's gone for them, I'd be happy to hear about it!"
    }
  ],
  "comments": [
    {
      "id": 2576620,
      "author_name": "Cristian Cuadrado",
      "author_url": "",
      "post_date": "2023-12-28T00:06:09.437000",
      "content": "<p>I had a similar approach as you, with similar results (above 99% of accuracy on training), but the problem is that one type of cancer contains tissues that belong to others kinds of cancer. For example, the image could be labeled as LGSC and contain only LGSC tissue, but other could be labeled as EC and contain tissues like LGSC and CC. Sadly, the data is not labeled for a segmentation task, but in my opinion, it would be the easiest way to solve the lack of data.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2576679,
          "author_name": "Varun Manoj Gupta",
          "author_url": "",
          "post_date": "2023-12-28T03:03:11.600000",
          "content": "<p>I was trying the very same approach, i used u net and it would score, 99.6% accuracy on validation, but it only had 0.33 IoU. </p>\n<p>I wonder what other approach could be used for this challenge with an incorrect scoring metric</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2577370,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "2023-12-28T14:44:11.657000",
              "content": "<p>I suspect that when we build classifiers on tiles from the 152 segmented images, the model learns something about the image as well as the cancer phenotype and our validation accuracy is a bit misleading (even if the training set is heavily augmented). If I take a mosaic of 64 tiles from the WSIs and count votes where the model is &gt;90% confident the tile is of class y, I get ~70% balanced accuracy, but that only translates to 38% on the leaderboard. If I predict on WSIs that have tiles as part of the training data, I get near perfect balanced accuracy. If I predict on the 25 TMAs (not used for training) I get ~65% balanced accuracy. If I were to get 60% correct on the test set without predicting \"other\", I think the LB score should be around 0.44.</p>\n<p>I wish we had info on where the slides were from. It would be awesome to compare splits where we hold different clinics out vs when we compare splits that hold samples out from the same clinics.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2576091,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "2023-12-27T13:32:56.270000",
      "content": "<p>It's worth noting that when I say \"transfer this instance based classifier to a slide level MIL like context\", what I mean is using a simple aggregation scoring method (max logits, sum logits, etc) to take the tile level predictions and come up with a slide level prediction. I haven't had time yet to try training a MIL classifier with the instance based classifier as an embedding layer just yet. If anyone wants to chime in on how that's gone for them, I'd be happy to hear about it!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2576620": "I had a similar approach as you, with similar results (above 99% of accuracy on training), but the problem is that one type of cancer contains tissues that belong to others kinds of cancer. For example, the image could be labeled as LGSC and contain only LGSC tissue, but other could be labeled as EC and contain tissues like LGSC and CC. Sadly, the data is not labeled for a segmentation task, but in my opinion, it would be the easiest way to solve the lack of data.",
    "2576080": "Hi Everyone,\n\nHas anyone had any luck with segmentation models to extract more tile level annotations for the data that was not masked for us?\n\nI've used the tiles that @jirkaborovec processed from the ~152 images with masks to train an instance classifier that has ~98% accuracy on a validation set taken from the same set (i.e. the tiles were from slides that have tiles in the training set, but the validation tiles are not in the training set). However, when I transfer this instance based classifier to a slide level MIL like context on slides it has seen vs slides it hasn't (i.e. the 152 WSIs with masks vs the ones without) I see a pretty steep drop off in accuracy even when the model has some pretty heavy augmentation and regularization techniques used during training. This shows me there's a lot of information left on the table if I could train with closer to the full image set, and I'm curious what others experiences have been.",
    "2576091": "It's worth noting that when I say \"transfer this instance based classifier to a slide level MIL like context\", what I mean is using a simple aggregation scoring method (max logits, sum logits, etc) to take the tile level predictions and come up with a slide level prediction. I haven't had time yet to try training a MIL classifier with the instance based classifier as an embedding layer just yet. If anyone wants to chime in on how that's gone for them, I'd be happy to hear about it!"
  }
}