{
  "id": 346850,
  "title": "\"Other\" images",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/346850",
  "author_name": "Nikita Glazunov",
  "post_date": "2022-08-21T17:51:57.006000",
  "votes": 4,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I wonder if anyone here tried to add the images from \"other\" directory?<br>\nIf you did so how that affected an overall score?<br>\nOr probably this is bad idea? Why?</p>",
  "messages": [
    {
      "id": 1908466,
      "postDate": "2022-08-21T17:51:57.007Z",
      "content": "<p>I wonder if anyone here tried to add the images from \"other\" directory?<br>\nIf you did so how that affected an overall score?<br>\nOr probably this is bad idea? Why?</p>",
      "rawMarkdown": "I wonder if anyone here tried to add the images from \"other\" directory?\nIf you did so how that affected an overall score?\nOr probably this is bad idea? Why?",
      "votes": 4
    },
    {
      "id": 1908469,
      "postDate": "2022-08-21T17:56:06.777Z",
      "content": "<p>Now that you reminded me of this I'm curious as well</p>",
      "rawMarkdown": "Now that you reminded me of this I'm curious as well",
      "votes": 1
    },
    {
      "id": 1920049,
      "postDate": "2022-08-30T20:25:50.407Z",
      "content": "<p>Haven't used them yet, but probably the best thing to do is to apply pseudolabelling to these image and later used them for training. An extra +10% of data can be really useful in this competition.</p>",
      "rawMarkdown": "Haven't used them yet, but probably the best thing to do is to apply pseudolabelling to these image and later used them for training. An extra +10% of data can be really useful in this competition.",
      "votes": 2,
      "replies": [
        {
          "id": 1920411,
          "postDate": "2022-08-31T06:39:20.033Z",
          "content": "<p>What do you mean by pseudolabelling? I've heard about that but don't know the details. Can you help me with this?</p>",
          "rawMarkdown": "What do you mean by pseudolabelling? I've heard about that but don't know the details. Can you help me with this?",
          "votes": 1
        },
        {
          "id": 1920931,
          "postDate": "2022-08-31T13:35:59.700Z",
          "content": "<p>On pseudo-labelling you:</p>\n<ol>\n<li>Train model on a batch of labeled data</li>\n<li>Use the trained model to predict labels on a batch of unlabeled data (soft labels).</li>\n<li>Use the predicted labels to calculate the loss on unlabeled data (compare soft vs hard labels).</li>\n<li>Combine labeled loss with unlabeled loss and backpropagate.</li>\n</ol>\n<p>This article is the best one I have read about the topic:<br>\n<a href=\"https://towardsdatascience.com/pseudo-labeling-to-deal-with-small-datasets-what-why-how-fd6f903213af\" target=\"_blank\">https://towardsdatascience.com/pseudo-labeling-to-deal-with-small-datasets-what-why-how-fd6f903213af</a></p>",
          "rawMarkdown": "On pseudo-labelling you:\n1. Train model on a batch of labeled data\n2. Use the trained model to predict labels on a batch of unlabeled data (soft labels).\n3. Use the predicted labels to calculate the loss on unlabeled data (compare soft vs hard labels).\n4. Combine labeled loss with unlabeled loss and backpropagate.\n\nThis article is the best one I have read about the topic:\nhttps://towardsdatascience.com/pseudo-labeling-to-deal-with-small-datasets-what-why-how-fd6f903213af",
          "votes": 2
        },
        {
          "id": 1921007,
          "postDate": "2022-08-31T14:33:19.413Z",
          "content": "<p>To add to <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a>'s nice explanation - the specific problem in this case is that for most people labeling a handful of positive tiles itself is not easy. In the MIL problem we know that if the slide is positive for one class, there is at least one positive tile in that image, but there are many negative tiles that do not contribute to the classification. So, as I see it, in this case it is easier to do negative tile pseudo labeling, i.e. tell you network what is not classifier material. In this aspect the 'Other' tiles can play a role, because they do not belong to either of the two important classes. And the hope is that in all this mostly negative information, you find a way to contrast positive tiles pseudo labeling by the methods in the article, and your network/feature extractor learns enough to discriminate between useful and non-useful tiles.</p>\n<p>One e.g. I used with 'Other' tiles - if CE class target is [0,1] and LAA target is [1,0], I used 'Other' class target as [0,0]. It did help overall a little. Hope this helps and some of you can share your ways of pseudo labeling. I'm currently trying clustering based methods, will update once something works :)</p>\n<p>Edit: Shared <a href=\"https://www.kaggle.com/code/icemantd/feature-cluster-tile-psuedo-label-train-mayo\" target=\"_blank\">notebook</a> - please feel free to comment.</p>",
          "rawMarkdown": "To add to @alejopaullier's nice explanation - the specific problem in this case is that for most people labeling a handful of positive tiles itself is not easy. In the MIL problem we know that if the slide is positive for one class, there is at least one positive tile in that image, but there are many negative tiles that do not contribute to the classification. So, as I see it, in this case it is easier to do negative tile pseudo labeling, i.e. tell you network what is not classifier material. In this aspect the 'Other' tiles can play a role, because they do not belong to either of the two important classes. And the hope is that in all this mostly negative information, you find a way to contrast positive tiles pseudo labeling by the methods in the article, and your network/feature extractor learns enough to discriminate between useful and non-useful tiles.\n\nOne e.g. I used with 'Other' tiles - if CE class target is [0,1] and LAA target is [1,0], I used 'Other' class target as [0,0]. It did help overall a little. Hope this helps and some of you can share your ways of pseudo labeling. I'm currently trying clustering based methods, will update once something works :)\n\nEdit: Shared [notebook](https://www.kaggle.com/code/icemantd/feature-cluster-tile-psuedo-label-train-mayo) - please feel free to comment.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1909773,
      "postDate": "2022-08-22T23:21:58.863Z",
      "content": "<p>I have tried other image tiles from cases with label 'Other', leaving out 'Unknown' labels to mean it could include the two classes we are after. This 'Other' label becomes a negative case for training. Unfortunately we have very few of these 'definitely negative' images. It did improve the MIL based classification in my case to some extent, although it becomes a bit easier to overfit as <a href=\"https://www.kaggle.com/nikitaglazunov\" target=\"_blank\">@nikitaglazunov</a> pointed out.</p>",
      "rawMarkdown": "I have tried other image tiles from cases with label 'Other', leaving out 'Unknown' labels to mean it could include the two classes we are after. This 'Other' label becomes a negative case for training. Unfortunately we have very few of these 'definitely negative' images. It did improve the MIL based classification in my case to some extent, although it becomes a bit easier to overfit as @nikitaglazunov pointed out.",
      "votes": 2
    },
    {
      "id": 1908870,
      "postDate": "2022-08-22T05:38:17.640Z",
      "content": "<p>Well, I have just tried it. I added the extra data (around 65 images with \"Other\" label). Despite low validation loss (<strong>~ .820</strong>) I got a huge log loss score on unseen test data. I think main issues with my approach are:</p>\n<ul>\n<li>our training data became more imbalanced;</li>\n<li>\"other\" class is just irrelevant for the main task of competition.</li>\n</ul>",
      "rawMarkdown": "Well, I have just tried it. I added the extra data (around 65 images with \"Other\" label). Despite low validation loss (**~ .820**) I got a huge log loss score on unseen test data. I think main issues with my approach are:\n- our training data became more imbalanced;\n- \"other\" class is just irrelevant for the main task of competition.",
      "votes": 2
    },
    {
      "id": 1908746,
      "postDate": "2022-08-22T01:26:47.683Z",
      "content": "<p>I think I read in another paper somewhere that the authors used extra data (similar to our \"other\" directory) for the purpose of training the encoder. I think this is supposed to give more data so that they encoder can better understand the histopathology data. Sorry for the lack of a source.</p>",
      "rawMarkdown": "I think I read in another paper somewhere that the authors used extra data (similar to our \"other\" directory) for the purpose of training the encoder. I think this is supposed to give more data so that they encoder can better understand the histopathology data. Sorry for the lack of a source.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1908469,
      "author_name": "Rishav Nandi",
      "author_url": "",
      "post_date": "2022-08-21T17:56:06.777000",
      "content": "<p>Now that you reminded me of this I'm curious as well</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1920049,
      "author_name": "moth",
      "author_url": "",
      "post_date": "2022-08-30T20:25:50.407000",
      "content": "<p>Haven't used them yet, but probably the best thing to do is to apply pseudolabelling to these image and later used them for training. An extra +10% of data can be really useful in this competition.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1920411,
          "author_name": "Nikita Glazunov",
          "author_url": "",
          "post_date": "2022-08-31T06:39:20.033000",
          "content": "<p>What do you mean by pseudolabelling? I've heard about that but don't know the details. Can you help me with this?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1920931,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2022-08-31T13:35:59.700000",
          "content": "<p>On pseudo-labelling you:</p>\n<ol>\n<li>Train model on a batch of labeled data</li>\n<li>Use the trained model to predict labels on a batch of unlabeled data (soft labels).</li>\n<li>Use the predicted labels to calculate the loss on unlabeled data (compare soft vs hard labels).</li>\n<li>Combine labeled loss with unlabeled loss and backpropagate.</li>\n</ol>\n<p>This article is the best one I have read about the topic:<br>\n<a href=\"https://towardsdatascience.com/pseudo-labeling-to-deal-with-small-datasets-what-why-how-fd6f903213af\" target=\"_blank\">https://towardsdatascience.com/pseudo-labeling-to-deal-with-small-datasets-what-why-how-fd6f903213af</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1921007,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-08-31T14:33:19.413000",
          "content": "<p>To add to <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a>'s nice explanation - the specific problem in this case is that for most people labeling a handful of positive tiles itself is not easy. In the MIL problem we know that if the slide is positive for one class, there is at least one positive tile in that image, but there are many negative tiles that do not contribute to the classification. So, as I see it, in this case it is easier to do negative tile pseudo labeling, i.e. tell you network what is not classifier material. In this aspect the 'Other' tiles can play a role, because they do not belong to either of the two important classes. And the hope is that in all this mostly negative information, you find a way to contrast positive tiles pseudo labeling by the methods in the article, and your network/feature extractor learns enough to discriminate between useful and non-useful tiles.</p>\n<p>One e.g. I used with 'Other' tiles - if CE class target is [0,1] and LAA target is [1,0], I used 'Other' class target as [0,0]. It did help overall a little. Hope this helps and some of you can share your ways of pseudo labeling. I'm currently trying clustering based methods, will update once something works :)</p>\n<p>Edit: Shared <a href=\"https://www.kaggle.com/code/icemantd/feature-cluster-tile-psuedo-label-train-mayo\" target=\"_blank\">notebook</a> - please feel free to comment.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1909773,
      "author_name": "tdiceman",
      "author_url": "",
      "post_date": "2022-08-22T23:21:58.863000",
      "content": "<p>I have tried other image tiles from cases with label 'Other', leaving out 'Unknown' labels to mean it could include the two classes we are after. This 'Other' label becomes a negative case for training. Unfortunately we have very few of these 'definitely negative' images. It did improve the MIL based classification in my case to some extent, although it becomes a bit easier to overfit as <a href=\"https://www.kaggle.com/nikitaglazunov\" target=\"_blank\">@nikitaglazunov</a> pointed out.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1908870,
      "author_name": "Nikita Glazunov",
      "author_url": "",
      "post_date": "2022-08-22T05:38:17.640000",
      "content": "<p>Well, I have just tried it. I added the extra data (around 65 images with \"Other\" label). Despite low validation loss (<strong>~ .820</strong>) I got a huge log loss score on unseen test data. I think main issues with my approach are:</p>\n<ul>\n<li>our training data became more imbalanced;</li>\n<li>\"other\" class is just irrelevant for the main task of competition.</li>\n</ul>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1908746,
      "author_name": "yqz",
      "author_url": "",
      "post_date": "2022-08-22T01:26:47.683000",
      "content": "<p>I think I read in another paper somewhere that the authors used extra data (similar to our \"other\" directory) for the purpose of training the encoder. I think this is supposed to give more data so that they encoder can better understand the histopathology data. Sorry for the lack of a source.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1908466": "I wonder if anyone here tried to add the images from \"other\" directory?\nIf you did so how that affected an overall score?\nOr probably this is bad idea? Why?",
    "1908469": "Now that you reminded me of this I'm curious as well",
    "1920049": "Haven't used them yet, but probably the best thing to do is to apply pseudolabelling to these image and later used them for training. An extra +10% of data can be really useful in this competition.",
    "1909773": "I have tried other image tiles from cases with label 'Other', leaving out 'Unknown' labels to mean it could include the two classes we are after. This 'Other' label becomes a negative case for training. Unfortunately we have very few of these 'definitely negative' images. It did improve the MIL based classification in my case to some extent, although it becomes a bit easier to overfit as @nikitaglazunov pointed out.",
    "1908870": "Well, I have just tried it. I added the extra data (around 65 images with \"Other\" label). Despite low validation loss (**~ .820**) I got a huge log loss score on unseen test data. I think main issues with my approach are:\n- our training data became more imbalanced;\n- \"other\" class is just irrelevant for the main task of competition.",
    "1908746": "I think I read in another paper somewhere that the authors used extra data (similar to our \"other\" directory) for the purpose of training the encoder. I think this is supposed to give more data so that they encoder can better understand the histopathology data. Sorry for the lack of a source."
  }
}