{
  "id": 465368,
  "title": "38th solution (Private 0.52, Hight score 0.55)",
  "url": "/competitions/UBC-OCEAN/discussion/465368",
  "author_name": "devchopin",
  "post_date": "2024-01-04T01:59:06.025000",
  "votes": 14,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Since I first started learning about data analytics, I've heard about Kaggle from many people and have come to admire them. If this competition ends successfully, I will become a competition master two years after starting Kaggle! Thanks everyone!</p>\n<p>And 'Gunes Evitan''s pyvips code was very helpful during the competition. Thank you.<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/gunesevitan/libvips-pyvips-installation-and-getting-started</a></p>\n<h3>Preprocessing</h3>\n<p>After using the back ground provided by the competition, we applied the otsu threshold. And I cut the image to size 512 x 512 and saved it.</p>\n<h3>Training</h3>\n<ul>\n<li>Model : VIT-s + TransMIL</li>\n<li>Augmentation : VerticalFlip, HorizontalFlip, CLAHE, RandomGamma, GridDistortion, ShiftScaleRotate</li>\n<li>Optimizer &amp; learning rate: Since vis-s was already pre-trained and MIL was prone to overfitting, vit-s was trained with a learning rate of 1e-6 and MIL was trained at a learning rate of 1e-5, and EMA was applied to each. AdamW and CE were used.</li>\n</ul>\n<blockquote>\n  <p>optimizer = torch.optim.AdamW([{'params': model.image_extractor.parameters(),'lr':1e-6}, {'params': model.mil.parameters()}], lr=1e-5, weight_decay=1e-3)<br>\n  extractor_ema = ModelEma(model.image_extractor, decay=ema_decay, device=None, resume='')<br>\n  mil_ema = ModelEma(model.mil, decay=ema_decay, device=None, resume='')</p>\n</blockquote>\n<p>I experimented with two methods.</p>\n<ol>\n<li>Traning only MIL: A weakly supervised method that extracts features from patch images using the vit-s model and then learns using only those features.</li>\n<li>Training with image encoder (vit-s) together: We randomly selected 100 images from a 512x512 patch for learning and evaluation.</li>\n</ol>\n<p>Of the two, method 2 showed better pb score.</p>\n<h3>Tried(helpful)</h3>\n<ul>\n<li>Pseudo-labeling 1536x1536 : After pseudo labeling the 1536x1536 image using MIL learned at 512x512, we learned a model for TMA prediction using images with a probability of 0.5 or higher. Although it was not good in pb score, it achieved 0.55 in private.</li>\n<li>Outlier detect: Each class was learned using binary cross entropy. After applying sigmoid, if all class predictions were less than 0.5, it was predicted as 'Other'. It's not exact, but there was an increase of about 0.1.</li>\n<li>Upscaling : It was better than applying weights to cross entropy.</li>\n</ul>\n<h3>Tried(but didn't help)</h3>\n<ul>\n<li>staintools: augmentation with staintools. But it didn't help much.</li>\n<li>Other dataset(external) : <a href=\"url\" target=\"_blank\">https://www.cancerimagingarchive.net/collection/ovarian-bevacizumab-response/</a> In this dataset, I trained a model with the class corresponding to 'UC' as other, but it did not help at all.</li>\n</ul>",
  "messages": [
    {
      "id": 2586129,
      "postDate": "2024-01-04T01:59:06.027Z",
      "content": "<p>Since I first started learning about data analytics, I've heard about Kaggle from many people and have come to admire them. If this competition ends successfully, I will become a competition master two years after starting Kaggle! Thanks everyone!</p>\n<p>And 'Gunes Evitan''s pyvips code was very helpful during the competition. Thank you.<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/gunesevitan/libvips-pyvips-installation-and-getting-started</a></p>\n<h3>Preprocessing</h3>\n<p>After using the back ground provided by the competition, we applied the otsu threshold. And I cut the image to size 512 x 512 and saved it.</p>\n<h3>Training</h3>\n<ul>\n<li>Model : VIT-s + TransMIL</li>\n<li>Augmentation : VerticalFlip, HorizontalFlip, CLAHE, RandomGamma, GridDistortion, ShiftScaleRotate</li>\n<li>Optimizer &amp; learning rate: Since vis-s was already pre-trained and MIL was prone to overfitting, vit-s was trained with a learning rate of 1e-6 and MIL was trained at a learning rate of 1e-5, and EMA was applied to each. AdamW and CE were used.</li>\n</ul>\n<blockquote>\n  <p>optimizer = torch.optim.AdamW([{'params': model.image_extractor.parameters(),'lr':1e-6}, {'params': model.mil.parameters()}], lr=1e-5, weight_decay=1e-3)<br>\n  extractor_ema = ModelEma(model.image_extractor, decay=ema_decay, device=None, resume='')<br>\n  mil_ema = ModelEma(model.mil, decay=ema_decay, device=None, resume='')</p>\n</blockquote>\n<p>I experimented with two methods.</p>\n<ol>\n<li>Traning only MIL: A weakly supervised method that extracts features from patch images using the vit-s model and then learns using only those features.</li>\n<li>Training with image encoder (vit-s) together: We randomly selected 100 images from a 512x512 patch for learning and evaluation.</li>\n</ol>\n<p>Of the two, method 2 showed better pb score.</p>\n<h3>Tried(helpful)</h3>\n<ul>\n<li>Pseudo-labeling 1536x1536 : After pseudo labeling the 1536x1536 image using MIL learned at 512x512, we learned a model for TMA prediction using images with a probability of 0.5 or higher. Although it was not good in pb score, it achieved 0.55 in private.</li>\n<li>Outlier detect: Each class was learned using binary cross entropy. After applying sigmoid, if all class predictions were less than 0.5, it was predicted as 'Other'. It's not exact, but there was an increase of about 0.1.</li>\n<li>Upscaling : It was better than applying weights to cross entropy.</li>\n</ul>\n<h3>Tried(but didn't help)</h3>\n<ul>\n<li>staintools: augmentation with staintools. But it didn't help much.</li>\n<li>Other dataset(external) : <a href=\"url\" target=\"_blank\">https://www.cancerimagingarchive.net/collection/ovarian-bevacizumab-response/</a> In this dataset, I trained a model with the class corresponding to 'UC' as other, but it did not help at all.</li>\n</ul>",
      "rawMarkdown": "Since I first started learning about data analytics, I've heard about Kaggle from many people and have come to admire them. If this competition ends successfully, I will become a competition master two years after starting Kaggle! Thanks everyone!\n\nAnd 'Gunes Evitan''s pyvips code was very helpful during the competition. Thank you.\n[https://www.kaggle.com/code/gunesevitan/libvips-pyvips-installation-and-getting-started](url)\n\n### Preprocessing\nAfter using the back ground provided by the competition, we applied the otsu threshold. And I cut the image to size 512 x 512 and saved it.\n\n### Training\n- Model : VIT-s + TransMIL\n- Augmentation : VerticalFlip, HorizontalFlip, CLAHE, RandomGamma, GridDistortion, ShiftScaleRotate\n- Optimizer & learning rate: Since vis-s was already pre-trained and MIL was prone to overfitting, vit-s was trained with a learning rate of 1e-6 and MIL was trained at a learning rate of 1e-5, and EMA was applied to each. AdamW and CE were used.\n\n>optimizer = torch.optim.AdamW([{'params': model.image_extractor.parameters(),'lr':1e-6}, {'params': model.mil.parameters()}], lr=1e-5, weight_decay=1e-3)\nextractor_ema = ModelEma(model.image_extractor, decay=ema_decay, device=None, resume='')\nmil_ema = ModelEma(model.mil, decay=ema_decay, device=None, resume='')\n\nI experimented with two methods.\n\n1. Traning only MIL: A weakly supervised method that extracts features from patch images using the vit-s model and then learns using only those features.\n2. Training with image encoder (vit-s) together: We randomly selected 100 images from a 512x512 patch for learning and evaluation.\n\nOf the two, method 2 showed better pb score.\n\n### Tried(helpful)\n- Pseudo-labeling 1536x1536 : After pseudo labeling the 1536x1536 image using MIL learned at 512x512, we learned a model for TMA prediction using images with a probability of 0.5 or higher. Although it was not good in pb score, it achieved 0.55 in private.\n- Outlier detect: Each class was learned using binary cross entropy. After applying sigmoid, if all class predictions were less than 0.5, it was predicted as 'Other'. It's not exact, but there was an increase of about 0.1.\n- Upscaling : It was better than applying weights to cross entropy.\n\n### Tried(but didn't help)\n- staintools: augmentation with staintools. But it didn't help much.\n-  Other dataset(external) : [https://www.cancerimagingarchive.net/collection/ovarian-bevacizumab-response/](url) In this dataset, I trained a model with the class corresponding to 'UC' as other, but it did not help at all.",
      "votes": 14
    },
    {
      "id": 2586787,
      "postDate": "2024-01-04T12:06:17.690Z",
      "content": "<p>This si great learning; curious what implementation did you use fro TransMIL?</p>",
      "rawMarkdown": "This si great learning; curious what implementation did you use fro TransMIL?",
      "votes": 1
    },
    {
      "id": 2586540,
      "postDate": "2024-01-04T09:40:02.537Z",
      "content": "<p>Congrats silver medal and becoming on competition Master! I'm interesting about your method of outlier detection .</p>\n<p>\"Each class was learned using binary cross entropy.\"</p>\n<p>Could you explain more about this? I want to implement your method, which can be solid on this competition. How you use the losses when you train the models? I guess you make separated head for each class. </p>",
      "rawMarkdown": "Congrats silver medal and becoming on competition Master! I'm interesting about your method of outlier detection .\n\n\"Each class was learned using binary cross entropy.\"\n\nCould you explain more about this? I want to implement your method, which can be solid on this competition. How you use the losses when you train the models? I guess you make separated head for each class. \n",
      "votes": 1,
      "replies": [
        {
          "id": 2586698,
          "postDate": "2024-01-04T10:54:33.607Z",
          "content": "<p>The label was used by one-hot encoding.</p>\n<p>ex) HGSC -&gt; [1, 0, 0, 0, 0], EC -&gt; [0, 0, 1, 0, 0] -&gt; training with BCEWithLogitsLoss</p>\n<p>By doing this, unlike Cross Entropy, I thought I would be able to detect outliers because I could find the probability of each class alone, and there was an increase of about 0.1 points.</p>",
          "rawMarkdown": "The label was used by one-hot encoding.\n\nex) HGSC -> [1, 0, 0, 0, 0], EC -> [0, 0, 1, 0, 0] -> training with BCEWithLogitsLoss\n\nBy doing this, unlike Cross Entropy, I thought I would be able to detect outliers because I could find the probability of each class alone, and there was an increase of about 0.1 points.",
          "votes": 2,
          "replies": [
            {
              "id": 2586735,
              "postDate": "2024-01-04T11:23:52.653Z",
              "content": "<p>Thanks for the share, I totally understood. Simple but I think it's very meaningful.</p>\n<p>Actually I almost gave up to find reasonable method for detect outliers and did nothing until final submission. I'm going to use your method for late submission.</p>\n<p>See you next competition and happy new year!</p>",
              "rawMarkdown": "Thanks for the share, I totally understood. Simple but I think it's very meaningful.\n\nActually I almost gave up to find reasonable method for detect outliers and did nothing until final submission. I'm going to use your method for late submission.\n\nSee you next competition and happy new year!\n",
              "votes": 1
            },
            {
              "id": 2586738,
              "postDate": "2024-01-04T11:26:58.317Z",
              "content": "<p>That's a very interesting approach. Can you still argmax those probabilities and get their class?</p>",
              "rawMarkdown": "That's a very interesting approach. Can you still argmax those probabilities and get their class?",
              "votes": 1
            },
            {
              "id": 2586767,
              "postDate": "2024-01-04T11:51:34.550Z",
              "content": "<p>No, applying argmax to the BCE model reduced the positive detection ability in LGSC. I used two models</p>\n<ol>\n<li>Model trained with BCE for outlier detection</li>\n<li>This is a model learned with CE for class classification.</li>\n</ol>\n<p>If there were no classes exceeding the threshold (0.4~0.5) in the BCE model, it was passed on to the CE model to classify the correct subclass.</p>\n<p>(This didn't work for pb, but it worked for lb)</p>",
              "rawMarkdown": "No, applying argmax to the BCE model reduced the positive detection ability in LGSC. I used two models\n\n1. Model trained with BCE for outlier detection\n2. This is a model learned with CE for class classification.\n\nIf there were no classes exceeding the threshold (0.4~0.5) in the BCE model, it was passed on to the CE model to classify the correct subclass.\n\n(This didn't work for pb, but it worked for lb)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2586275,
      "postDate": "2024-01-04T05:22:41.860Z",
      "content": "<p><code>Training with image encoder (vit-s) together: We randomly selected 100 images from a 512x512 patch for learning and evaluation.</code></p>\n<p>My approach is quite similar to your Method 1. I'm intrigued by your Method 2. Can you explain how you train ViT-S and MIL together?</p>",
      "rawMarkdown": "`Training with image encoder (vit-s) together: We randomly selected 100 images from a 512x512 patch for learning and evaluation.`\n\nMy approach is quite similar to your Method 1. I'm intrigued by your Method 2. Can you explain how you train ViT-S and MIL together?",
      "votes": 1,
      "replies": [
        {
          "id": 2586287,
          "postDate": "2024-01-04T05:36:10.530Z",
          "content": "<pre><code> (nn.Module):\n     ():\n        self.image_extractor = ViT(size=)\n        self.mil = TransMIL()\n\n     ():\n        x = self.image_extractor(x)\n        x = self.mil(x)\n         x\n</code></pre>\n<p>This is a simple code for my model. </p>\n<p>The backpropagation of the loss is transmitted to both ViT, which extracted the image, and MIL, which makes the final prediction.</p>\n<p>However, it requires a lot of VRAM and there is an overfitting problem, so different learning rates and ema are essential.</p>",
          "rawMarkdown": "```python\n\nclass UBCModel(nn.Module):\n    def __init__(self):\n        self.image_extractor = ViT(size='small')\n        self.mil = TransMIL()\n\n    def forward(self, x):\n        x = self.image_extractor(x)\n        x = self.mil(x)\n        return x\n```\n\nThis is a simple code for my model. \n\nThe backpropagation of the loss is transmitted to both ViT, which extracted the image, and MIL, which makes the final prediction.\n\nHowever, it requires a lot of VRAM and there is an overfitting problem, so different learning rates and ema are essential.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2586227,
      "postDate": "2024-01-04T04:07:08.850Z",
      "content": "<p>Congratulations on getting the 40th rank in this competition. Thanks for sharing the solution details. </p>",
      "rawMarkdown": "Congratulations on getting the 40th rank in this competition. Thanks for sharing the solution details. ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2586787,
      "author_name": "Jirka",
      "author_url": "",
      "post_date": "2024-01-04T12:06:17.690000",
      "content": "<p>This si great learning; curious what implementation did you use fro TransMIL?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2586540,
      "author_name": "olivepicker",
      "author_url": "",
      "post_date": "2024-01-04T09:40:02.537000",
      "content": "<p>Congrats silver medal and becoming on competition Master! I'm interesting about your method of outlier detection .</p>\n<p>\"Each class was learned using binary cross entropy.\"</p>\n<p>Could you explain more about this? I want to implement your method, which can be solid on this competition. How you use the losses when you train the models? I guess you make separated head for each class. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2586698,
          "author_name": "devchopin",
          "author_url": "",
          "post_date": "2024-01-04T10:54:33.607000",
          "content": "<p>The label was used by one-hot encoding.</p>\n<p>ex) HGSC -&gt; [1, 0, 0, 0, 0], EC -&gt; [0, 0, 1, 0, 0] -&gt; training with BCEWithLogitsLoss</p>\n<p>By doing this, unlike Cross Entropy, I thought I would be able to detect outliers because I could find the probability of each class alone, and there was an increase of about 0.1 points.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2586735,
              "author_name": "olivepicker",
              "author_url": "",
              "post_date": "2024-01-04T11:23:52.653000",
              "content": "<p>Thanks for the share, I totally understood. Simple but I think it's very meaningful.</p>\n<p>Actually I almost gave up to find reasonable method for detect outliers and did nothing until final submission. I'm going to use your method for late submission.</p>\n<p>See you next competition and happy new year!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2586738,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2024-01-04T11:26:58.317000",
              "content": "<p>That's a very interesting approach. Can you still argmax those probabilities and get their class?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2586767,
              "author_name": "devchopin",
              "author_url": "",
              "post_date": "2024-01-04T11:51:34.550000",
              "content": "<p>No, applying argmax to the BCE model reduced the positive detection ability in LGSC. I used two models</p>\n<ol>\n<li>Model trained with BCE for outlier detection</li>\n<li>This is a model learned with CE for class classification.</li>\n</ol>\n<p>If there were no classes exceeding the threshold (0.4~0.5) in the BCE model, it was passed on to the CE model to classify the correct subclass.</p>\n<p>(This didn't work for pb, but it worked for lb)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2586275,
      "author_name": "Huang Jin Feng",
      "author_url": "",
      "post_date": "2024-01-04T05:22:41.860000",
      "content": "<p><code>Training with image encoder (vit-s) together: We randomly selected 100 images from a 512x512 patch for learning and evaluation.</code></p>\n<p>My approach is quite similar to your Method 1. I'm intrigued by your Method 2. Can you explain how you train ViT-S and MIL together?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2586287,
          "author_name": "devchopin",
          "author_url": "",
          "post_date": "2024-01-04T05:36:10.530000",
          "content": "<pre><code> (nn.Module):\n     ():\n        self.image_extractor = ViT(size=)\n        self.mil = TransMIL()\n\n     ():\n        x = self.image_extractor(x)\n        x = self.mil(x)\n         x\n</code></pre>\n<p>This is a simple code for my model. </p>\n<p>The backpropagation of the loss is transmitted to both ViT, which extracted the image, and MIL, which makes the final prediction.</p>\n<p>However, it requires a lot of VRAM and there is an overfitting problem, so different learning rates and ema are essential.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2586227,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2024-01-04T04:07:08.850000",
      "content": "<p>Congratulations on getting the 40th rank in this competition. Thanks for sharing the solution details. </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2586129": "Since I first started learning about data analytics, I've heard about Kaggle from many people and have come to admire them. If this competition ends successfully, I will become a competition master two years after starting Kaggle! Thanks everyone!\n\nAnd 'Gunes Evitan''s pyvips code was very helpful during the competition. Thank you.\n[https://www.kaggle.com/code/gunesevitan/libvips-pyvips-installation-and-getting-started](url)\n\n### Preprocessing\nAfter using the back ground provided by the competition, we applied the otsu threshold. And I cut the image to size 512 x 512 and saved it.\n\n### Training\n- Model : VIT-s + TransMIL\n- Augmentation : VerticalFlip, HorizontalFlip, CLAHE, RandomGamma, GridDistortion, ShiftScaleRotate\n- Optimizer & learning rate: Since vis-s was already pre-trained and MIL was prone to overfitting, vit-s was trained with a learning rate of 1e-6 and MIL was trained at a learning rate of 1e-5, and EMA was applied to each. AdamW and CE were used.\n\n>optimizer = torch.optim.AdamW([{'params': model.image_extractor.parameters(),'lr':1e-6}, {'params': model.mil.parameters()}], lr=1e-5, weight_decay=1e-3)\nextractor_ema = ModelEma(model.image_extractor, decay=ema_decay, device=None, resume='')\nmil_ema = ModelEma(model.mil, decay=ema_decay, device=None, resume='')\n\nI experimented with two methods.\n\n1. Traning only MIL: A weakly supervised method that extracts features from patch images using the vit-s model and then learns using only those features.\n2. Training with image encoder (vit-s) together: We randomly selected 100 images from a 512x512 patch for learning and evaluation.\n\nOf the two, method 2 showed better pb score.\n\n### Tried(helpful)\n- Pseudo-labeling 1536x1536 : After pseudo labeling the 1536x1536 image using MIL learned at 512x512, we learned a model for TMA prediction using images with a probability of 0.5 or higher. Although it was not good in pb score, it achieved 0.55 in private.\n- Outlier detect: Each class was learned using binary cross entropy. After applying sigmoid, if all class predictions were less than 0.5, it was predicted as 'Other'. It's not exact, but there was an increase of about 0.1.\n- Upscaling : It was better than applying weights to cross entropy.\n\n### Tried(but didn't help)\n- staintools: augmentation with staintools. But it didn't help much.\n-  Other dataset(external) : [https://www.cancerimagingarchive.net/collection/ovarian-bevacizumab-response/](url) In this dataset, I trained a model with the class corresponding to 'UC' as other, but it did not help at all.",
    "2586787": "This si great learning; curious what implementation did you use fro TransMIL?",
    "2586540": "Congrats silver medal and becoming on competition Master! I'm interesting about your method of outlier detection .\n\n\"Each class was learned using binary cross entropy.\"\n\nCould you explain more about this? I want to implement your method, which can be solid on this competition. How you use the losses when you train the models? I guess you make separated head for each class. \n",
    "2586275": "`Training with image encoder (vit-s) together: We randomly selected 100 images from a 512x512 patch for learning and evaluation.`\n\nMy approach is quite similar to your Method 1. I'm intrigued by your Method 2. Can you explain how you train ViT-S and MIL together?",
    "2586227": "Congratulations on getting the 40th rank in this competition. Thanks for sharing the solution details. "
  }
}