{
  "id": 465410,
  "title": "2nd Place Solution - UBC-OCEAN",
  "url": "/competitions/UBC-OCEAN/discussion/465410",
  "author_name": "zznznb",
  "post_date": "2024-01-04T06:26:53.910000",
  "votes": 43,
  "comment_count": 26,
  "views": 0,
  "content": "<h1>Preface</h1>\n<p>The most significant difficulty in whole slide image (WSI) classification is the extremely high resolution, which should have been experienced by all competitors. Although the organizers of the competition provided a data type difficult to process, fortunately, the resolution of the data is much lower than that of typical WSI. In this discussion, we will provide a detailed introduction to our method.</p>\n<h1>Overview</h1>\n<p>Following the commonly used methods in academia, we toke the following steps:</p>\n<ol>\n<li><strong>Crop</strong> an entire WSI into thousands of <strong>patches</strong>;</li>\n<li>Use extractors to <strong>extract the features</strong>;</li>\n<li>Train the <strong>MIL</strong> models.</li>\n</ol>\n<h1>External Data</h1>\n<p>We used two external data with labels. All competitors can download without payment. We found that although more external data and Other class were used for training, there was no significant improvement in scores. We believe this is due to quality issues with external data or significant differences from competition data. Just as some competitors can achieve high scores without using external data, we believe that the external data is not necessary in this competition.</p>\n<ul>\n<li><a href=\"https://wirtualnymikroskop.mostwiedzy.pl/list/\" target=\"_blank\">https://wirtualnymikroskop.mostwiedzy.pl/list/</a></li>\n<li><a href=\"https://www.cancerimagingarchive.net/collection/ptrc-hgsoc/\" target=\"_blank\">https://www.cancerimagingarchive.net/collection/ptrc-hgsoc/</a></li>\n</ul>\n<h1>Crop Patches and Extract Features</h1>\n<p>We create one Dataset for one WSI. Code is here:</p>\n<pre><code> ():\n     ():\n        ().__init__()\n        self.data_path = data_path\n        self.wsi_name = wsi_name\n        self.ratio = ratio\n         mode  [, ]\n        self.mode = mode\n        self.wsi = pyvips.Image.new_from_file(os.path.join(data_path, , wsi_name + ))\n        self.is_tma = self.wsi.height &lt;   self.wsi.width &lt; \n        self.patch_size = patch_size\n        self.transform = T.Compose([T.ToTensor(), T.Resize((, ), antialias=), T.Normalize(mean=[, , ], std=[, , ])])\n        self.cor_list = self.get_patch()\n\n     ():\n        cor_list = []\n         self.is_tma:\n            thumbnail = self.wsi\n        :\n            thumbnail = pyvips.Image.new_from_file(os.path.join(self.data_path, , self.wsi_name + ))\n        wsi_width, wsi_height = self.wsi.width, self.wsi.height\n        thu_width, thu_height = thumbnail.width, thumbnail.height\n        h_r, w_r = wsi_height / thu_height, wsi_width / thu_width\n        down_h, down_w = (self.patch_size / h_r), (self.patch_size / w_r)\n        cors = [(x, y)  y  (, thu_height, down_h)  x  (, thu_width, down_w)]\n         x, y  cors:\n            tile = thumbnail.crop(x, y, (down_w, thu_width - x), (down_h, thu_height - y)).numpy()[..., :]\n            black_bg = np.mean(tile, axis=) &lt; \n            tile[black_bg, :] = \n            mask_bg = np.mean(tile, axis=) &gt; \n             np.(mask_bg) &lt; (down_h, thu_height - y) * (down_w, thu_width - x) *   (cor_list) ==   self.is_tma:\n                cor_list.append(((x * w_r), (y * h_r)))\n         self.is_tma:\n             cor_list\n         self.wsi.height &lt;   self.wsi.width &lt; :\n            R_ratio = \n         self.wsi.height &lt;   self.wsi.width &lt; :\n            R_ratio = \n        :\n            R_ratio = \n        random.shuffle(cor_list)\n        cor_list = cor_list[:(((cor_list) * R_ratio), )]\n         cor_list\n\n     ():\n         (self.cor_list)\n\n     ():\n        x, y = self.cor_list[idx]\n        tile = self.wsi.crop(x, y, (self.patch_size, self.wsi.width - x), (self.patch_size, self.wsi.height - y)).numpy()[..., :]\n        tile = self.transform(tile)\n         tile\n</code></pre>\n<h1>Feature Extraction Model</h1>\n<p>We used <strong>dino_vit_small_patch16_200ep.torch</strong> and <strong>dino_vit_small_patch8_200ep.torch</strong>.</p>\n<ul>\n<li><a href=\"https://github.com/lunit-io/benchmark-ssl-pathology/releases/tag/pretrained-weights\" target=\"_blank\">https://github.com/lunit-io/benchmark-ssl-pathology/releases/tag/pretrained-weights</a></li>\n</ul>\n<h1>MIL Model</h1>\n<ul>\n<li>ABMIL</li>\n<li>DSMIL</li>\n<li>TransMIL</li>\n</ul>\n<h1>Codes</h1>\n<p>Simplified Version</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/zznznb/wsi-train\" target=\"_blank\">https://www.kaggle.com/code/zznznb/wsi-train</a></li>\n<li><a href=\"https://www.kaggle.com/code/zznznb/wsi-inference-public-0-6-private-0-58\" target=\"_blank\">https://www.kaggle.com/code/zznznb/wsi-inference-public-0-6-private-0-58</a></li>\n</ul>\n<p>Final Version</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hustzx/2nd-0-61-train-abmil-dsmil-transmil\" target=\"_blank\">https://www.kaggle.com/code/hustzx/2nd-0-61-train-abmil-dsmil-transmil</a></li>\n<li><a href=\"https://www.kaggle.com/code/hustzx/2nd-0-61-infernece-abmil-dsmil-transmil\" target=\"_blank\">https://www.kaggle.com/code/hustzx/2nd-0-61-infernece-abmil-dsmil-transmil</a></li>\n</ul>\n<p>Feature Extraction Codes</p>\n<ul>\n<li><a href=\"https://github.com/ZeningZeng/UBC-OCEAN\" target=\"_blank\">https://github.com/ZeningZeng/UBC-OCEAN</a></li>\n</ul>",
  "messages": [
    {
      "id": 2586351,
      "postDate": "2024-01-04T06:26:53.910Z",
      "content": "<h1>Preface</h1>\n<p>The most significant difficulty in whole slide image (WSI) classification is the extremely high resolution, which should have been experienced by all competitors. Although the organizers of the competition provided a data type difficult to process, fortunately, the resolution of the data is much lower than that of typical WSI. In this discussion, we will provide a detailed introduction to our method.</p>\n<h1>Overview</h1>\n<p>Following the commonly used methods in academia, we toke the following steps:</p>\n<ol>\n<li><strong>Crop</strong> an entire WSI into thousands of <strong>patches</strong>;</li>\n<li>Use extractors to <strong>extract the features</strong>;</li>\n<li>Train the <strong>MIL</strong> models.</li>\n</ol>\n<h1>External Data</h1>\n<p>We used two external data with labels. All competitors can download without payment. We found that although more external data and Other class were used for training, there was no significant improvement in scores. We believe this is due to quality issues with external data or significant differences from competition data. Just as some competitors can achieve high scores without using external data, we believe that the external data is not necessary in this competition.</p>\n<ul>\n<li><a href=\"https://wirtualnymikroskop.mostwiedzy.pl/list/\" target=\"_blank\">https://wirtualnymikroskop.mostwiedzy.pl/list/</a></li>\n<li><a href=\"https://www.cancerimagingarchive.net/collection/ptrc-hgsoc/\" target=\"_blank\">https://www.cancerimagingarchive.net/collection/ptrc-hgsoc/</a></li>\n</ul>\n<h1>Crop Patches and Extract Features</h1>\n<p>We create one Dataset for one WSI. Code is here:</p>\n<pre><code> ():\n     ():\n        ().__init__()\n        self.data_path = data_path\n        self.wsi_name = wsi_name\n        self.ratio = ratio\n         mode  [, ]\n        self.mode = mode\n        self.wsi = pyvips.Image.new_from_file(os.path.join(data_path, , wsi_name + ))\n        self.is_tma = self.wsi.height &lt;   self.wsi.width &lt; \n        self.patch_size = patch_size\n        self.transform = T.Compose([T.ToTensor(), T.Resize((, ), antialias=), T.Normalize(mean=[, , ], std=[, , ])])\n        self.cor_list = self.get_patch()\n\n     ():\n        cor_list = []\n         self.is_tma:\n            thumbnail = self.wsi\n        :\n            thumbnail = pyvips.Image.new_from_file(os.path.join(self.data_path, , self.wsi_name + ))\n        wsi_width, wsi_height = self.wsi.width, self.wsi.height\n        thu_width, thu_height = thumbnail.width, thumbnail.height\n        h_r, w_r = wsi_height / thu_height, wsi_width / thu_width\n        down_h, down_w = (self.patch_size / h_r), (self.patch_size / w_r)\n        cors = [(x, y)  y  (, thu_height, down_h)  x  (, thu_width, down_w)]\n         x, y  cors:\n            tile = thumbnail.crop(x, y, (down_w, thu_width - x), (down_h, thu_height - y)).numpy()[..., :]\n            black_bg = np.mean(tile, axis=) &lt; \n            tile[black_bg, :] = \n            mask_bg = np.mean(tile, axis=) &gt; \n             np.(mask_bg) &lt; (down_h, thu_height - y) * (down_w, thu_width - x) *   (cor_list) ==   self.is_tma:\n                cor_list.append(((x * w_r), (y * h_r)))\n         self.is_tma:\n             cor_list\n         self.wsi.height &lt;   self.wsi.width &lt; :\n            R_ratio = \n         self.wsi.height &lt;   self.wsi.width &lt; :\n            R_ratio = \n        :\n            R_ratio = \n        random.shuffle(cor_list)\n        cor_list = cor_list[:(((cor_list) * R_ratio), )]\n         cor_list\n\n     ():\n         (self.cor_list)\n\n     ():\n        x, y = self.cor_list[idx]\n        tile = self.wsi.crop(x, y, (self.patch_size, self.wsi.width - x), (self.patch_size, self.wsi.height - y)).numpy()[..., :]\n        tile = self.transform(tile)\n         tile\n</code></pre>\n<h1>Feature Extraction Model</h1>\n<p>We used <strong>dino_vit_small_patch16_200ep.torch</strong> and <strong>dino_vit_small_patch8_200ep.torch</strong>.</p>\n<ul>\n<li><a href=\"https://github.com/lunit-io/benchmark-ssl-pathology/releases/tag/pretrained-weights\" target=\"_blank\">https://github.com/lunit-io/benchmark-ssl-pathology/releases/tag/pretrained-weights</a></li>\n</ul>\n<h1>MIL Model</h1>\n<ul>\n<li>ABMIL</li>\n<li>DSMIL</li>\n<li>TransMIL</li>\n</ul>\n<h1>Codes</h1>\n<p>Simplified Version</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/zznznb/wsi-train\" target=\"_blank\">https://www.kaggle.com/code/zznznb/wsi-train</a></li>\n<li><a href=\"https://www.kaggle.com/code/zznznb/wsi-inference-public-0-6-private-0-58\" target=\"_blank\">https://www.kaggle.com/code/zznznb/wsi-inference-public-0-6-private-0-58</a></li>\n</ul>\n<p>Final Version</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/hustzx/2nd-0-61-train-abmil-dsmil-transmil\" target=\"_blank\">https://www.kaggle.com/code/hustzx/2nd-0-61-train-abmil-dsmil-transmil</a></li>\n<li><a href=\"https://www.kaggle.com/code/hustzx/2nd-0-61-infernece-abmil-dsmil-transmil\" target=\"_blank\">https://www.kaggle.com/code/hustzx/2nd-0-61-infernece-abmil-dsmil-transmil</a></li>\n</ul>\n<p>Feature Extraction Codes</p>\n<ul>\n<li><a href=\"https://github.com/ZeningZeng/UBC-OCEAN\" target=\"_blank\">https://github.com/ZeningZeng/UBC-OCEAN</a></li>\n</ul>",
      "rawMarkdown": "# Preface\nThe most significant difficulty in whole slide image (WSI) classification is the extremely high resolution, which should have been experienced by all competitors. Although the organizers of the competition provided a data type difficult to process, fortunately, the resolution of the data is much lower than that of typical WSI. In this discussion, we will provide a detailed introduction to our method.\n\n# Overview\nFollowing the commonly used methods in academia, we toke the following steps:\n1. **Crop** an entire WSI into thousands of **patches**;\n2. Use extractors to **extract the features**;\n3. Train the **MIL** models.\n\n# External Data\nWe used two external data with labels. All competitors can download without payment. We found that although more external data and Other class were used for training, there was no significant improvement in scores. We believe this is due to quality issues with external data or significant differences from competition data. Just as some competitors can achieve high scores without using external data, we believe that the external data is not necessary in this competition.\n- https://wirtualnymikroskop.mostwiedzy.pl/list/\n- https://www.cancerimagingarchive.net/collection/ptrc-hgsoc/\n\n# Crop Patches and Extract Features\nWe create one Dataset for one WSI. Code is here:\n```python\nclass SingleWSIDataset(Dataset):\n    def __init__(self, data_path: str, wsi_name: str, patch_size: int, mode: str):\n        super().__init__()\n        self.data_path = data_path\n        self.wsi_name = wsi_name\n        self.ratio = ratio\n        assert mode in ['train', 'test']\n        self.mode = mode\n        self.wsi = pyvips.Image.new_from_file(os.path.join(data_path, f'{mode}_images', wsi_name + '.png'))\n        self.is_tma = self.wsi.height < 5000 and self.wsi.width < 5000\n        self.patch_size = patch_size\n        self.transform = T.Compose([T.ToTensor(), T.Resize((224, 224), antialias=True), T.Normalize(mean=[0.2585, 0.2556, 0.2506], std=[0.229, 0.224, 0.225])])\n        self.cor_list = self.get_patch()\n\n    def get_patch(self):\n        cor_list = []\n        if self.is_tma:\n            thumbnail = self.wsi\n        else:\n            thumbnail = pyvips.Image.new_from_file(os.path.join(self.data_path, f'{self.mode}_thumbnails', self.wsi_name + '_thumbnail.png'))\n        wsi_width, wsi_height = self.wsi.width, self.wsi.height\n        thu_width, thu_height = thumbnail.width, thumbnail.height\n        h_r, w_r = wsi_height / thu_height, wsi_width / thu_width\n        down_h, down_w = int(self.patch_size / h_r), int(self.patch_size / w_r)\n        cors = [(x, y) for y in range(0, thu_height, down_h) for x in range(0, thu_width, down_w)]\n        for x, y in cors:\n            tile = thumbnail.crop(x, y, min(down_w, thu_width - x), min(down_h, thu_height - y)).numpy()[..., :3]\n            black_bg = np.mean(tile, axis=2) < 20\n            tile[black_bg, :] = 255\n            mask_bg = np.mean(tile, axis=2) > 235\n            if np.sum(mask_bg) < min(down_h, thu_height - y) * min(down_w, thu_width - x) * 0.7 or len(cor_list) == 0 or self.is_tma:\n                cor_list.append((int(x * w_r), int(y * h_r)))\n        if self.is_tma:\n            return cor_list\n        if self.wsi.height < 40000 and self.wsi.width < 40000:\n            R_ratio = 0.8\n        elif self.wsi.height < 80000 and self.wsi.width < 80000:\n            R_ratio = 0.6\n        else:\n            R_ratio = 0.5\n        random.shuffle(cor_list)\n        cor_list = cor_list[:max(int(len(cor_list) * R_ratio), 1)]\n        return cor_list\n\n    def __len__(self):\n        return len(self.cor_list)\n\n    def __getitem__(self, idx):\n        x, y = self.cor_list[idx]\n        tile = self.wsi.crop(x, y, min(self.patch_size, self.wsi.width - x), min(self.patch_size, self.wsi.height - y)).numpy()[..., :3]\n        tile = self.transform(tile)\n        return tile\n```\n# Feature Extraction Model\nWe used **dino_vit_small_patch16_200ep.torch** and **dino_vit_small_patch8_200ep.torch**.\n- https://github.com/lunit-io/benchmark-ssl-pathology/releases/tag/pretrained-weights\n# MIL Model\n- ABMIL\n- DSMIL\n- TransMIL\n# Codes\nSimplified Version\n- https://www.kaggle.com/code/zznznb/wsi-train\n- https://www.kaggle.com/code/zznznb/wsi-inference-public-0-6-private-0-58\n\nFinal Version\n- https://www.kaggle.com/code/hustzx/2nd-0-61-train-abmil-dsmil-transmil\n- https://www.kaggle.com/code/hustzx/2nd-0-61-infernece-abmil-dsmil-transmil\n\nFeature Extraction Codes\n- https://github.com/ZeningZeng/UBC-OCEAN",
      "votes": 43
    },
    {
      "id": 2586626,
      "postDate": "2024-01-04T10:26:45.637Z",
      "content": "<p>I'm sorry for the mistake I made when inserting the website address before. It has been corrected and all can be opened.🙂</p>",
      "rawMarkdown": "I'm sorry for the mistake I made when inserting the website address before. It has been corrected and all can be opened.🙂",
      "votes": 6
    },
    {
      "id": 2591091,
      "postDate": "2024-01-07T16:55:37.183Z",
      "content": "<p>Congratulations!<br>\nCurious about the dataset from Poland (<a href=\"https://wirtualnymikroskop.mostwiedzy.pl/list/)\" target=\"_blank\">https://wirtualnymikroskop.mostwiedzy.pl/list/)</a>. Did you download the files manually or create a script for that?</p>",
      "rawMarkdown": "Congratulations!\nCurious about the dataset from Poland (https://wirtualnymikroskop.mostwiedzy.pl/list/). Did you download the files manually or create a script for that?",
      "votes": 1,
      "replies": [
        {
          "id": 2591470,
          "postDate": "2024-01-08T02:02:20.283Z",
          "content": "<p>Just manually. So in the end, I didn't get too many samples from here.😂</p>",
          "rawMarkdown": "Just manually. So in the end, I didn't get too many samples from here.😂",
          "replies": [
            {
              "id": 3159923,
              "postDate": "2025-03-26T04:43:16.730Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2586475,
      "postDate": "2024-01-04T08:44:51.557Z",
      "content": "<p>Sweet, I was thinking about the Multi-Instance Learning all the time but just did not have time to make it…<br>\nDid you train on the thumbnails, patches, or random crops from WSI?</p>\n<p>btw, the last link is broken URL seems fine, but after clicking it goes nowhere path</p>",
      "rawMarkdown": "Sweet, I was thinking about the Multi-Instance Learning all the time but just did not have time to make it...\nDid you train on the thumbnails, patches, or random crops from WSI?\n\n\nbtw, the last link is broken URL seems fine, but after clicking it goes nowhere path",
      "votes": 1,
      "replies": [
        {
          "id": 2586496,
          "postDate": "2024-01-04T08:56:26.473Z",
          "content": "<p>As shown in my code above, firstly filter out the coordinates of effective patches on the thumbnails, then crop and extract features on the images, and finally train MIL.</p>",
          "rawMarkdown": "As shown in my code above, firstly filter out the coordinates of effective patches on the thumbnails, then crop and extract features on the images, and finally train MIL.",
          "votes": 1
        },
        {
          "id": 2589043,
          "postDate": "2024-01-06T02:37:23.423Z",
          "content": "<p>We have tried CLAM and TransMIL and obtained low scores. Initially, we thought it was due to our model, but it turned out to be the feature extraction part😂</p>",
          "rawMarkdown": "We have tried CLAM and TransMIL and obtained low scores. Initially, we thought it was due to our model, but it turned out to be the feature extraction part😂",
          "votes": 2
        }
      ]
    },
    {
      "id": 2590342,
      "postDate": "2024-01-07T06:23:15.257Z",
      "content": "<p>Thank you all for your attention. I provide a simplified version of the codes that is close to our final score. I hope they are helpful to you.</p>",
      "rawMarkdown": "Thank you all for your attention. I provide a simplified version of the codes that is close to our final score. I hope they are helpful to you.",
      "votes": 2
    },
    {
      "id": 2586591,
      "postDate": "2024-01-04T10:10:27.120Z",
      "content": "<p>Congratulations on the 2nd position. This is first time I am hearing about Multi-Instance Learning and learning about it.<br>\nCompetitions are all about learning new stuff. If possible will you be able to put the whole training notebook, so that we can fork it.</p>\n<p>And also I can't open the external data links and inference notebook links.</p>",
      "rawMarkdown": "Congratulations on the 2nd position. This is first time I am hearing about Multi-Instance Learning and learning about it.\nCompetitions are all about learning new stuff. If possible will you be able to put the whole training notebook, so that we can fork it.\n\nAnd also I can't open the external data links and inference notebook links.",
      "votes": 2,
      "replies": [
        {
          "id": 2586645,
          "postDate": "2024-01-04T10:32:37.093Z",
          "content": "<p>Thank you for your attention. We will release all the code and data in a few days. And all links have been corrected.</p>",
          "rawMarkdown": "Thank you for your attention. We will release all the code and data in a few days. And all links have been corrected.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2619022,
      "postDate": "2024-01-25T06:53:39.303Z",
      "content": "<p>Hi! I want to understand how you worked with the signs and I'm trying to read your code to get the same data, you have \"extrain.csv\" written. How do I understand? Where can I get these files? So that I can build the signs from beginning to end the same way as you, then apply a model to them?</p>",
      "rawMarkdown": "Hi! I want to understand how you worked with the signs and I'm trying to read your code to get the same data, you have \"extrain.csv\" written. How do I understand? Where can I get these files? So that I can build the signs from beginning to end the same way as you, then apply a model to them?",
      "replies": [
        {
          "id": 2619244,
          "postDate": "2024-01-25T10:49:08.773Z",
          "content": "<p>Please see <strong>\"beifen.csv\"</strong> at <a href=\"https://www.kaggle.com/datasets/zznznb/checkpoints\" target=\"_blank\">https://www.kaggle.com/datasets/zznznb/checkpoints</a>. \"extrain.csv\" is a subset of \"beifen.csv\". It contains the WSI information, including external data. Generating codes can be found at <a href=\"https://github.com/ZeningZeng/UBC-OCEAN\" target=\"_blank\">https://github.com/ZeningZeng/UBC-OCEAN</a>.</p>",
          "rawMarkdown": "Please see **\"beifen.csv\"** at https://www.kaggle.com/datasets/zznznb/checkpoints. \"extrain.csv\" is a subset of \"beifen.csv\". It contains the WSI information, including external data. Generating codes can be found at https://github.com/ZeningZeng/UBC-OCEAN."
        }
      ]
    },
    {
      "id": 2590181,
      "postDate": "2024-01-07T00:55:49.630Z",
      "content": "<p>Congratulations！ Could you fix the train notebook link ？😁</p>",
      "rawMarkdown": "Congratulations！ Could you fix the train notebook link ？😁"
    },
    {
      "id": 2589676,
      "postDate": "2024-01-06T14:54:13.060Z",
      "content": "<p>Congratulations!<br>\nI'm just wondering, what would happen if your model would predict 'Other' when there were multiple classes in the image? For example, if the confidence of HGSC was 50%, EC was 49%, and the rest were minimal, then the output would be 'Other'</p>",
      "rawMarkdown": "Congratulations!\nI'm just wondering, what would happen if your model would predict 'Other' when there were multiple classes in the image? For example, if the confidence of HGSC was 50%, EC was 49%, and the rest were minimal, then the output would be 'Other'",
      "replies": [
        {
          "id": 2589727,
          "postDate": "2024-01-06T15:51:10.477Z",
          "content": "<p>We collected some samples that did not belong to the original five classes as independent Other class, so we only selected the one with the highest score as the prediction result. As for multiple classes in the image, I think it depends on which patches the model considers to have more prominent features.</p>",
          "rawMarkdown": "We collected some samples that did not belong to the original five classes as independent Other class, so we only selected the one with the highest score as the prediction result. As for multiple classes in the image, I think it depends on which patches the model considers to have more prominent features.",
          "replies": [
            {
              "id": 2605195,
              "postDate": "2024-01-16T23:14:58.487Z",
              "content": "<p>Looking at your submission notebook, I modified it to get information about the output confidence percentage for each class. I noticed that when it predicts a class, the confidence value is about 25-30% while the other classes are at about 10%. So I modified the code and added something like this:<br>\n<code>if (max_probability &gt; 0.20):\n            pred_label = max_class\n        else:\n            pred_label = 'Other'\n</code><br>\nThe resulting private score was 62%. </p>",
              "rawMarkdown": "Looking at your submission notebook, I modified it to get information about the output confidence percentage for each class. I noticed that when it predicts a class, the confidence value is about 25-30% while the other classes are at about 10%. So I modified the code and added something like this:\n` if (max_probability > 0.20):\n            pred_label = max_class\n        else:\n            pred_label = 'Other'\n`\nThe resulting private score was 62%. "
            },
            {
              "id": 2605268,
              "postDate": "2024-01-17T01:47:24.853Z",
              "content": "<p>Thank you for your information. We have tried this method before, but it cannot be determined whether it is effective in public score during the competition😄</p>",
              "rawMarkdown": "Thank you for your information. We have tried this method before, but it cannot be determined whether it is effective in public score during the competition😄"
            }
          ]
        }
      ]
    },
    {
      "id": 2588812,
      "postDate": "2024-01-05T18:33:38.483Z",
      "content": "<p>great work</p>",
      "rawMarkdown": "great work"
    },
    {
      "id": 2588270,
      "postDate": "2024-01-05T11:50:22.930Z",
      "content": "<p>Congratulations!!! Thanks for sharing details of your project.</p>",
      "rawMarkdown": "Congratulations!!! Thanks for sharing details of your project.\n"
    },
    {
      "id": 2586954,
      "postDate": "2024-01-04T13:45:10.547Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations"
    },
    {
      "id": 2586791,
      "postDate": "2024-01-04T12:07:25.770Z",
      "content": "<p>Congratulations for the 2nd place in this competition. I am waiting for your training solution. This is my first competition on kaggle and learning MIL approach is good for me.</p>",
      "rawMarkdown": "Congratulations for the 2nd place in this competition. I am waiting for your training solution. This is my first competition on kaggle and learning MIL approach is good for me."
    },
    {
      "id": 2586472,
      "postDate": "2024-01-04T08:44:24.077Z",
      "content": "<p>Hi, congratulations, nice and clean solution. About <code>R_ratio</code>, why did you make it different? Was it to overcome memory issue and speed inference or not?</p>",
      "rawMarkdown": "Hi, congratulations, nice and clean solution. About `R_ratio`, why did you make it different? Was it to overcome memory issue and speed inference or not?",
      "replies": [
        {
          "id": 2586508,
          "postDate": "2024-01-04T09:07:40.293Z",
          "content": "<p>Thanks🥳. This is indeed due to time limit. TMA has less than a hundred patches, and of course, all of them can be used for prediction. Therefore, discarding patches from some large images is always better than discarding patches from all images at a fixed ratio. Moreover, TMA accounts for a considerable proportion in private data.</p>",
          "rawMarkdown": "Thanks🥳. This is indeed due to time limit. TMA has less than a hundred patches, and of course, all of them can be used for prediction. Therefore, discarding patches from some large images is always better than discarding patches from all images at a fixed ratio. Moreover, TMA accounts for a considerable proportion in private data.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2586355,
      "postDate": "2024-01-04T06:31:00.123Z",
      "content": "<p>Congratulations on achieving second position in this prestigious competition.  Thanks for sharing details of your model.  Did the external data have \"Other\" class label data?</p>",
      "rawMarkdown": "Congratulations on achieving second position in this prestigious competition.  Thanks for sharing details of your model.  Did the external data have \"Other\" class label data?\n",
      "replies": [
        {
          "id": 2586518,
          "postDate": "2024-01-04T09:14:15.403Z",
          "content": "<p>It really has some data that does not belong to the five classes. However, there is almost no significant improvement compared to not using the Other class.😂</p>",
          "rawMarkdown": "It really has some data that does not belong to the five classes. However, there is almost no significant improvement compared to not using the Other class.😂",
          "votes": 1
        }
      ]
    },
    {
      "id": 2586398,
      "postDate": "2024-01-04T07:16:49.937Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2586626,
      "author_name": "zznznb",
      "author_url": "",
      "post_date": "2024-01-04T10:26:45.637000",
      "content": "<p>I'm sorry for the mistake I made when inserting the website address before. It has been corrected and all can be opened.🙂</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2591091,
      "author_name": "agape",
      "author_url": "",
      "post_date": "2024-01-07T16:55:37.183000",
      "content": "<p>Congratulations!<br>\nCurious about the dataset from Poland (<a href=\"https://wirtualnymikroskop.mostwiedzy.pl/list/)\" target=\"_blank\">https://wirtualnymikroskop.mostwiedzy.pl/list/)</a>. Did you download the files manually or create a script for that?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2591470,
          "author_name": "zznznb",
          "author_url": "",
          "post_date": "2024-01-08T02:02:20.283000",
          "content": "<p>Just manually. So in the end, I didn't get too many samples from here.😂</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3159923,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-26T04:43:16.730000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2586475,
      "author_name": "Jirka",
      "author_url": "",
      "post_date": "2024-01-04T08:44:51.557000",
      "content": "<p>Sweet, I was thinking about the Multi-Instance Learning all the time but just did not have time to make it…<br>\nDid you train on the thumbnails, patches, or random crops from WSI?</p>\n<p>btw, the last link is broken URL seems fine, but after clicking it goes nowhere path</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2586496,
          "author_name": "zznznb",
          "author_url": "",
          "post_date": "2024-01-04T08:56:26.473000",
          "content": "<p>As shown in my code above, firstly filter out the coordinates of effective patches on the thumbnails, then crop and extract features on the images, and finally train MIL.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2589043,
          "author_name": "Seeing Times",
          "author_url": "",
          "post_date": "2024-01-06T02:37:23.423000",
          "content": "<p>We have tried CLAM and TransMIL and obtained low scores. Initially, we thought it was due to our model, but it turned out to be the feature extraction part😂</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2590342,
      "author_name": "zznznb",
      "author_url": "",
      "post_date": "2024-01-07T06:23:15.257000",
      "content": "<p>Thank you all for your attention. I provide a simplified version of the codes that is close to our final score. I hope they are helpful to you.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2586591,
      "author_name": "Pritam Sinha",
      "author_url": "",
      "post_date": "2024-01-04T10:10:27.120000",
      "content": "<p>Congratulations on the 2nd position. This is first time I am hearing about Multi-Instance Learning and learning about it.<br>\nCompetitions are all about learning new stuff. If possible will you be able to put the whole training notebook, so that we can fork it.</p>\n<p>And also I can't open the external data links and inference notebook links.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2586645,
          "author_name": "zznznb",
          "author_url": "",
          "post_date": "2024-01-04T10:32:37.093000",
          "content": "<p>Thank you for your attention. We will release all the code and data in a few days. And all links have been corrected.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2619022,
      "author_name": "Rinat",
      "author_url": "",
      "post_date": "2024-01-25T06:53:39.303000",
      "content": "<p>Hi! I want to understand how you worked with the signs and I'm trying to read your code to get the same data, you have \"extrain.csv\" written. How do I understand? Where can I get these files? So that I can build the signs from beginning to end the same way as you, then apply a model to them?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2619244,
          "author_name": "zznznb",
          "author_url": "",
          "post_date": "2024-01-25T10:49:08.773000",
          "content": "<p>Please see <strong>\"beifen.csv\"</strong> at <a href=\"https://www.kaggle.com/datasets/zznznb/checkpoints\" target=\"_blank\">https://www.kaggle.com/datasets/zznznb/checkpoints</a>. \"extrain.csv\" is a subset of \"beifen.csv\". It contains the WSI information, including external data. Generating codes can be found at <a href=\"https://github.com/ZeningZeng/UBC-OCEAN\" target=\"_blank\">https://github.com/ZeningZeng/UBC-OCEAN</a>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2590181,
      "author_name": "Seeing Times",
      "author_url": "",
      "post_date": "2024-01-07T00:55:49.630000",
      "content": "<p>Congratulations！ Could you fix the train notebook link ？😁</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2589676,
      "author_name": "Noli Alonso",
      "author_url": "",
      "post_date": "2024-01-06T14:54:13.060000",
      "content": "<p>Congratulations!<br>\nI'm just wondering, what would happen if your model would predict 'Other' when there were multiple classes in the image? For example, if the confidence of HGSC was 50%, EC was 49%, and the rest were minimal, then the output would be 'Other'</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2589727,
          "author_name": "zznznb",
          "author_url": "",
          "post_date": "2024-01-06T15:51:10.477000",
          "content": "<p>We collected some samples that did not belong to the original five classes as independent Other class, so we only selected the one with the highest score as the prediction result. As for multiple classes in the image, I think it depends on which patches the model considers to have more prominent features.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2605195,
              "author_name": "Noli Alonso",
              "author_url": "",
              "post_date": "2024-01-16T23:14:58.487000",
              "content": "<p>Looking at your submission notebook, I modified it to get information about the output confidence percentage for each class. I noticed that when it predicts a class, the confidence value is about 25-30% while the other classes are at about 10%. So I modified the code and added something like this:<br>\n<code>if (max_probability &gt; 0.20):\n            pred_label = max_class\n        else:\n            pred_label = 'Other'\n</code><br>\nThe resulting private score was 62%. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2605268,
              "author_name": "zznznb",
              "author_url": "",
              "post_date": "2024-01-17T01:47:24.853000",
              "content": "<p>Thank you for your information. We have tried this method before, but it cannot be determined whether it is effective in public score during the competition😄</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2588812,
      "author_name": "Md Nazrul Islam",
      "author_url": "",
      "post_date": "2024-01-05T18:33:38.483000",
      "content": "<p>great work</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2588270,
      "author_name": "Rishabh0517",
      "author_url": "",
      "post_date": "2024-01-05T11:50:22.930000",
      "content": "<p>Congratulations!!! Thanks for sharing details of your project.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2586954,
      "author_name": "Jiaxuan Zhang",
      "author_url": "",
      "post_date": "2024-01-04T13:45:10.547000",
      "content": "<p>Congratulations</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2586791,
      "author_name": "Huma Perveen",
      "author_url": "",
      "post_date": "2024-01-04T12:07:25.770000",
      "content": "<p>Congratulations for the 2nd place in this competition. I am waiting for your training solution. This is my first competition on kaggle and learning MIL approach is good for me.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2586472,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2024-01-04T08:44:24.077000",
      "content": "<p>Hi, congratulations, nice and clean solution. About <code>R_ratio</code>, why did you make it different? Was it to overcome memory issue and speed inference or not?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2586508,
          "author_name": "zznznb",
          "author_url": "",
          "post_date": "2024-01-04T09:07:40.293000",
          "content": "<p>Thanks🥳. This is indeed due to time limit. TMA has less than a hundred patches, and of course, all of them can be used for prediction. Therefore, discarding patches from some large images is always better than discarding patches from all images at a fixed ratio. Moreover, TMA accounts for a considerable proportion in private data.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2586355,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2024-01-04T06:31:00.123000",
      "content": "<p>Congratulations on achieving second position in this prestigious competition.  Thanks for sharing details of your model.  Did the external data have \"Other\" class label data?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2586518,
          "author_name": "zznznb",
          "author_url": "",
          "post_date": "2024-01-04T09:14:15.403000",
          "content": "<p>It really has some data that does not belong to the five classes. However, there is almost no significant improvement compared to not using the Other class.😂</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2586398,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-04T07:16:49.937000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2586351": "# Preface\nThe most significant difficulty in whole slide image (WSI) classification is the extremely high resolution, which should have been experienced by all competitors. Although the organizers of the competition provided a data type difficult to process, fortunately, the resolution of the data is much lower than that of typical WSI. In this discussion, we will provide a detailed introduction to our method.\n\n# Overview\nFollowing the commonly used methods in academia, we toke the following steps:\n1. **Crop** an entire WSI into thousands of **patches**;\n2. Use extractors to **extract the features**;\n3. Train the **MIL** models.\n\n# External Data\nWe used two external data with labels. All competitors can download without payment. We found that although more external data and Other class were used for training, there was no significant improvement in scores. We believe this is due to quality issues with external data or significant differences from competition data. Just as some competitors can achieve high scores without using external data, we believe that the external data is not necessary in this competition.\n- https://wirtualnymikroskop.mostwiedzy.pl/list/\n- https://www.cancerimagingarchive.net/collection/ptrc-hgsoc/\n\n# Crop Patches and Extract Features\nWe create one Dataset for one WSI. Code is here:\n```python\nclass SingleWSIDataset(Dataset):\n    def __init__(self, data_path: str, wsi_name: str, patch_size: int, mode: str):\n        super().__init__()\n        self.data_path = data_path\n        self.wsi_name = wsi_name\n        self.ratio = ratio\n        assert mode in ['train', 'test']\n        self.mode = mode\n        self.wsi = pyvips.Image.new_from_file(os.path.join(data_path, f'{mode}_images', wsi_name + '.png'))\n        self.is_tma = self.wsi.height < 5000 and self.wsi.width < 5000\n        self.patch_size = patch_size\n        self.transform = T.Compose([T.ToTensor(), T.Resize((224, 224), antialias=True), T.Normalize(mean=[0.2585, 0.2556, 0.2506], std=[0.229, 0.224, 0.225])])\n        self.cor_list = self.get_patch()\n\n    def get_patch(self):\n        cor_list = []\n        if self.is_tma:\n            thumbnail = self.wsi\n        else:\n            thumbnail = pyvips.Image.new_from_file(os.path.join(self.data_path, f'{self.mode}_thumbnails', self.wsi_name + '_thumbnail.png'))\n        wsi_width, wsi_height = self.wsi.width, self.wsi.height\n        thu_width, thu_height = thumbnail.width, thumbnail.height\n        h_r, w_r = wsi_height / thu_height, wsi_width / thu_width\n        down_h, down_w = int(self.patch_size / h_r), int(self.patch_size / w_r)\n        cors = [(x, y) for y in range(0, thu_height, down_h) for x in range(0, thu_width, down_w)]\n        for x, y in cors:\n            tile = thumbnail.crop(x, y, min(down_w, thu_width - x), min(down_h, thu_height - y)).numpy()[..., :3]\n            black_bg = np.mean(tile, axis=2) < 20\n            tile[black_bg, :] = 255\n            mask_bg = np.mean(tile, axis=2) > 235\n            if np.sum(mask_bg) < min(down_h, thu_height - y) * min(down_w, thu_width - x) * 0.7 or len(cor_list) == 0 or self.is_tma:\n                cor_list.append((int(x * w_r), int(y * h_r)))\n        if self.is_tma:\n            return cor_list\n        if self.wsi.height < 40000 and self.wsi.width < 40000:\n            R_ratio = 0.8\n        elif self.wsi.height < 80000 and self.wsi.width < 80000:\n            R_ratio = 0.6\n        else:\n            R_ratio = 0.5\n        random.shuffle(cor_list)\n        cor_list = cor_list[:max(int(len(cor_list) * R_ratio), 1)]\n        return cor_list\n\n    def __len__(self):\n        return len(self.cor_list)\n\n    def __getitem__(self, idx):\n        x, y = self.cor_list[idx]\n        tile = self.wsi.crop(x, y, min(self.patch_size, self.wsi.width - x), min(self.patch_size, self.wsi.height - y)).numpy()[..., :3]\n        tile = self.transform(tile)\n        return tile\n```\n# Feature Extraction Model\nWe used **dino_vit_small_patch16_200ep.torch** and **dino_vit_small_patch8_200ep.torch**.\n- https://github.com/lunit-io/benchmark-ssl-pathology/releases/tag/pretrained-weights\n# MIL Model\n- ABMIL\n- DSMIL\n- TransMIL\n# Codes\nSimplified Version\n- https://www.kaggle.com/code/zznznb/wsi-train\n- https://www.kaggle.com/code/zznznb/wsi-inference-public-0-6-private-0-58\n\nFinal Version\n- https://www.kaggle.com/code/hustzx/2nd-0-61-train-abmil-dsmil-transmil\n- https://www.kaggle.com/code/hustzx/2nd-0-61-infernece-abmil-dsmil-transmil\n\nFeature Extraction Codes\n- https://github.com/ZeningZeng/UBC-OCEAN",
    "2586626": "I'm sorry for the mistake I made when inserting the website address before. It has been corrected and all can be opened.🙂",
    "2591091": "Congratulations!\nCurious about the dataset from Poland (https://wirtualnymikroskop.mostwiedzy.pl/list/). Did you download the files manually or create a script for that?",
    "2586475": "Sweet, I was thinking about the Multi-Instance Learning all the time but just did not have time to make it...\nDid you train on the thumbnails, patches, or random crops from WSI?\n\n\nbtw, the last link is broken URL seems fine, but after clicking it goes nowhere path",
    "2590342": "Thank you all for your attention. I provide a simplified version of the codes that is close to our final score. I hope they are helpful to you.",
    "2586591": "Congratulations on the 2nd position. This is first time I am hearing about Multi-Instance Learning and learning about it.\nCompetitions are all about learning new stuff. If possible will you be able to put the whole training notebook, so that we can fork it.\n\nAnd also I can't open the external data links and inference notebook links.",
    "2619022": "Hi! I want to understand how you worked with the signs and I'm trying to read your code to get the same data, you have \"extrain.csv\" written. How do I understand? Where can I get these files? So that I can build the signs from beginning to end the same way as you, then apply a model to them?",
    "2590181": "Congratulations！ Could you fix the train notebook link ？😁",
    "2589676": "Congratulations!\nI'm just wondering, what would happen if your model would predict 'Other' when there were multiple classes in the image? For example, if the confidence of HGSC was 50%, EC was 49%, and the rest were minimal, then the output would be 'Other'",
    "2588812": "great work",
    "2588270": "Congratulations!!! Thanks for sharing details of your project.\n",
    "2586954": "Congratulations",
    "2586791": "Congratulations for the 2nd place in this competition. I am waiting for your training solution. This is my first competition on kaggle and learning MIL approach is good for me.",
    "2586472": "Hi, congratulations, nice and clean solution. About `R_ratio`, why did you make it different? Was it to overcome memory issue and speed inference or not?",
    "2586355": "Congratulations on achieving second position in this prestigious competition.  Thanks for sharing details of your model.  Did the external data have \"Other\" class label data?\n",
    "2586398": ""
  }
}