{
  "id": 465697,
  "title": "7th place solution",
  "url": "/competitions/UBC-OCEAN/discussion/465697",
  "author_name": "m1dsolo",
  "post_date": "2024-01-05T09:37:04.807000",
  "votes": 36,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Thanks to kaggle and UBC for hosting this interesting competition and congrats to all the winners for their hard work! I would also like to thank my teammates and everyone in the discussion forum for their help!</p>\n<h1>Method</h1>\n<h2>Summary</h2>\n<p>Our final solution is based on multiple instance learning(MIL) for <strong>ovarian cancer subtype classification</strong> and use <code>sigmoid</code> and thresholding for <strong>outlier detection</strong>.<br>\nWe did not use mask annotation and additional datasets in the final submission.</p>\n<h3>1. preprocess</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15788664%2F16e938db8ccddff048fef4f4b9f306a3%2FUBC-OCEAN-1.jpeg?generation=1704447344840948&amp;alt=media\"></p>\n<p>1. Use <code>pyvips</code> to speed up png reading speed. (Thanks for <a href=\"https://www.kaggle.com/code/gunesevitan/libvips-pyvips-installation-and-getting-started\" target=\"_blank\">GUNES EVITAN's pyvips notebook</a>.)</p>\n<pre><code>image = pyvips.Image.new_from_file(image_id, access=).numpy()\nis_tma = image.shape[] &lt;=   image.shape[] &lt;= \n</code></pre>\n<p>2. Downsample WSI and TMA from x20 and x40 to x10 respectively. (Maybe x20 results will be better, but I can't submit due to resource constraints.)</p>\n<pre><code> is_tma:\n    resize = A.Resize(image.shape[] // , image.shape[] // )\n:\n    resize = A.Resize(image.shape[] // , image.shape[] // )\nimage = resize(image=image)[]\n</code></pre>\n<p>3. Deduplicate the identical tissue areas for WSI. (I'm not sure if this contributed to the results, but it saved me a lot of local memory.)</p>\n<pre><code> ():\n    image = image.astype(np.float16)\n    image = (image[..., ] *  + image[..., ] *  + image[..., ] * ) / \n     image.astype(np.uint8)\n\n  is_tma:\n    resize = A.Resize(image.shape[] // , image.shape[] // )\n    thumbnail = resize(image=image)[].astype(np.float16)\n    mask = rgb2gray(thumbnail) &gt; \n    x0, y0, x1, y1 = get_biggest_component_box(mask)\n\n    scale_h = image.shape[] / thumbnail.shape[]\n    scale_w = image.shape[] / thumbnail.shape[]\n\n    x0 = (, math.floor(x0 * scale_w))\n    y0 = (, math.floor(y0 * scale_h))\n    x1 = (image.shape[] - , math.ceil(x1 * scale_w))\n    y1 = (image.shape[] - , math.ceil(y1 * scale_h))\n    image = image[y0: y1 + , x0: x1 + ]\n</code></pre>\n<p>4. Use the non-overlapping sliding window method to tile the tissue area into 256x256 patches. (For TMA I used overlap, but not sure if that would have an impact on the results.)</p>\n<pre><code> ():\n    patches = []\n     i  (, image.shape[], step):\n         j  (, image.shape[], step):\n            patch = image[i: i + patch_size, j: j + patch_size, :]\n             patch.shape != (patch_size, patch_size, ):\n                patch = np.pad(patch, ((, patch_size - patch.shape[]), (, patch_size - patch.shape[]), (, )))\n\n             is_tma:\n                patch = transform(image=patch)[]\n                patches.append(patch)\n            :\n                patch_gray = rgb2gray(patch)  \n                patch_binary = (patch_gray &lt;= ) &amp; (patch_gray &gt; )\n\n                 np.count_nonzero(patch_binary) / patch_binary.size &gt;= ratio:\n                    patch = transform(image=patch)[]\n                    patches.append(patch)\n\n     (patches) != :\n        patches = torch.stack(patches, dim=)\n    :\n        patches = torch.zeros(, dtype=torch.uint8)\n\n     patches\n\nimage2patches(image, , [, ][is_tma], , transform, is_tma)\n</code></pre>\n<h3>2. Subtype classification</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15788664%2Ff7021178fffcbd920a2bbf69c47f4cf2%2FUBC-OCEAN-2.jpeg?generation=1704447379688941&amp;alt=media\"></p>\n<p>Cancer subtype classification method is mainly based on multiple instance learning(MIL).<br>\nAfter trying various backbone and MIL methods, <code>CTransPath</code> and <code>LunitDINO</code> were finally selected as the backbone, <code>DSMIL</code> and <code>Perceiver</code> were selected as the MIL classifier. For their specific information, please refer to:</p>\n<ol>\n<li><a href=\"https://github.com/Xiyue-Wang/TransPath\" target=\"_blank\">CTransPath, MIA2022</a></li>\n<li><a href=\"https://github.com/lunit-io/benchmark-ssl-pathology\" target=\"_blank\">LunitDINO, CVPR2023</a></li>\n<li><a href=\"https://github.com/binli123/dsmil-wsi\" target=\"_blank\">DSMIL, CVPR2021</a></li>\n<li><a href=\"https://github.com/cgtuebingen/DualQueryMIL\" target=\"_blank\">Perceiver, BMVA2023</a></li>\n</ol>\n<p>Local CV results:</p>\n<table>\n<thead>\n<tr>\n<th>exp</th>\n<th>CC</th>\n<th>EC</th>\n<th>HGSC</th>\n<th>LGSC</th>\n<th>MC</th>\n<th>mean</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CTransPath + DSMIL</td>\n<td>0.9300</td>\n<td>0.7657</td>\n<td>0.8909</td>\n<td>0.7822</td>\n<td>0.7911</td>\n<td>0.8320</td>\n</tr>\n<tr>\n<td>CTransPath + Perceiver</td>\n<td>0.9695</td>\n<td>0.8147</td>\n<td>0.8818</td>\n<td>0.8044</td>\n<td>0.9156</td>\n<td>0.8772</td>\n</tr>\n<tr>\n<td>LunitDINO + DSMIL</td>\n<td>0.9400</td>\n<td>0.7240</td>\n<td>0.8864</td>\n<td>0.8244</td>\n<td>0.9356</td>\n<td>0.8621</td>\n</tr>\n<tr>\n<td>LunitDINO + Perceiver</td>\n<td>0.9300</td>\n<td>0.7983</td>\n<td>0.8591</td>\n<td>0.8711</td>\n<td>0.8933</td>\n<td>0.8704</td>\n</tr>\n</tbody>\n</table>\n<p>Leaderboard results:</p>\n<table>\n<thead>\n<tr>\n<th>exp</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CTransPath + LunitDINO + DSMIL</td>\n<td>0.57</td>\n<td>0.54</td>\n</tr>\n<tr>\n<td>CTransPath + LunitDINO + Perceiver</td>\n<td>0.58</td>\n<td>0.57</td>\n</tr>\n<tr>\n<td>CTransPath + LunitDINO + DSMIL + Perceiver</td>\n<td>0.6</td>\n<td>0.58</td>\n</tr>\n</tbody>\n</table>\n<p>I almost didn't adjust the MIL hyperparameters because I found that high CV score tended to be low public score.</p>\n<ol>\n<li>For <code>DSMIL</code>, we use <code>nn.CrossEntropyLoss</code> as loss function.</li>\n<li>For <code>Perceiver</code>, we use <code>nn.BCEWithLogitsLoss</code> as loss function and use <code>mixup</code>, <code>label smoothing</code> to alleviate overfitting.</li>\n</ol>\n<h3>3. Outlier detection</h3>\n<p>We tried many methods, two of which can get a private score of 0.6. (Private score 0.58 if not use outlier detection.)</p>\n<h4>1. BCE + Thresholding</h4>\n<p>Score: public 0.6 and private 0.6.</p>\n<p>This method is very simple. Use <code>nn.BCEWithLogitsLoss</code> as the loss function to train the model, and then for the maximum prediction probability, if it is less than 0.4, it is considered an outlier.</p>\n<pre><code>logits = self.model(x)\nprobs = F.sigmoid(logits)  \npred = probs.argmax(dim=).item()\n (probs) &lt; PROB_THRESH:  \n    pred =   \n</code></pre>\n<h4>2. Probability entropy</h4>\n<p>Score: public 0.54 and private 0.6.</p>\n<p>This method is also very simple. Compared to setting a probability threshold, this method detects outliers by calculating the entropy of the probability.</p>\n<pre><code>logits = self.model(x)\nprobs = F.sigmoid(logits)  \npred = probs.argmax(dim=).item()\nentropy = (probs * torch.log2(probs)).mean(dim=)\n entropy &gt; ENTROPY_THRESH:   \n    pred =   \n</code></pre>\n<h1>Summary</h1>\n<h2>which didn't work</h2>\n<ol>\n<li>Extra dataset: ATEC, PTRC-HGSOC, CPTAC-OV, TCGA-OV, Bevacizumab.</li>\n<li>End-to-end finetune the backbone and MIL together by selecting cancer areas through attention or mask.</li>\n<li>Select only patches in cancer areas for MIL.</li>\n<li>Detect outliers based on patch prediction probability entropy. (<a href=\"https://www.sciencedirect.com/science/article/pii/S1361841522002833\" target=\"_blank\">MIA2023</a>)</li>\n<li>Detect outliers based on KNN classifier. (<a href=\"https://arxiv.org/abs/2309.05528\" target=\"_blank\">Arxiv2023</a>)</li>\n</ol>\n<h1>Supplementary</h1>\n<p>All pytorch codes(include submission notebook) are built based on <a href=\"https://github.com/m1dsolo/yangdl\" target=\"_blank\">a simple pytorch-based deep learning framework</a>.<br>\nThis framework has only a few hundred lines of code and I think it is very suitable for beginners to learn.</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/code/m1dsolo/ubc-ocean-7th-submission\" target=\"_blank\">submission notebook</a></li>\n<li><a href=\"https://github.com/m1dsolo/UBC-OCEAN-7th\" target=\"_blank\">Training code</a></li>\n</ol>",
  "messages": [
    {
      "id": 2588131,
      "postDate": "2024-01-05T09:37:04.807Z",
      "content": "<p>Thanks to kaggle and UBC for hosting this interesting competition and congrats to all the winners for their hard work! I would also like to thank my teammates and everyone in the discussion forum for their help!</p>\n<h1>Method</h1>\n<h2>Summary</h2>\n<p>Our final solution is based on multiple instance learning(MIL) for <strong>ovarian cancer subtype classification</strong> and use <code>sigmoid</code> and thresholding for <strong>outlier detection</strong>.<br>\nWe did not use mask annotation and additional datasets in the final submission.</p>\n<h3>1. preprocess</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15788664%2F16e938db8ccddff048fef4f4b9f306a3%2FUBC-OCEAN-1.jpeg?generation=1704447344840948&amp;alt=media\"></p>\n<p>1. Use <code>pyvips</code> to speed up png reading speed. (Thanks for <a href=\"https://www.kaggle.com/code/gunesevitan/libvips-pyvips-installation-and-getting-started\" target=\"_blank\">GUNES EVITAN's pyvips notebook</a>.)</p>\n<pre><code>image = pyvips.Image.new_from_file(image_id, access=).numpy()\nis_tma = image.shape[] &lt;=   image.shape[] &lt;= \n</code></pre>\n<p>2. Downsample WSI and TMA from x20 and x40 to x10 respectively. (Maybe x20 results will be better, but I can't submit due to resource constraints.)</p>\n<pre><code> is_tma:\n    resize = A.Resize(image.shape[] // , image.shape[] // )\n:\n    resize = A.Resize(image.shape[] // , image.shape[] // )\nimage = resize(image=image)[]\n</code></pre>\n<p>3. Deduplicate the identical tissue areas for WSI. (I'm not sure if this contributed to the results, but it saved me a lot of local memory.)</p>\n<pre><code> ():\n    image = image.astype(np.float16)\n    image = (image[..., ] *  + image[..., ] *  + image[..., ] * ) / \n     image.astype(np.uint8)\n\n  is_tma:\n    resize = A.Resize(image.shape[] // , image.shape[] // )\n    thumbnail = resize(image=image)[].astype(np.float16)\n    mask = rgb2gray(thumbnail) &gt; \n    x0, y0, x1, y1 = get_biggest_component_box(mask)\n\n    scale_h = image.shape[] / thumbnail.shape[]\n    scale_w = image.shape[] / thumbnail.shape[]\n\n    x0 = (, math.floor(x0 * scale_w))\n    y0 = (, math.floor(y0 * scale_h))\n    x1 = (image.shape[] - , math.ceil(x1 * scale_w))\n    y1 = (image.shape[] - , math.ceil(y1 * scale_h))\n    image = image[y0: y1 + , x0: x1 + ]\n</code></pre>\n<p>4. Use the non-overlapping sliding window method to tile the tissue area into 256x256 patches. (For TMA I used overlap, but not sure if that would have an impact on the results.)</p>\n<pre><code> ():\n    patches = []\n     i  (, image.shape[], step):\n         j  (, image.shape[], step):\n            patch = image[i: i + patch_size, j: j + patch_size, :]\n             patch.shape != (patch_size, patch_size, ):\n                patch = np.pad(patch, ((, patch_size - patch.shape[]), (, patch_size - patch.shape[]), (, )))\n\n             is_tma:\n                patch = transform(image=patch)[]\n                patches.append(patch)\n            :\n                patch_gray = rgb2gray(patch)  \n                patch_binary = (patch_gray &lt;= ) &amp; (patch_gray &gt; )\n\n                 np.count_nonzero(patch_binary) / patch_binary.size &gt;= ratio:\n                    patch = transform(image=patch)[]\n                    patches.append(patch)\n\n     (patches) != :\n        patches = torch.stack(patches, dim=)\n    :\n        patches = torch.zeros(, dtype=torch.uint8)\n\n     patches\n\nimage2patches(image, , [, ][is_tma], , transform, is_tma)\n</code></pre>\n<h3>2. Subtype classification</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15788664%2Ff7021178fffcbd920a2bbf69c47f4cf2%2FUBC-OCEAN-2.jpeg?generation=1704447379688941&amp;alt=media\"></p>\n<p>Cancer subtype classification method is mainly based on multiple instance learning(MIL).<br>\nAfter trying various backbone and MIL methods, <code>CTransPath</code> and <code>LunitDINO</code> were finally selected as the backbone, <code>DSMIL</code> and <code>Perceiver</code> were selected as the MIL classifier. For their specific information, please refer to:</p>\n<ol>\n<li><a href=\"https://github.com/Xiyue-Wang/TransPath\" target=\"_blank\">CTransPath, MIA2022</a></li>\n<li><a href=\"https://github.com/lunit-io/benchmark-ssl-pathology\" target=\"_blank\">LunitDINO, CVPR2023</a></li>\n<li><a href=\"https://github.com/binli123/dsmil-wsi\" target=\"_blank\">DSMIL, CVPR2021</a></li>\n<li><a href=\"https://github.com/cgtuebingen/DualQueryMIL\" target=\"_blank\">Perceiver, BMVA2023</a></li>\n</ol>\n<p>Local CV results:</p>\n<table>\n<thead>\n<tr>\n<th>exp</th>\n<th>CC</th>\n<th>EC</th>\n<th>HGSC</th>\n<th>LGSC</th>\n<th>MC</th>\n<th>mean</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CTransPath + DSMIL</td>\n<td>0.9300</td>\n<td>0.7657</td>\n<td>0.8909</td>\n<td>0.7822</td>\n<td>0.7911</td>\n<td>0.8320</td>\n</tr>\n<tr>\n<td>CTransPath + Perceiver</td>\n<td>0.9695</td>\n<td>0.8147</td>\n<td>0.8818</td>\n<td>0.8044</td>\n<td>0.9156</td>\n<td>0.8772</td>\n</tr>\n<tr>\n<td>LunitDINO + DSMIL</td>\n<td>0.9400</td>\n<td>0.7240</td>\n<td>0.8864</td>\n<td>0.8244</td>\n<td>0.9356</td>\n<td>0.8621</td>\n</tr>\n<tr>\n<td>LunitDINO + Perceiver</td>\n<td>0.9300</td>\n<td>0.7983</td>\n<td>0.8591</td>\n<td>0.8711</td>\n<td>0.8933</td>\n<td>0.8704</td>\n</tr>\n</tbody>\n</table>\n<p>Leaderboard results:</p>\n<table>\n<thead>\n<tr>\n<th>exp</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CTransPath + LunitDINO + DSMIL</td>\n<td>0.57</td>\n<td>0.54</td>\n</tr>\n<tr>\n<td>CTransPath + LunitDINO + Perceiver</td>\n<td>0.58</td>\n<td>0.57</td>\n</tr>\n<tr>\n<td>CTransPath + LunitDINO + DSMIL + Perceiver</td>\n<td>0.6</td>\n<td>0.58</td>\n</tr>\n</tbody>\n</table>\n<p>I almost didn't adjust the MIL hyperparameters because I found that high CV score tended to be low public score.</p>\n<ol>\n<li>For <code>DSMIL</code>, we use <code>nn.CrossEntropyLoss</code> as loss function.</li>\n<li>For <code>Perceiver</code>, we use <code>nn.BCEWithLogitsLoss</code> as loss function and use <code>mixup</code>, <code>label smoothing</code> to alleviate overfitting.</li>\n</ol>\n<h3>3. Outlier detection</h3>\n<p>We tried many methods, two of which can get a private score of 0.6. (Private score 0.58 if not use outlier detection.)</p>\n<h4>1. BCE + Thresholding</h4>\n<p>Score: public 0.6 and private 0.6.</p>\n<p>This method is very simple. Use <code>nn.BCEWithLogitsLoss</code> as the loss function to train the model, and then for the maximum prediction probability, if it is less than 0.4, it is considered an outlier.</p>\n<pre><code>logits = self.model(x)\nprobs = F.sigmoid(logits)  \npred = probs.argmax(dim=).item()\n (probs) &lt; PROB_THRESH:  \n    pred =   \n</code></pre>\n<h4>2. Probability entropy</h4>\n<p>Score: public 0.54 and private 0.6.</p>\n<p>This method is also very simple. Compared to setting a probability threshold, this method detects outliers by calculating the entropy of the probability.</p>\n<pre><code>logits = self.model(x)\nprobs = F.sigmoid(logits)  \npred = probs.argmax(dim=).item()\nentropy = (probs * torch.log2(probs)).mean(dim=)\n entropy &gt; ENTROPY_THRESH:   \n    pred =   \n</code></pre>\n<h1>Summary</h1>\n<h2>which didn't work</h2>\n<ol>\n<li>Extra dataset: ATEC, PTRC-HGSOC, CPTAC-OV, TCGA-OV, Bevacizumab.</li>\n<li>End-to-end finetune the backbone and MIL together by selecting cancer areas through attention or mask.</li>\n<li>Select only patches in cancer areas for MIL.</li>\n<li>Detect outliers based on patch prediction probability entropy. (<a href=\"https://www.sciencedirect.com/science/article/pii/S1361841522002833\" target=\"_blank\">MIA2023</a>)</li>\n<li>Detect outliers based on KNN classifier. (<a href=\"https://arxiv.org/abs/2309.05528\" target=\"_blank\">Arxiv2023</a>)</li>\n</ol>\n<h1>Supplementary</h1>\n<p>All pytorch codes(include submission notebook) are built based on <a href=\"https://github.com/m1dsolo/yangdl\" target=\"_blank\">a simple pytorch-based deep learning framework</a>.<br>\nThis framework has only a few hundred lines of code and I think it is very suitable for beginners to learn.</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/code/m1dsolo/ubc-ocean-7th-submission\" target=\"_blank\">submission notebook</a></li>\n<li><a href=\"https://github.com/m1dsolo/UBC-OCEAN-7th\" target=\"_blank\">Training code</a></li>\n</ol>",
      "rawMarkdown": "Thanks to kaggle and UBC for hosting this interesting competition and congrats to all the winners for their hard work! I would also like to thank my teammates and everyone in the discussion forum for their help!\n\n# Method\n\n## Summary\n\nOur final solution is based on multiple instance learning(MIL) for **ovarian cancer subtype classification** and use `sigmoid` and thresholding for **outlier detection**.\nWe did not use mask annotation and additional datasets in the final submission.\n\n### 1. preprocess\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15788664%2F16e938db8ccddff048fef4f4b9f306a3%2FUBC-OCEAN-1.jpeg?generation=1704447344840948&alt=media)\n\n1\\. Use `pyvips` to speed up png reading speed. (Thanks for [GUNES EVITAN's pyvips notebook](https://www.kaggle.com/code/gunesevitan/libvips-pyvips-installation-and-getting-started).)\n\n```python\nimage = pyvips.Image.new_from_file(image_id, access='sequential').numpy()\nis_tma = image.shape[0] <= 5000 and image.shape[1] <= 5000\n```\n\n2\\. Downsample WSI and TMA from x20 and x40 to x10 respectively. (Maybe x20 results will be better, but I can't submit due to resource constraints.)\n\n```python\nif is_tma:\n    resize = A.Resize(image.shape[0] // 4, image.shape[1] // 4)\nelse:\n    resize = A.Resize(image.shape[0] // 2, image.shape[1] // 2)\nimage = resize(image=image)['image']\n```\n\n3\\. Deduplicate the identical tissue areas for WSI. (I'm not sure if this contributed to the results, but it saved me a lot of local memory.)\n\n```python\ndef rgb2gray(image: np.ndarray):\n    image = image.astype(np.float16)\n    image = (image[..., 0] * 299 + image[..., 1] * 587 + image[..., 2] * 114) / 1000\n    return image.astype(np.uint8)\n\nif not is_tma:\n    resize = A.Resize(image.shape[0] // 16, image.shape[1] // 16)\n    thumbnail = resize(image=image)['image'].astype(np.float16)\n    mask = rgb2gray(thumbnail) > 0\n    x0, y0, x1, y1 = get_biggest_component_box(mask)\n\n    scale_h = image.shape[0] / thumbnail.shape[0]\n    scale_w = image.shape[1] / thumbnail.shape[1]\n\n    x0 = max(0, math.floor(x0 * scale_w))\n    y0 = max(0, math.floor(y0 * scale_h))\n    x1 = min(image.shape[1] - 1, math.ceil(x1 * scale_w))\n    y1 = min(image.shape[0] - 1, math.ceil(y1 * scale_h))\n    image = image[y0: y1 + 1, x0: x1 + 1]\n```\n\n4\\. Use the non-overlapping sliding window method to tile the tissue area into 256x256 patches. (For TMA I used overlap, but not sure if that would have an impact on the results.)\n\n```python\ndef image2patches(image: np.ndarray, patch_size: int, step: int, ratio: float, transform, is_tma: bool):\n    patches = []\n    for i in range(0, image.shape[0], step):\n        for j in range(0, image.shape[1], step):\n            patch = image[i: i + patch_size, j: j + patch_size, :]\n            if patch.shape != (patch_size, patch_size, 3):\n                patch = np.pad(patch, ((0, patch_size - patch.shape[0]), (0, patch_size - patch.shape[1]), (0, 0)))\n\n            if is_tma:\n                patch = transform(image=patch)['image']\n                patches.append(patch)\n            else:\n                patch_gray = rgb2gray(patch)  # (patch_size, patch_size)\n                patch_binary = (patch_gray <= 220) & (patch_gray > 0)\n\n                if np.count_nonzero(patch_binary) / patch_binary.size >= ratio:\n                    patch = transform(image=patch)['image']\n                    patches.append(patch)\n\n    if len(patches) != 0:\n        patches = torch.stack(patches, dim=0)\n    else:\n        patches = torch.zeros(0, dtype=torch.uint8)\n\n    return patches\n\nimage2patches(image, 256, [256, 128][is_tma], 0.25, transform, is_tma)\n```\n\n### 2. Subtype classification\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15788664%2Ff7021178fffcbd920a2bbf69c47f4cf2%2FUBC-OCEAN-2.jpeg?generation=1704447379688941&alt=media)\n\nCancer subtype classification method is mainly based on multiple instance learning(MIL).\nAfter trying various backbone and MIL methods, `CTransPath` and `LunitDINO` were finally selected as the backbone, `DSMIL` and `Perceiver` were selected as the MIL classifier. For their specific information, please refer to:\n\n1. [CTransPath, MIA2022](https://github.com/Xiyue-Wang/TransPath)\n2. [LunitDINO, CVPR2023](https://github.com/lunit-io/benchmark-ssl-pathology)\n3. [DSMIL, CVPR2021](https://github.com/binli123/dsmil-wsi)\n4. [Perceiver, BMVA2023](https://github.com/cgtuebingen/DualQueryMIL)\n\nLocal CV results:\n\n| exp | CC | EC | HGSC | LGSC | MC | mean |\n| --- | --- | --- | --- | --- | --- | --- |\n| CTransPath + DSMIL | 0.9300 | 0.7657 | 0.8909 | 0.7822 | 0.7911 | 0.8320 |\n| CTransPath + Perceiver | 0.9695 | 0.8147 | 0.8818 | 0.8044 | 0.9156 | 0.8772 |\n| LunitDINO + DSMIL | 0.9400 | 0.7240 | 0.8864 | 0.8244 | 0.9356 | 0.8621 |\n| LunitDINO + Perceiver | 0.9300 | 0.7983 | 0.8591 | 0.8711 | 0.8933 | 0.8704 |\n\nLeaderboard results:\n\n| exp | public | private |\n| --- | --- | --- |\n| CTransPath + LunitDINO + DSMIL | 0.57 | 0.54 |\n| CTransPath + LunitDINO + Perceiver | 0.58 | 0.57 |\n| CTransPath + LunitDINO + DSMIL + Perceiver | 0.6 | 0.58 |\n\nI almost didn't adjust the MIL hyperparameters because I found that high CV score tended to be low public score.\n\n1. For `DSMIL`, we use `nn.CrossEntropyLoss` as loss function.\n2. For `Perceiver`, we use `nn.BCEWithLogitsLoss` as loss function and use `mixup`, `label smoothing` to alleviate overfitting.\n\n### 3. Outlier detection\n\nWe tried many methods, two of which can get a private score of 0.6. (Private score 0.58 if not use outlier detection.)\n\n#### 1. BCE + Thresholding\n\nScore: public 0.6 and private 0.6.\n\nThis method is very simple. Use `nn.BCEWithLogitsLoss` as the loss function to train the model, and then for the maximum prediction probability, if it is less than 0.4, it is considered an outlier.\n\n```python\nlogits = self.model(x)\nprobs = F.sigmoid(logits)  # (C,)\npred = probs.argmax(dim=0).item()\nif max(probs) < PROB_THRESH:  # Choose the threshold based on the validation set\n    pred = 5  # 'Other' class\n```\n\n#### 2. Probability entropy\n\nScore: public 0.54 and private 0.6.\n\nThis method is also very simple. Compared to setting a probability threshold, this method detects outliers by calculating the entropy of the probability.\n\n```python\nlogits = self.model(x)\nprobs = F.sigmoid(logits)  # (C,)\npred = probs.argmax(dim=0).item()\nentropy = (probs * torch.log2(probs)).mean(dim=0)\nif entropy > ENTROPY_THRESH:   # Choose the threshold based on the validation set \n    pred = 5  # 'Other' class\n```\n\n# Summary\n\n## which didn't work\n\n1. Extra dataset: ATEC, PTRC-HGSOC, CPTAC-OV, TCGA-OV, Bevacizumab.\n2. End-to-end finetune the backbone and MIL together by selecting cancer areas through attention or mask.\n3. Select only patches in cancer areas for MIL.\n4. Detect outliers based on patch prediction probability entropy. ([MIA2023](https://www.sciencedirect.com/science/article/pii/S1361841522002833))\n5. Detect outliers based on KNN classifier. ([Arxiv2023](https://arxiv.org/abs/2309.05528))\n\n# Supplementary\n\nAll pytorch codes(include submission notebook) are built based on [a simple pytorch-based deep learning framework](https://github.com/m1dsolo/yangdl).\nThis framework has only a few hundred lines of code and I think it is very suitable for beginners to learn.\n\n1. [submission notebook](https://www.kaggle.com/code/m1dsolo/ubc-ocean-7th-submission)\n2. [Training code](https://github.com/m1dsolo/UBC-OCEAN-7th)\n\n",
      "votes": 35
    },
    {
      "id": 2589611,
      "postDate": "2024-01-06T14:03:46.720Z",
      "content": "<p>Training code has been released to <a href=\"https://github.com/m1dsolo/UBC-OCEAN-7th\" target=\"_blank\">github</a>.</p>",
      "rawMarkdown": "Training code has been released to [github](https://github.com/m1dsolo/UBC-OCEAN-7th).",
      "votes": 3,
      "replies": [
        {
          "id": 2831361,
          "postDate": "2024-05-23T16:32:01.560Z",
          "content": "<p>Thanks for sharing and beautiful explanation!</p>",
          "rawMarkdown": "Thanks for sharing and beautiful explanation!"
        }
      ]
    },
    {
      "id": 2588501,
      "postDate": "2024-01-05T14:37:33.620Z",
      "content": "<p>Really brief and clear explanation, it's so helpful. Thanks for sharing your excellent work. I'm looking forward to your training notebook, as I'm not familiar with MIL training😆. Anyway, congrats and Happy New Year!</p>",
      "rawMarkdown": "Really brief and clear explanation, it's so helpful. Thanks for sharing your excellent work. I'm looking forward to your training notebook, as I'm not familiar with MIL training😆. Anyway, congrats and Happy New Year!",
      "votes": 3,
      "replies": [
        {
          "id": 2588532,
          "postDate": "2024-01-05T15:01:51.707Z",
          "content": "<p>Thanks! Happy new year to you too!</p>",
          "rawMarkdown": "Thanks! Happy new year to you too!",
          "votes": 2,
          "replies": [
            {
              "id": 2588944,
              "postDate": "2024-01-05T22:06:36.280Z",
              "content": "<p>I'm wondering about the deduplication part:</p>\n<p>\"if is_tma:<br>\n    …<br>\n    x1 = min(thumbnail.shape[1] - 1, math.ceil(x1 * scale_w))<br>\n    y1 = min(thumbnail.shape[0] - 1, math.ceil(y1 * scale_h))\"</p>\n<p>Should 'if is_tma' actually be 'if not is_tma' since it's meant for WSI? Also, should 'thumbnail' be changed to 'image' in this context, since the cut is applied to image not thumbnail.</p>\n<p>I really appreciate your solution; it's nearly everything I need. Thank you very much.</p>",
              "rawMarkdown": "I'm wondering about the deduplication part:\n\n\"if is_tma:\n    ...\n    x1 = min(thumbnail.shape[1] - 1, math.ceil(x1 * scale_w))\n    y1 = min(thumbnail.shape[0] - 1, math.ceil(y1 * scale_h))\"\n\nShould 'if is_tma' actually be 'if not is_tma' since it's meant for WSI? Also, should 'thumbnail' be changed to 'image' in this context, since the cut is applied to image not thumbnail.\n\nI really appreciate your solution; it's nearly everything I need. Thank you very much.",
              "votes": 1
            },
            {
              "id": 2589042,
              "postDate": "2024-01-06T02:33:21.253Z",
              "content": "<p>Yes, it is <code>if not is_tma</code>, I made a mistake when organizing the code. I've corrected it. The purpose of <code>thumbnail</code> is only to speed up the process of finding the bounding box <code>[x0, y0, x1, y1]</code>, the bounding box will be scaled in <code>x0 = max(0, math.floor(x0 * scale_w)) ...</code>. Selecting ROI through bounding boxes can achieve the purpose of deduplication.</p>",
              "rawMarkdown": "Yes, it is `if not is_tma`, I made a mistake when organizing the code. I've corrected it. The purpose of `thumbnail` is only to speed up the process of finding the bounding box `[x0, y0, x1, y1]`, the bounding box will be scaled in `x0 = max(0, math.floor(x0 * scale_w)) ...`. Selecting ROI through bounding boxes can achieve the purpose of deduplication.",
              "votes": 1
            },
            {
              "id": 2589116,
              "postDate": "2024-01-06T05:12:27.237Z",
              "content": "<p>[Edited] I have seen the revised version and there are no other issues.😀</p>\n<p>[Solved] Thanks for the response. I understand now that the cut is initially made on the thumbnail and then scaled up to the WSI. So, I'm considering adjusting the new boundary coordinates as <code>new x1 = min(image.shape[1] - 1, math.ceil(x1 * scale_w))</code> to ensure the scaled boundary does not exceed the dimensions of the WSI image. Is that correct?</p>",
              "rawMarkdown": "[Edited] I have seen the revised version and there are no other issues.😀\n\n[Solved] Thanks for the response. I understand now that the cut is initially made on the thumbnail and then scaled up to the WSI. So, I'm considering adjusting the new boundary coordinates as `new x1 = min(image.shape[1] - 1, math.ceil(x1 * scale_w))` to ensure the scaled boundary does not exceed the dimensions of the WSI image. Is that correct?",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2588200,
      "postDate": "2024-01-05T10:41:46.720Z",
      "content": "<p>Congrats! Thanks for the nice and clean explanation. Would it be possible to share the training notebook/code as well?</p>",
      "rawMarkdown": "Congrats! Thanks for the nice and clean explanation. Would it be possible to share the training notebook/code as well?",
      "votes": 2,
      "replies": [
        {
          "id": 2588310,
          "postDate": "2024-01-05T12:25:15.637Z",
          "content": "<p>I will start cleaning up the training code tomorrow, I hope it can be helpful to everyone.</p>",
          "rawMarkdown": "I will start cleaning up the training code tomorrow, I hope it can be helpful to everyone.",
          "votes": 4
        }
      ]
    },
    {
      "id": 2604702,
      "postDate": "2024-01-16T16:01:44.020Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/m1dsolo\" target=\"_blank\">@m1dsolo</a> Can you provide your CTransPath training book I tried to emulate your work from inference and github but unable to get it right somehow. Thanks in advance.</p>",
      "rawMarkdown": "Hi @m1dsolo Can you provide your CTransPath training book I tried to emulate your work from inference and github but unable to get it right somehow. Thanks in advance.\n",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 2605237,
          "postDate": "2024-01-17T00:47:50.177Z",
          "content": "<p>The <code>Ctranspath</code> weight I used is the <a href=\"https://drive.google.com/file/d/1DoDx_70_TLj98gTf6YTXnu4tFhsFocDX/view\" target=\"_blank\">official weight</a>. I only changed the file name.</p>",
          "rawMarkdown": "The `Ctranspath` weight I used is the [official weight](https://drive.google.com/file/d/1DoDx_70_TLj98gTf6YTXnu4tFhsFocDX/view). I only changed the file name.",
          "replies": [
            {
              "id": 2606439,
              "postDate": "2024-01-17T16:22:41.167Z",
              "content": "<p>thanks for sharing. <a href=\"https://www.kaggle.com/m1dsolo\" target=\"_blank\">@m1dsolo</a> </p>",
              "rawMarkdown": "thanks for sharing. @m1dsolo ",
              "votes": 1,
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2589611,
      "author_name": "m1dsolo",
      "author_url": "",
      "post_date": "2024-01-06T14:03:46.720000",
      "content": "<p>Training code has been released to <a href=\"https://github.com/m1dsolo/UBC-OCEAN-7th\" target=\"_blank\">github</a>.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2831361,
          "author_name": "bijhatu",
          "author_url": "",
          "post_date": "2024-05-23T16:32:01.560000",
          "content": "<p>Thanks for sharing and beautiful explanation!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2588501,
      "author_name": "snowmanfat",
      "author_url": "",
      "post_date": "2024-01-05T14:37:33.620000",
      "content": "<p>Really brief and clear explanation, it's so helpful. Thanks for sharing your excellent work. I'm looking forward to your training notebook, as I'm not familiar with MIL training😆. Anyway, congrats and Happy New Year!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2588532,
          "author_name": "m1dsolo",
          "author_url": "",
          "post_date": "2024-01-05T15:01:51.707000",
          "content": "<p>Thanks! Happy new year to you too!</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2588944,
              "author_name": "snowmanfat",
              "author_url": "",
              "post_date": "2024-01-05T22:06:36.280000",
              "content": "<p>I'm wondering about the deduplication part:</p>\n<p>\"if is_tma:<br>\n    …<br>\n    x1 = min(thumbnail.shape[1] - 1, math.ceil(x1 * scale_w))<br>\n    y1 = min(thumbnail.shape[0] - 1, math.ceil(y1 * scale_h))\"</p>\n<p>Should 'if is_tma' actually be 'if not is_tma' since it's meant for WSI? Also, should 'thumbnail' be changed to 'image' in this context, since the cut is applied to image not thumbnail.</p>\n<p>I really appreciate your solution; it's nearly everything I need. Thank you very much.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2589042,
              "author_name": "m1dsolo",
              "author_url": "",
              "post_date": "2024-01-06T02:33:21.253000",
              "content": "<p>Yes, it is <code>if not is_tma</code>, I made a mistake when organizing the code. I've corrected it. The purpose of <code>thumbnail</code> is only to speed up the process of finding the bounding box <code>[x0, y0, x1, y1]</code>, the bounding box will be scaled in <code>x0 = max(0, math.floor(x0 * scale_w)) ...</code>. Selecting ROI through bounding boxes can achieve the purpose of deduplication.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2589116,
              "author_name": "snowmanfat",
              "author_url": "",
              "post_date": "2024-01-06T05:12:27.237000",
              "content": "<p>[Edited] I have seen the revised version and there are no other issues.😀</p>\n<p>[Solved] Thanks for the response. I understand now that the cut is initially made on the thumbnail and then scaled up to the WSI. So, I'm considering adjusting the new boundary coordinates as <code>new x1 = min(image.shape[1] - 1, math.ceil(x1 * scale_w))</code> to ensure the scaled boundary does not exceed the dimensions of the WSI image. Is that correct?</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2588200,
      "author_name": "Jinho Park",
      "author_url": "",
      "post_date": "2024-01-05T10:41:46.720000",
      "content": "<p>Congrats! Thanks for the nice and clean explanation. Would it be possible to share the training notebook/code as well?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2588310,
          "author_name": "m1dsolo",
          "author_url": "",
          "post_date": "2024-01-05T12:25:15.637000",
          "content": "<p>I will start cleaning up the training code tomorrow, I hope it can be helpful to everyone.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2604702,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-16T16:01:44.020000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/m1dsolo\" target=\"_blank\">@m1dsolo</a> Can you provide your CTransPath training book I tried to emulate your work from inference and github but unable to get it right somehow. Thanks in advance.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2605237,
          "author_name": "m1dsolo",
          "author_url": "",
          "post_date": "2024-01-17T00:47:50.177000",
          "content": "<p>The <code>Ctranspath</code> weight I used is the <a href=\"https://drive.google.com/file/d/1DoDx_70_TLj98gTf6YTXnu4tFhsFocDX/view\" target=\"_blank\">official weight</a>. I only changed the file name.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2606439,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-01-17T16:22:41.167000",
              "content": "<p>thanks for sharing. <a href=\"https://www.kaggle.com/m1dsolo\" target=\"_blank\">@m1dsolo</a> </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2588131": "Thanks to kaggle and UBC for hosting this interesting competition and congrats to all the winners for their hard work! I would also like to thank my teammates and everyone in the discussion forum for their help!\n\n# Method\n\n## Summary\n\nOur final solution is based on multiple instance learning(MIL) for **ovarian cancer subtype classification** and use `sigmoid` and thresholding for **outlier detection**.\nWe did not use mask annotation and additional datasets in the final submission.\n\n### 1. preprocess\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15788664%2F16e938db8ccddff048fef4f4b9f306a3%2FUBC-OCEAN-1.jpeg?generation=1704447344840948&alt=media)\n\n1\\. Use `pyvips` to speed up png reading speed. (Thanks for [GUNES EVITAN's pyvips notebook](https://www.kaggle.com/code/gunesevitan/libvips-pyvips-installation-and-getting-started).)\n\n```python\nimage = pyvips.Image.new_from_file(image_id, access='sequential').numpy()\nis_tma = image.shape[0] <= 5000 and image.shape[1] <= 5000\n```\n\n2\\. Downsample WSI and TMA from x20 and x40 to x10 respectively. (Maybe x20 results will be better, but I can't submit due to resource constraints.)\n\n```python\nif is_tma:\n    resize = A.Resize(image.shape[0] // 4, image.shape[1] // 4)\nelse:\n    resize = A.Resize(image.shape[0] // 2, image.shape[1] // 2)\nimage = resize(image=image)['image']\n```\n\n3\\. Deduplicate the identical tissue areas for WSI. (I'm not sure if this contributed to the results, but it saved me a lot of local memory.)\n\n```python\ndef rgb2gray(image: np.ndarray):\n    image = image.astype(np.float16)\n    image = (image[..., 0] * 299 + image[..., 1] * 587 + image[..., 2] * 114) / 1000\n    return image.astype(np.uint8)\n\nif not is_tma:\n    resize = A.Resize(image.shape[0] // 16, image.shape[1] // 16)\n    thumbnail = resize(image=image)['image'].astype(np.float16)\n    mask = rgb2gray(thumbnail) > 0\n    x0, y0, x1, y1 = get_biggest_component_box(mask)\n\n    scale_h = image.shape[0] / thumbnail.shape[0]\n    scale_w = image.shape[1] / thumbnail.shape[1]\n\n    x0 = max(0, math.floor(x0 * scale_w))\n    y0 = max(0, math.floor(y0 * scale_h))\n    x1 = min(image.shape[1] - 1, math.ceil(x1 * scale_w))\n    y1 = min(image.shape[0] - 1, math.ceil(y1 * scale_h))\n    image = image[y0: y1 + 1, x0: x1 + 1]\n```\n\n4\\. Use the non-overlapping sliding window method to tile the tissue area into 256x256 patches. (For TMA I used overlap, but not sure if that would have an impact on the results.)\n\n```python\ndef image2patches(image: np.ndarray, patch_size: int, step: int, ratio: float, transform, is_tma: bool):\n    patches = []\n    for i in range(0, image.shape[0], step):\n        for j in range(0, image.shape[1], step):\n            patch = image[i: i + patch_size, j: j + patch_size, :]\n            if patch.shape != (patch_size, patch_size, 3):\n                patch = np.pad(patch, ((0, patch_size - patch.shape[0]), (0, patch_size - patch.shape[1]), (0, 0)))\n\n            if is_tma:\n                patch = transform(image=patch)['image']\n                patches.append(patch)\n            else:\n                patch_gray = rgb2gray(patch)  # (patch_size, patch_size)\n                patch_binary = (patch_gray <= 220) & (patch_gray > 0)\n\n                if np.count_nonzero(patch_binary) / patch_binary.size >= ratio:\n                    patch = transform(image=patch)['image']\n                    patches.append(patch)\n\n    if len(patches) != 0:\n        patches = torch.stack(patches, dim=0)\n    else:\n        patches = torch.zeros(0, dtype=torch.uint8)\n\n    return patches\n\nimage2patches(image, 256, [256, 128][is_tma], 0.25, transform, is_tma)\n```\n\n### 2. Subtype classification\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15788664%2Ff7021178fffcbd920a2bbf69c47f4cf2%2FUBC-OCEAN-2.jpeg?generation=1704447379688941&alt=media)\n\nCancer subtype classification method is mainly based on multiple instance learning(MIL).\nAfter trying various backbone and MIL methods, `CTransPath` and `LunitDINO` were finally selected as the backbone, `DSMIL` and `Perceiver` were selected as the MIL classifier. For their specific information, please refer to:\n\n1. [CTransPath, MIA2022](https://github.com/Xiyue-Wang/TransPath)\n2. [LunitDINO, CVPR2023](https://github.com/lunit-io/benchmark-ssl-pathology)\n3. [DSMIL, CVPR2021](https://github.com/binli123/dsmil-wsi)\n4. [Perceiver, BMVA2023](https://github.com/cgtuebingen/DualQueryMIL)\n\nLocal CV results:\n\n| exp | CC | EC | HGSC | LGSC | MC | mean |\n| --- | --- | --- | --- | --- | --- | --- |\n| CTransPath + DSMIL | 0.9300 | 0.7657 | 0.8909 | 0.7822 | 0.7911 | 0.8320 |\n| CTransPath + Perceiver | 0.9695 | 0.8147 | 0.8818 | 0.8044 | 0.9156 | 0.8772 |\n| LunitDINO + DSMIL | 0.9400 | 0.7240 | 0.8864 | 0.8244 | 0.9356 | 0.8621 |\n| LunitDINO + Perceiver | 0.9300 | 0.7983 | 0.8591 | 0.8711 | 0.8933 | 0.8704 |\n\nLeaderboard results:\n\n| exp | public | private |\n| --- | --- | --- |\n| CTransPath + LunitDINO + DSMIL | 0.57 | 0.54 |\n| CTransPath + LunitDINO + Perceiver | 0.58 | 0.57 |\n| CTransPath + LunitDINO + DSMIL + Perceiver | 0.6 | 0.58 |\n\nI almost didn't adjust the MIL hyperparameters because I found that high CV score tended to be low public score.\n\n1. For `DSMIL`, we use `nn.CrossEntropyLoss` as loss function.\n2. For `Perceiver`, we use `nn.BCEWithLogitsLoss` as loss function and use `mixup`, `label smoothing` to alleviate overfitting.\n\n### 3. Outlier detection\n\nWe tried many methods, two of which can get a private score of 0.6. (Private score 0.58 if not use outlier detection.)\n\n#### 1. BCE + Thresholding\n\nScore: public 0.6 and private 0.6.\n\nThis method is very simple. Use `nn.BCEWithLogitsLoss` as the loss function to train the model, and then for the maximum prediction probability, if it is less than 0.4, it is considered an outlier.\n\n```python\nlogits = self.model(x)\nprobs = F.sigmoid(logits)  # (C,)\npred = probs.argmax(dim=0).item()\nif max(probs) < PROB_THRESH:  # Choose the threshold based on the validation set\n    pred = 5  # 'Other' class\n```\n\n#### 2. Probability entropy\n\nScore: public 0.54 and private 0.6.\n\nThis method is also very simple. Compared to setting a probability threshold, this method detects outliers by calculating the entropy of the probability.\n\n```python\nlogits = self.model(x)\nprobs = F.sigmoid(logits)  # (C,)\npred = probs.argmax(dim=0).item()\nentropy = (probs * torch.log2(probs)).mean(dim=0)\nif entropy > ENTROPY_THRESH:   # Choose the threshold based on the validation set \n    pred = 5  # 'Other' class\n```\n\n# Summary\n\n## which didn't work\n\n1. Extra dataset: ATEC, PTRC-HGSOC, CPTAC-OV, TCGA-OV, Bevacizumab.\n2. End-to-end finetune the backbone and MIL together by selecting cancer areas through attention or mask.\n3. Select only patches in cancer areas for MIL.\n4. Detect outliers based on patch prediction probability entropy. ([MIA2023](https://www.sciencedirect.com/science/article/pii/S1361841522002833))\n5. Detect outliers based on KNN classifier. ([Arxiv2023](https://arxiv.org/abs/2309.05528))\n\n# Supplementary\n\nAll pytorch codes(include submission notebook) are built based on [a simple pytorch-based deep learning framework](https://github.com/m1dsolo/yangdl).\nThis framework has only a few hundred lines of code and I think it is very suitable for beginners to learn.\n\n1. [submission notebook](https://www.kaggle.com/code/m1dsolo/ubc-ocean-7th-submission)\n2. [Training code](https://github.com/m1dsolo/UBC-OCEAN-7th)\n\n",
    "2589611": "Training code has been released to [github](https://github.com/m1dsolo/UBC-OCEAN-7th).",
    "2588501": "Really brief and clear explanation, it's so helpful. Thanks for sharing your excellent work. I'm looking forward to your training notebook, as I'm not familiar with MIL training😆. Anyway, congrats and Happy New Year!",
    "2588200": "Congrats! Thanks for the nice and clean explanation. Would it be possible to share the training notebook/code as well?",
    "2604702": "Hi @m1dsolo Can you provide your CTransPath training book I tried to emulate your work from inference and github but unable to get it right somehow. Thanks in advance.\n"
  }
}