{
  "id": 465356,
  "title": "20th Place Solution - UBC-OCEAN",
  "url": "/competitions/UBC-OCEAN/discussion/465356",
  "author_name": "Bartley",
  "post_date": "2024-01-04T00:10:52.799000",
  "votes": 19,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Although I am posting the write-up, this was a great team effort by <a href=\"https://www.kaggle.com/kevin0912\" target=\"_blank\">@kevin0912</a> and me. Also, thanks to UBC for hosting this competition, it was a fun competition, and interesting working with such large images! </p>\n<p>Our solution is based on a multiple instance learning (MIL) architecture with attention pooling. We use an ensemble of <code>efficientnet_b2</code>, <code>tf_efficientnetv2_b2.in1k</code> and <code>regnety_016.tv2_in1k</code> backbones trained on sequences of 8 x 1280 x 1280 images, and ignore the <code>other</code> class. We also apply light TTA during inference (rot90, flips, transpose, random image order).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fcd22d830c5e4d19b09cb8af0a71a65be%2Farchitecture.jpg?generation=1704326896955132&amp;alt=media\" alt=\"Cropper\"></p>\n<h2>Strategies</h2>\n<p><strong>Efficient Tiling</strong></p>\n<p>We select tiles from WSIs based on the darkest median pixel value. To make the pipeline more efficient, we use multiprocessing on 3 CPU cores, and prefilter crop locations using the smaller thumbnail images. This prefiltering selects the largest area of tissue on the slide and ignores other smaller areas of tissue.</p>\n<p>For TMAs, we take 5 central crops of size 2560 x 2560 and resize to 1280 x 1280 to match WSI magnification. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F983d5d3df388060aaceed59f780c2f1e%2Fsmart_cropper.jpg?generation=1704327003387823&amp;alt=media\" alt=\"Cropper\"></p>\n<p>Although efficient, a limitation of the pipeline is that it may not extract informative tiles from each image. We also experimented with a lightweight tile classifier trained on the ~150 segmentation masks, but this did not improve tile selection.</p>\n<p><strong>Modeling</strong></p>\n<p>We trained each model for 20-30 epochs with heavy augmentations and SWA (Stochastic Weight Averaging). Most models were trained on all the WSIs and TMAs, but some were trained using synthetically generated TMAs (aka. TMA Planets) from the <a href=\"https://www.kaggle.com/datasets/sohier/ubc-ovarian-cancer-competition-supplemental-masks\" target=\"_blank\">supplemental masks</a>. We would likely have explored TMA planets further but we were skeptical of the mask quality, and low count relative to the total number of WSIs.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F63447a1f1ace988b300842a49ec40c53%2Ftma_planet.JPG?generation=1704326966847250&amp;alt=media\" alt=\"Cropper\"></p>\n<p><strong>OOF Relabel + Remove</strong></p>\n<p>Based on <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/445804#2559062\" target=\"_blank\">Noli Alonso's comments</a>, we removed ~5% of the images and relabelled 8 images. We used a similar denoising method to that in the <a href=\"https://www.kaggle.com/competitions/prostate-cancer-grade-assessment/discussion/169143\" target=\"_blank\">1st place solution</a> of the <a href=\"https://www.kaggle.com/competitions/prostate-cancer-grade-assessment/overview\" target=\"_blank\">PANDA Competition</a>.</p>\n<pre><code>relabel_dict = {\n    '3': 'MC',\n    '5': 'LGSC', \n    '2': 'CC',\n    '8': 'LGSC',\n    '9': 'MC',\n    '7': 'EC',\n    '4': 'CC',\n    '6': 'LGSC',\n}\n</code></pre>\n<h2>External Data</h2>\n<p>The only external dataset we used was the <a href=\"https://www.medicalimageanalysis.com/data/ovarian-carcinomas-histopathology-dataset\" target=\"_blank\">Ovarian Carcinoma Histopathology Dataset (SFU)</a>. This dataset had 80 WSIs at 40x magnification from 6 different pathology centers.</p>\n<p>Class distribution: <code>{'HGSC': 30, 'CC': 20, 'EC': 11, 'MC': 10, 'LGSC': 9}</code></p>\n<h2>Did not work for us</h2>\n<ul>\n<li>Larger backbones</li>\n<li>Lightweight tile classifier</li>\n<li>Stain normalization (staintools, stainnet, etc.)</li>\n<li>JPGs</li>\n</ul>\n<h2>Frameworks</h2>\n<ul>\n<li><a href=\"https://lightning.ai/docs/pytorch/stable/\" target=\"_blank\">Pytorch Lightning</a> (training)</li>\n<li><a href=\"https://wandb.ai/site\" target=\"_blank\">Weights + Biases</a> (logging)</li>\n<li><a href=\"https://huggingface.co/timm\" target=\"_blank\">Timm</a> (backbones)</li>\n</ul>",
  "messages": [
    {
      "id": 2586057,
      "postDate": "2024-01-04T00:10:52.800Z",
      "content": "<p>Although I am posting the write-up, this was a great team effort by <a href=\"https://www.kaggle.com/kevin0912\" target=\"_blank\">@kevin0912</a> and me. Also, thanks to UBC for hosting this competition, it was a fun competition, and interesting working with such large images! </p>\n<p>Our solution is based on a multiple instance learning (MIL) architecture with attention pooling. We use an ensemble of <code>efficientnet_b2</code>, <code>tf_efficientnetv2_b2.in1k</code> and <code>regnety_016.tv2_in1k</code> backbones trained on sequences of 8 x 1280 x 1280 images, and ignore the <code>other</code> class. We also apply light TTA during inference (rot90, flips, transpose, random image order).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fcd22d830c5e4d19b09cb8af0a71a65be%2Farchitecture.jpg?generation=1704326896955132&amp;alt=media\" alt=\"Cropper\"></p>\n<h2>Strategies</h2>\n<p><strong>Efficient Tiling</strong></p>\n<p>We select tiles from WSIs based on the darkest median pixel value. To make the pipeline more efficient, we use multiprocessing on 3 CPU cores, and prefilter crop locations using the smaller thumbnail images. This prefiltering selects the largest area of tissue on the slide and ignores other smaller areas of tissue.</p>\n<p>For TMAs, we take 5 central crops of size 2560 x 2560 and resize to 1280 x 1280 to match WSI magnification. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F983d5d3df388060aaceed59f780c2f1e%2Fsmart_cropper.jpg?generation=1704327003387823&amp;alt=media\" alt=\"Cropper\"></p>\n<p>Although efficient, a limitation of the pipeline is that it may not extract informative tiles from each image. We also experimented with a lightweight tile classifier trained on the ~150 segmentation masks, but this did not improve tile selection.</p>\n<p><strong>Modeling</strong></p>\n<p>We trained each model for 20-30 epochs with heavy augmentations and SWA (Stochastic Weight Averaging). Most models were trained on all the WSIs and TMAs, but some were trained using synthetically generated TMAs (aka. TMA Planets) from the <a href=\"https://www.kaggle.com/datasets/sohier/ubc-ovarian-cancer-competition-supplemental-masks\" target=\"_blank\">supplemental masks</a>. We would likely have explored TMA planets further but we were skeptical of the mask quality, and low count relative to the total number of WSIs.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F63447a1f1ace988b300842a49ec40c53%2Ftma_planet.JPG?generation=1704326966847250&amp;alt=media\" alt=\"Cropper\"></p>\n<p><strong>OOF Relabel + Remove</strong></p>\n<p>Based on <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/445804#2559062\" target=\"_blank\">Noli Alonso's comments</a>, we removed ~5% of the images and relabelled 8 images. We used a similar denoising method to that in the <a href=\"https://www.kaggle.com/competitions/prostate-cancer-grade-assessment/discussion/169143\" target=\"_blank\">1st place solution</a> of the <a href=\"https://www.kaggle.com/competitions/prostate-cancer-grade-assessment/overview\" target=\"_blank\">PANDA Competition</a>.</p>\n<pre><code>relabel_dict = {\n    '3': 'MC',\n    '5': 'LGSC', \n    '2': 'CC',\n    '8': 'LGSC',\n    '9': 'MC',\n    '7': 'EC',\n    '4': 'CC',\n    '6': 'LGSC',\n}\n</code></pre>\n<h2>External Data</h2>\n<p>The only external dataset we used was the <a href=\"https://www.medicalimageanalysis.com/data/ovarian-carcinomas-histopathology-dataset\" target=\"_blank\">Ovarian Carcinoma Histopathology Dataset (SFU)</a>. This dataset had 80 WSIs at 40x magnification from 6 different pathology centers.</p>\n<p>Class distribution: <code>{'HGSC': 30, 'CC': 20, 'EC': 11, 'MC': 10, 'LGSC': 9}</code></p>\n<h2>Did not work for us</h2>\n<ul>\n<li>Larger backbones</li>\n<li>Lightweight tile classifier</li>\n<li>Stain normalization (staintools, stainnet, etc.)</li>\n<li>JPGs</li>\n</ul>\n<h2>Frameworks</h2>\n<ul>\n<li><a href=\"https://lightning.ai/docs/pytorch/stable/\" target=\"_blank\">Pytorch Lightning</a> (training)</li>\n<li><a href=\"https://wandb.ai/site\" target=\"_blank\">Weights + Biases</a> (logging)</li>\n<li><a href=\"https://huggingface.co/timm\" target=\"_blank\">Timm</a> (backbones)</li>\n</ul>",
      "rawMarkdown": "Although I am posting the write-up, this was a great team effort by @kevin0912 and me. Also, thanks to UBC for hosting this competition, it was a fun competition, and interesting working with such large images! \n\nOur solution is based on a multiple instance learning (MIL) architecture with attention pooling. We use an ensemble of `efficientnet_b2`, `tf_efficientnetv2_b2.in1k` and `regnety_016.tv2_in1k` backbones trained on sequences of 8 x 1280 x 1280 images, and ignore the `other` class. We also apply light TTA during inference (rot90, flips, transpose, random image order).\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fcd22d830c5e4d19b09cb8af0a71a65be%2Farchitecture.jpg?generation=1704326896955132&alt=media\" alt=\"Cropper\" style=\"max-width: 75%;\">\n\n## Strategies\n\n**Efficient Tiling**\n\nWe select tiles from WSIs based on the darkest median pixel value. To make the pipeline more efficient, we use multiprocessing on 3 CPU cores, and prefilter crop locations using the smaller thumbnail images. This prefiltering selects the largest area of tissue on the slide and ignores other smaller areas of tissue.\n\nFor TMAs, we take 5 central crops of size 2560 x 2560 and resize to 1280 x 1280 to match WSI magnification. \n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F983d5d3df388060aaceed59f780c2f1e%2Fsmart_cropper.jpg?generation=1704327003387823&alt=media\" alt=\"Cropper\" style=\"max-width: 25%;\">\n\nAlthough efficient, a limitation of the pipeline is that it may not extract informative tiles from each image. We also experimented with a lightweight tile classifier trained on the ~150 segmentation masks, but this did not improve tile selection.\n\n**Modeling**\n\nWe trained each model for 20-30 epochs with heavy augmentations and SWA (Stochastic Weight Averaging). Most models were trained on all the WSIs and TMAs, but some were trained using synthetically generated TMAs (aka. TMA Planets) from the [supplemental masks](https://www.kaggle.com/datasets/sohier/ubc-ovarian-cancer-competition-supplemental-masks). We would likely have explored TMA planets further but we were skeptical of the mask quality, and low count relative to the total number of WSIs.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F63447a1f1ace988b300842a49ec40c53%2Ftma_planet.JPG?generation=1704326966847250&alt=media\" alt=\"Cropper\" style=\"max-width: 25%;\">\n\n**OOF Relabel + Remove**\n\nBased on [Noli Alonso's comments](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/445804#2559062), we removed ~5% of the images and relabelled 8 images. We used a similar denoising method to that in the [1st place solution](https://www.kaggle.com/competitions/prostate-cancer-grade-assessment/discussion/169143) of the [PANDA Competition](https://www.kaggle.com/competitions/prostate-cancer-grade-assessment/overview).\n\n```\nrelabel_dict = {\n    '15583': 'MC',\n    '51215': 'LGSC', \n    '21432': 'CC',\n    '50878': 'LGSC',\n    '19569': 'MC',\n    '38097': 'EC',\n    '29084': 'CC',\n    '63836': 'LGSC',\n}\n```\n\n## External Data\n\nThe only external dataset we used was the [Ovarian Carcinoma Histopathology Dataset (SFU)](https://www.medicalimageanalysis.com/data/ovarian-carcinomas-histopathology-dataset). This dataset had 80 WSIs at 40x magnification from 6 different pathology centers.\n\nClass distribution: `{'HGSC': 30, 'CC': 20, 'EC': 11, 'MC': 10, 'LGSC': 9}`\n\n## Did not work for us\n\n- Larger backbones\n- Lightweight tile classifier\n- Stain normalization (staintools, stainnet, etc.)\n- JPGs\n\n## Frameworks\n\n- [Pytorch Lightning](https://lightning.ai/docs/pytorch/stable/) (training)\n- [Weights + Biases](https://wandb.ai/site) (logging)\n- [Timm](https://huggingface.co/timm) (backbones)\n",
      "votes": 19
    },
    {
      "id": 2827499,
      "postDate": "2024-05-21T14:14:53.910Z",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> I have a query regarding the Ovarian Carcinoma Histopathology Dataset (SFU) you used. I'm experiencing bandwidth issues while downloading the dataset, with frequent crashes and speed limitations. Did you encounter similar issues, and if so, do you know of any alternative sources to download the data? I'm working on an Ovarian Cancer model and urgently need the dataset validation.</p>",
      "rawMarkdown": "@brendanartley I have a query regarding the Ovarian Carcinoma Histopathology Dataset (SFU) you used. I'm experiencing bandwidth issues while downloading the dataset, with frequent crashes and speed limitations. Did you encounter similar issues, and if so, do you know of any alternative sources to download the data? I'm working on an Ovarian Cancer model and urgently need the dataset validation.",
      "votes": 1,
      "replies": [
        {
          "id": 2843387,
          "postDate": "2024-05-29T14:48:19.933Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/usmansafdar09\" target=\"_blank\">@usmansafdar09</a>. We downloaded the data from <a href=\"https://mia-www.cs.sfu.ca/download/\" target=\"_blank\">here</a>, but you can also try from <a href=\"https://mia-www.cs.sfu.ca/download/\" target=\"_blank\">here</a> as well.</p>",
          "rawMarkdown": "Hi @usmansafdar09. We downloaded the data from [here](https://mia-www.cs.sfu.ca/download/), but you can also try from [here](https://mia-www.cs.sfu.ca/download/) as well."
        }
      ]
    },
    {
      "id": 2586956,
      "postDate": "2024-01-04T13:46:10.967Z",
      "content": "<p>Congratulations！Could you share the code for Efficient Tiling？</p>",
      "rawMarkdown": "Congratulations！Could you share the code for Efficient Tiling？",
      "votes": 1,
      "replies": [
        {
          "id": 2587089,
          "postDate": "2024-01-04T15:01:52.413Z",
          "content": "<p>Sure, I just created a notebook <a href=\"https://www.kaggle.com/code/brendanartley/ubco-efficient-tiling-code/notebook\" target=\"_blank\">here</a> with the tiling code. <a href=\"https://www.kaggle.com/seeingtimes\" target=\"_blank\">@seeingtimes</a>, <a href=\"https://www.kaggle.com/samu2505\" target=\"_blank\">@samu2505</a>, <a href=\"https://www.kaggle.com/jirkaborovec\" target=\"_blank\">@jirkaborovec</a>.</p>",
          "rawMarkdown": "Sure, I just created a notebook [here](https://www.kaggle.com/code/brendanartley/ubco-efficient-tiling-code/notebook) with the tiling code. @seeingtimes, @samu2505, @jirkaborovec.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2586801,
      "postDate": "2024-01-04T12:12:29.287Z",
      "content": "<p>Awesome to see using lightning with great ranking!</p>\n<blockquote>\n  <p>To make the pipeline more efficient, we use multiprocessing on 3 CPU cores, and prefilter crop locations using the smaller thumbnail images.</p>\n</blockquote>\n<p>Was thinking maybe if the sampling would be sequential tile-by-tile not always specific positions it would be also much faster, right?</p>",
      "rawMarkdown": "Awesome to see using lightning with great ranking!\n\n> To make the pipeline more efficient, we use multiprocessing on 3 CPU cores, and prefilter crop locations using the smaller thumbnail images.\n\nWas thinking maybe if the sampling would be sequential tile-by-tile not always specific positions it would be also much faster, right?",
      "votes": 1,
      "replies": [
        {
          "id": 2587080,
          "postDate": "2024-01-04T14:55:47.780Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/jirkaborovec\" target=\"_blank\">@jirkaborovec</a> for all your contributions to lightning! It is a great package.</p>\n<p>We did select the tiles somewhat sequentially by creating a list of sequential x,y pairs, and then iterating over the list and calling the <code>.crop()</code> function. Not sure how the processing time would change if we shuffled the pairs.. </p>",
          "rawMarkdown": "Thanks @jirkaborovec for all your contributions to lightning! It is a great package.\n\nWe did select the tiles somewhat sequentially by creating a list of sequential x,y pairs, and then iterating over the list and calling the `.crop()` function. Not sure how the processing time would change if we shuffled the pairs.. "
        }
      ]
    },
    {
      "id": 2586404,
      "postDate": "2024-01-04T07:23:34.227Z",
      "content": "<p>Thank you for sharing the solution. <br>\nI have a question to ensure my understanding is correct. <br>\nSince our image has 3 channels, <br>\nshould the input be in the format [batch * 24 * 1280 * 1280]? Thank you</p>",
      "rawMarkdown": "Thank you for sharing the solution. \nI have a question to ensure my understanding is correct. \nSince our image has 3 channels, \nshould the input be in the format [batch * 24 * 1280 * 1280]? Thank you",
      "votes": 1,
      "replies": [
        {
          "id": 2587063,
          "postDate": "2024-01-04T14:39:26.810Z",
          "content": "<p>Close! The input_shape during training is <code>batch_size x 8 x 3 x 1280 x 1280</code>. We used a package called <a href=\"https://einops.rocks/\" target=\"_blank\">einops</a> to reshape the tensors in the <code>forward</code> function. Something like this.</p>\n<pre><code>def (self, x):\n    b, t = x.()[:]\n\n    # Feature extractor\n    x = (x, )\n    x = self.(x)\n    x = self.(x)\n    x = (x, , b=b, t=t)\n\n    # Attention Pooling\n    a = self.(x)\n    a = torch.(a, dim=)\n    x = torch.(x * a, dim=)\n    x = self.(x)\n    return x\n</code></pre>",
          "rawMarkdown": "Close! The input_shape during training is `batch_size x 8 x 3 x 1280 x 1280`. We used a package called [einops](https://einops.rocks/) to reshape the tensors in the `forward` function. Something like this.\n\n```\ndef forward(self, x):\n    b, t = x.size()[:2]\n\n    # Feature extractor\n    x = rearrange(x, \"b t c h w -> (b t) c h w\")\n    x = self.backbone(x)\n    x = self.dropout(x)\n    x = rearrange(x, \"(b t) f -> b t f\", b=b, t=t)\n\n    # Attention Pooling\n    a = self.attention(x)\n    a = torch.softmax(a, dim=1)\n    x = torch.sum(x * a, dim=1)\n    x = self.fc(x)\n    return x\n```",
          "votes": 1
        }
      ]
    },
    {
      "id": 2586226,
      "postDate": "2024-01-04T04:03:52.387Z",
      "content": "<p>Congratulations on getting the 22nd rank in this competition.  Thanks for sharing the details of your solution with nice diagrams. Did you tried to include the sixth class i.e., \"other\".</p>",
      "rawMarkdown": "Congratulations on getting the 22nd rank in this competition.  Thanks for sharing the details of your solution with nice diagrams. Did you tried to include the sixth class i.e., \"other\".",
      "votes": 1,
      "replies": [
        {
          "id": 2586256,
          "postDate": "2024-01-04T04:54:54.147Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/crsuthikshnkumar\" target=\"_blank\">@crsuthikshnkumar</a>! We tried predicting the “other” class using prediction variance and a high threshold, but it did not work well. In the end we just ignored the “other” class.</p>",
          "rawMarkdown": "Thanks @crsuthikshnkumar! We tried predicting the “other” class using prediction variance and a high threshold, but it did not work well. In the end we just ignored the “other” class."
        }
      ]
    },
    {
      "id": 2586210,
      "postDate": "2024-01-04T03:53:23.060Z",
      "content": "<p>Congratulation on your silver medal! May I ask the scores of using \"OOF Relabel + Remove\" and without it?</p>",
      "rawMarkdown": "Congratulation on your silver medal! May I ask the scores of using \"OOF Relabel + Remove\" and without it?",
      "votes": 1,
      "replies": [
        {
          "id": 2586247,
          "postDate": "2024-01-04T04:47:36.370Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/huyduong7101\" target=\"_blank\">@huyduong7101</a>! </p>\n<p>We had submissions reach 0.54 with and without this. I suspect our ensemble of 16 models helped overcome the noisy labels.</p>",
          "rawMarkdown": "Thanks @huyduong7101! \n\nWe had submissions reach 0.54 with and without this. I suspect our ensemble of 16 models helped overcome the noisy labels."
        }
      ]
    },
    {
      "id": 2586120,
      "postDate": "2024-01-04T01:44:13.410Z",
      "content": "<p>would you mind sharing (code/pseudocode) the tiling strategy that you used</p>",
      "rawMarkdown": "would you mind sharing (code/pseudocode) the tiling strategy that you used",
      "votes": 1,
      "replies": [
        {
          "id": 2586180,
          "postDate": "2024-01-04T03:28:09.233Z",
          "content": "<p>Sure! Here is the pseudocode.</p>\n<pre><code>. Load thumbnail \n. Select areas (x,y) of the  that are cell tissue\n. Convert (x,y) pairs to full  \n. Load  tile   the  pixel value\n. Select the top  tiles with the darkest  pixel \n</code></pre>\n<p>This process took ~4hrs to save all tiles on submission.</p>",
          "rawMarkdown": "Sure! Here is the pseudocode.\n\n```\n1. Load thumbnail image\n2. Select areas (x,y) of the image that are cell tissue\n3. Convert (x,y) pairs to full image scale\n4. Load every tile and save the median pixel value\n5. Select the top 14 tiles with the darkest median pixel values\n```\n\nThis process took ~4hrs to save all tiles on submission."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2827499,
      "author_name": "USMAN SAFDAR",
      "author_url": "",
      "post_date": "2024-05-21T14:14:53.910000",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> I have a query regarding the Ovarian Carcinoma Histopathology Dataset (SFU) you used. I'm experiencing bandwidth issues while downloading the dataset, with frequent crashes and speed limitations. Did you encounter similar issues, and if so, do you know of any alternative sources to download the data? I'm working on an Ovarian Cancer model and urgently need the dataset validation.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2843387,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-05-29T14:48:19.933000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/usmansafdar09\" target=\"_blank\">@usmansafdar09</a>. We downloaded the data from <a href=\"https://mia-www.cs.sfu.ca/download/\" target=\"_blank\">here</a>, but you can also try from <a href=\"https://mia-www.cs.sfu.ca/download/\" target=\"_blank\">here</a> as well.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2586956,
      "author_name": "Seeing Times",
      "author_url": "",
      "post_date": "2024-01-04T13:46:10.967000",
      "content": "<p>Congratulations！Could you share the code for Efficient Tiling？</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2587089,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-01-04T15:01:52.413000",
          "content": "<p>Sure, I just created a notebook <a href=\"https://www.kaggle.com/code/brendanartley/ubco-efficient-tiling-code/notebook\" target=\"_blank\">here</a> with the tiling code. <a href=\"https://www.kaggle.com/seeingtimes\" target=\"_blank\">@seeingtimes</a>, <a href=\"https://www.kaggle.com/samu2505\" target=\"_blank\">@samu2505</a>, <a href=\"https://www.kaggle.com/jirkaborovec\" target=\"_blank\">@jirkaborovec</a>.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2586801,
      "author_name": "Jirka",
      "author_url": "",
      "post_date": "2024-01-04T12:12:29.287000",
      "content": "<p>Awesome to see using lightning with great ranking!</p>\n<blockquote>\n  <p>To make the pipeline more efficient, we use multiprocessing on 3 CPU cores, and prefilter crop locations using the smaller thumbnail images.</p>\n</blockquote>\n<p>Was thinking maybe if the sampling would be sequential tile-by-tile not always specific positions it would be also much faster, right?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2587080,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-01-04T14:55:47.780000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/jirkaborovec\" target=\"_blank\">@jirkaborovec</a> for all your contributions to lightning! It is a great package.</p>\n<p>We did select the tiles somewhat sequentially by creating a list of sequential x,y pairs, and then iterating over the list and calling the <code>.crop()</code> function. Not sure how the processing time would change if we shuffled the pairs.. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2586404,
      "author_name": "ParkSom",
      "author_url": "",
      "post_date": "2024-01-04T07:23:34.227000",
      "content": "<p>Thank you for sharing the solution. <br>\nI have a question to ensure my understanding is correct. <br>\nSince our image has 3 channels, <br>\nshould the input be in the format [batch * 24 * 1280 * 1280]? Thank you</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2587063,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-01-04T14:39:26.810000",
          "content": "<p>Close! The input_shape during training is <code>batch_size x 8 x 3 x 1280 x 1280</code>. We used a package called <a href=\"https://einops.rocks/\" target=\"_blank\">einops</a> to reshape the tensors in the <code>forward</code> function. Something like this.</p>\n<pre><code>def (self, x):\n    b, t = x.()[:]\n\n    # Feature extractor\n    x = (x, )\n    x = self.(x)\n    x = self.(x)\n    x = (x, , b=b, t=t)\n\n    # Attention Pooling\n    a = self.(x)\n    a = torch.(a, dim=)\n    x = torch.(x * a, dim=)\n    x = self.(x)\n    return x\n</code></pre>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2586226,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2024-01-04T04:03:52.387000",
      "content": "<p>Congratulations on getting the 22nd rank in this competition.  Thanks for sharing the details of your solution with nice diagrams. Did you tried to include the sixth class i.e., \"other\".</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2586256,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-01-04T04:54:54.147000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/crsuthikshnkumar\" target=\"_blank\">@crsuthikshnkumar</a>! We tried predicting the “other” class using prediction variance and a high threshold, but it did not work well. In the end we just ignored the “other” class.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2586210,
      "author_name": "Dive Deeper",
      "author_url": "",
      "post_date": "2024-01-04T03:53:23.060000",
      "content": "<p>Congratulation on your silver medal! May I ask the scores of using \"OOF Relabel + Remove\" and without it?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2586247,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-01-04T04:47:36.370000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/huyduong7101\" target=\"_blank\">@huyduong7101</a>! </p>\n<p>We had submissions reach 0.54 with and without this. I suspect our ensemble of 16 models helped overcome the noisy labels.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2586120,
      "author_name": "samu2505",
      "author_url": "",
      "post_date": "2024-01-04T01:44:13.410000",
      "content": "<p>would you mind sharing (code/pseudocode) the tiling strategy that you used</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2586180,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-01-04T03:28:09.233000",
          "content": "<p>Sure! Here is the pseudocode.</p>\n<pre><code>. Load thumbnail \n. Select areas (x,y) of the  that are cell tissue\n. Convert (x,y) pairs to full  \n. Load  tile   the  pixel value\n. Select the top  tiles with the darkest  pixel \n</code></pre>\n<p>This process took ~4hrs to save all tiles on submission.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2586057": "Although I am posting the write-up, this was a great team effort by @kevin0912 and me. Also, thanks to UBC for hosting this competition, it was a fun competition, and interesting working with such large images! \n\nOur solution is based on a multiple instance learning (MIL) architecture with attention pooling. We use an ensemble of `efficientnet_b2`, `tf_efficientnetv2_b2.in1k` and `regnety_016.tv2_in1k` backbones trained on sequences of 8 x 1280 x 1280 images, and ignore the `other` class. We also apply light TTA during inference (rot90, flips, transpose, random image order).\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fcd22d830c5e4d19b09cb8af0a71a65be%2Farchitecture.jpg?generation=1704326896955132&alt=media\" alt=\"Cropper\" style=\"max-width: 75%;\">\n\n## Strategies\n\n**Efficient Tiling**\n\nWe select tiles from WSIs based on the darkest median pixel value. To make the pipeline more efficient, we use multiprocessing on 3 CPU cores, and prefilter crop locations using the smaller thumbnail images. This prefiltering selects the largest area of tissue on the slide and ignores other smaller areas of tissue.\n\nFor TMAs, we take 5 central crops of size 2560 x 2560 and resize to 1280 x 1280 to match WSI magnification. \n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F983d5d3df388060aaceed59f780c2f1e%2Fsmart_cropper.jpg?generation=1704327003387823&alt=media\" alt=\"Cropper\" style=\"max-width: 25%;\">\n\nAlthough efficient, a limitation of the pipeline is that it may not extract informative tiles from each image. We also experimented with a lightweight tile classifier trained on the ~150 segmentation masks, but this did not improve tile selection.\n\n**Modeling**\n\nWe trained each model for 20-30 epochs with heavy augmentations and SWA (Stochastic Weight Averaging). Most models were trained on all the WSIs and TMAs, but some were trained using synthetically generated TMAs (aka. TMA Planets) from the [supplemental masks](https://www.kaggle.com/datasets/sohier/ubc-ovarian-cancer-competition-supplemental-masks). We would likely have explored TMA planets further but we were skeptical of the mask quality, and low count relative to the total number of WSIs.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F63447a1f1ace988b300842a49ec40c53%2Ftma_planet.JPG?generation=1704326966847250&alt=media\" alt=\"Cropper\" style=\"max-width: 25%;\">\n\n**OOF Relabel + Remove**\n\nBased on [Noli Alonso's comments](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/445804#2559062), we removed ~5% of the images and relabelled 8 images. We used a similar denoising method to that in the [1st place solution](https://www.kaggle.com/competitions/prostate-cancer-grade-assessment/discussion/169143) of the [PANDA Competition](https://www.kaggle.com/competitions/prostate-cancer-grade-assessment/overview).\n\n```\nrelabel_dict = {\n    '15583': 'MC',\n    '51215': 'LGSC', \n    '21432': 'CC',\n    '50878': 'LGSC',\n    '19569': 'MC',\n    '38097': 'EC',\n    '29084': 'CC',\n    '63836': 'LGSC',\n}\n```\n\n## External Data\n\nThe only external dataset we used was the [Ovarian Carcinoma Histopathology Dataset (SFU)](https://www.medicalimageanalysis.com/data/ovarian-carcinomas-histopathology-dataset). This dataset had 80 WSIs at 40x magnification from 6 different pathology centers.\n\nClass distribution: `{'HGSC': 30, 'CC': 20, 'EC': 11, 'MC': 10, 'LGSC': 9}`\n\n## Did not work for us\n\n- Larger backbones\n- Lightweight tile classifier\n- Stain normalization (staintools, stainnet, etc.)\n- JPGs\n\n## Frameworks\n\n- [Pytorch Lightning](https://lightning.ai/docs/pytorch/stable/) (training)\n- [Weights + Biases](https://wandb.ai/site) (logging)\n- [Timm](https://huggingface.co/timm) (backbones)\n",
    "2827499": "@brendanartley I have a query regarding the Ovarian Carcinoma Histopathology Dataset (SFU) you used. I'm experiencing bandwidth issues while downloading the dataset, with frequent crashes and speed limitations. Did you encounter similar issues, and if so, do you know of any alternative sources to download the data? I'm working on an Ovarian Cancer model and urgently need the dataset validation.",
    "2586956": "Congratulations！Could you share the code for Efficient Tiling？",
    "2586801": "Awesome to see using lightning with great ranking!\n\n> To make the pipeline more efficient, we use multiprocessing on 3 CPU cores, and prefilter crop locations using the smaller thumbnail images.\n\nWas thinking maybe if the sampling would be sequential tile-by-tile not always specific positions it would be also much faster, right?",
    "2586404": "Thank you for sharing the solution. \nI have a question to ensure my understanding is correct. \nSince our image has 3 channels, \nshould the input be in the format [batch * 24 * 1280 * 1280]? Thank you",
    "2586226": "Congratulations on getting the 22nd rank in this competition.  Thanks for sharing the details of your solution with nice diagrams. Did you tried to include the sixth class i.e., \"other\".",
    "2586210": "Congratulation on your silver medal! May I ask the scores of using \"OOF Relabel + Remove\" and without it?",
    "2586120": "would you mind sharing (code/pseudocode) the tiling strategy that you used"
  }
}