{
  "id": 343665,
  "title": "Poor Results from Vision Transformers",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/343665",
  "author_name": "yqz",
  "post_date": "2022-08-12T04:13:25.448000",
  "votes": 17,
  "comment_count": 30,
  "views": 0,
  "content": "<p>Hi all, I've tried a few basic models like random forest (RF), and general CNN's like alexnet and resnet. I tried the ViT base model via HuggingFace, and the AUC score is 0.51/0.52 after 30 epochs, which is quite a bit lower than even RF with a 5-fold AUC of 0.5662. Is anyone else facing similar issues? It seems to me that the ViT can't learn the problem at all, and is similar to a random guess of 0.5.</p>",
  "messages": [
    {
      "id": 1895291,
      "postDate": "2022-08-12T04:13:25.450Z",
      "content": "<p>Hi all, I've tried a few basic models like random forest (RF), and general CNN's like alexnet and resnet. I tried the ViT base model via HuggingFace, and the AUC score is 0.51/0.52 after 30 epochs, which is quite a bit lower than even RF with a 5-fold AUC of 0.5662. Is anyone else facing similar issues? It seems to me that the ViT can't learn the problem at all, and is similar to a random guess of 0.5.</p>",
      "rawMarkdown": "Hi all, I've tried a few basic models like random forest (RF), and general CNN's like alexnet and resnet. I tried the ViT base model via HuggingFace, and the AUC score is 0.51/0.52 after 30 epochs, which is quite a bit lower than even RF with a 5-fold AUC of 0.5662. Is anyone else facing similar issues? It seems to me that the ViT can't learn the problem at all, and is similar to a random guess of 0.5.",
      "votes": 17
    },
    {
      "id": 1907091,
      "postDate": "2022-08-20T12:46:57.577Z",
      "content": "<p>Improve Vision Transformers Training by Suppressing Over-smoothing, May be this research paper would help in solving the issue. when we train directly it yields unstable results to take care of this some modification is required in model by incorporating some conv layer …\" <a href=\"https://www.arxiv-vanity.com/papers/2104.12753/\" target=\"_blank\">https://www.arxiv-vanity.com/papers/2104.12753/</a> \". may be it could help :)</p>",
      "rawMarkdown": "Improve Vision Transformers Training by Suppressing Over-smoothing, May be this research paper would help in solving the issue. when we train directly it yields unstable results to take care of this some modification is required in model by incorporating some conv layer ...\" https://www.arxiv-vanity.com/papers/2104.12753/ \". may be it could help :)",
      "votes": 1,
      "replies": [
        {
          "id": 1907547,
          "postDate": "2022-08-20T21:57:47.230Z",
          "content": "<p>Awesome, thanks for the idea. I will look into this.</p>",
          "rawMarkdown": "Awesome, thanks for the idea. I will look into this."
        }
      ]
    },
    {
      "id": 1906423,
      "postDate": "2022-08-19T21:37:03.943Z",
      "content": "<p>I get the same feeling from nearly all models.. </p>",
      "rawMarkdown": "I get the same feeling from nearly all models.. \n\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 1906427,
          "postDate": "2022-08-19T21:48:26.677Z",
          "content": "<p>Have you tried weighted random sampling? </p>",
          "rawMarkdown": "Have you tried weighted random sampling? ",
          "votes": 2
        },
        {
          "id": 1906452,
          "postDate": "2022-08-19T22:36:03.850Z",
          "content": "<p><a href=\"https://www.kaggle.com/jcerpentier\" target=\"_blank\">@jcerpentier</a> I did it initially with a resnet18 model, which so far has my best AUC score. The resulting score was a bit lower surprisingly. I had the weighted random sampler as something on my to-do list, but hadn't yet got around to working on until you mentioned it. I plan on trying it later with some more models though.</p>",
          "rawMarkdown": "@jcerpentier I did it initially with a resnet18 model, which so far has my best AUC score. The resulting score was a bit lower surprisingly. I had the weighted random sampler as something on my to-do list, but hadn't yet got around to working on until you mentioned it. I plan on trying it later with some more models though.",
          "votes": 1
        },
        {
          "id": 1907667,
          "postDate": "2022-08-21T00:50:07.387Z",
          "content": "<p><a href=\"https://www.kaggle.com/yuqizheng\" target=\"_blank\">@yuqizheng</a> Did you still tile the original images when using this approach?</p>",
          "rawMarkdown": "@yuqizheng Did you still tile the original images when using this approach?",
          "votes": 1
        },
        {
          "id": 1907700,
          "postDate": "2022-08-21T01:58:28.093Z",
          "content": "<p><a href=\"https://www.kaggle.com/hjunlee941\" target=\"_blank\">@hjunlee941</a> No, do you have a source for a good approach that utilizes tiling? I'm trying to work on a good pipeline right now for tiles.</p>",
          "rawMarkdown": "@hjunlee941 No, do you have a source for a good approach that utilizes tiling? I'm trying to work on a good pipeline right now for tiles."
        },
        {
          "id": 1907709,
          "postDate": "2022-08-21T02:19:32.490Z",
          "content": "<p><a href=\"https://www.kaggle.com/yuqizheng\" target=\"_blank\">@yuqizheng</a> I literally scratch built the pipeline to use patches. However, I still cannot find any good way to only remove background tiles.</p>",
          "rawMarkdown": "@yuqizheng I literally scratch built the pipeline to use patches. However, I still cannot find any good way to only remove background tiles.",
          "votes": 1
        },
        {
          "id": 1907721,
          "postDate": "2022-08-21T02:50:35.807Z",
          "content": "<p><a href=\"https://www.kaggle.com/hjunlee941\" target=\"_blank\">@hjunlee941</a> For background removal, a method that I'm using is as follows. Say you're reading patches of a WSI to save memory on your notebook (e.g., OpenSlide's read_region function). So for a given patch you read it of say size 512x512. You can first convert the RGB image to grayscale using Otsu's threshold (there are other choices for thresholding). From the grayscale image, you can count the ratio of black pixels to white pixels. If for example there are at least 50% or 90% of white pixels (indicating tissue regions), then you keep the tile (the untouched version that hasn't been given grayscale). I also noticed this can keep random regions with mostly whitespace, so you would probably also want to add another condition that there's at least 5% of black pixels in the image also.</p>\n<p>I think this method is quite different from what I've seen in many notebooks. I think many people are giving examples of first resizing the WSI image to something like 1024x1024, and then taking patches out of it. This however gives poor resolution patches that lack much information (based on just using my eye). I think the ideal method is to read the region, check if it has enough tissue area, and then resize the patch to something like 512x512.</p>\n<p>There are some other details too, like applying a sliding window with some multiplier (say 5) applied to the 512x512 patch window. This way the sliding window going through the 50k pixel WSI is 2560x2560, which in turn is the patch that is resized to 512x512.</p>\n<p>Sorry if this isn't super clear, I was thinking about sharing a notebook later once I got over the hurdle of ensuring that each image has at least X (e.g., 64) number of patches. When I did it without considering this, I had a minimum of 1 patch and maximum of ~500 patches for a given image.</p>",
          "rawMarkdown": "@hjunlee941 For background removal, a method that I'm using is as follows. Say you're reading patches of a WSI to save memory on your notebook (e.g., OpenSlide's read_region function). So for a given patch you read it of say size 512x512. You can first convert the RGB image to grayscale using Otsu's threshold (there are other choices for thresholding). From the grayscale image, you can count the ratio of black pixels to white pixels. If for example there are at least 50% or 90% of white pixels (indicating tissue regions), then you keep the tile (the untouched version that hasn't been given grayscale). I also noticed this can keep random regions with mostly whitespace, so you would probably also want to add another condition that there's at least 5% of black pixels in the image also.\n\nI think this method is quite different from what I've seen in many notebooks. I think many people are giving examples of first resizing the WSI image to something like 1024x1024, and then taking patches out of it. This however gives poor resolution patches that lack much information (based on just using my eye). I think the ideal method is to read the region, check if it has enough tissue area, and then resize the patch to something like 512x512.\n\nThere are some other details too, like applying a sliding window with some multiplier (say 5) applied to the 512x512 patch window. This way the sliding window going through the 50k pixel WSI is 2560x2560, which in turn is the patch that is resized to 512x512.\n\nSorry if this isn't super clear, I was thinking about sharing a notebook later once I got over the hurdle of ensuring that each image has at least X (e.g., 64) number of patches. When I did it without considering this, I had a minimum of 1 patch and maximum of ~500 patches for a given image.",
          "votes": 2
        },
        {
          "id": 1908488,
          "postDate": "2022-08-21T18:08:51.880Z",
          "content": "<p>I found <a href=\"https://www.kaggle.com/code/yasufuminakama/mayo-train-images-size-1024-n-16-1\" target=\"_blank\">this</a> notebook's tile creation very effective if you are looking at a top N tiles/patches sorted by signal in tile. This method was used effectively in previous competitions as well.</p>\n<p>Along with this I found that using unique value count threshold helped to remove less signal or mostly background tiles in a very simple manner. For e.g.,</p>\n<p>len(np.unique(tile)) &gt; 180 -- worked reasonably well, given tile is (s,s,3) shape. The threshold will probably depend on downscale factor and tile size. I guess this is ok if you are selecting top N tiles, but may be more challenging if you want to process all tiles from the image.</p>",
          "rawMarkdown": "I found [this](https://www.kaggle.com/code/yasufuminakama/mayo-train-images-size-1024-n-16-1) notebook's tile creation very effective if you are looking at a top N tiles/patches sorted by signal in tile. This method was used effectively in previous competitions as well.\n\nAlong with this I found that using unique value count threshold helped to remove less signal or mostly background tiles in a very simple manner. For e.g.,\n\nlen(np.unique(tile)) > 180 -- worked reasonably well, given tile is (s,s,3) shape. The threshold will probably depend on downscale factor and tile size. I guess this is ok if you are selecting top N tiles, but may be more challenging if you want to process all tiles from the image.",
          "votes": 2
        },
        {
          "id": 1909681,
          "postDate": "2022-08-22T20:43:12.307Z",
          "content": "<p>Something I've been eager to duplicate is the patching method found in this <a href=\"https://arxiv.org/pdf/1612.07180.pdf\" target=\"_blank\">paper</a>. I've included the image below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182873%2F03207d78b7078cfc5f335abfc03fd727%2FScreen%20Shot%202022-08-22%20at%201.37.06%20PM.png?generation=1661200689201901&amp;alt=media\" alt=\"\"><br>\nThey use a tool called CellProfiler, which is available on GitHub. I was wondering though if there's an ad-hoc way to implement it. Basically, for their breast cancer dataset, it makes sense for them to find the \"densest\" part of the image in terms of the tissue, since this is related to the cancer cells in breast cancer (if I read it correctly). I think if we can do something similar, by basically sampling from around the non-background regions and getting a set of e.g., N=64 patches then it would help us greatly.</p>\n<p>I tried a method to first read patches from the image, use Otsu's threshold to grayscale them. Then, I can determine which areas are tissue versus background. I tried to then randomly sample from these points, and filter the sampled points by first checking if there's at least e.g., 50% of white pixels in the image. The code works, but it's far too slow despite hours trying to vectorize and optimize the speed of the operations. Currently, it seems that even with the notebook you shared, the tiling can still at times have issues, giving patches that give little to no useful information.</p>",
          "rawMarkdown": "Something I've been eager to duplicate is the patching method found in this [paper](https://arxiv.org/pdf/1612.07180.pdf). I've included the image below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182873%2F03207d78b7078cfc5f335abfc03fd727%2FScreen%20Shot%202022-08-22%20at%201.37.06%20PM.png?generation=1661200689201901&alt=media)\nThey use a tool called CellProfiler, which is available on GitHub. I was wondering though if there's an ad-hoc way to implement it. Basically, for their breast cancer dataset, it makes sense for them to find the \"densest\" part of the image in terms of the tissue, since this is related to the cancer cells in breast cancer (if I read it correctly). I think if we can do something similar, by basically sampling from around the non-background regions and getting a set of e.g., N=64 patches then it would help us greatly.\n\nI tried a method to first read patches from the image, use Otsu's threshold to grayscale them. Then, I can determine which areas are tissue versus background. I tried to then randomly sample from these points, and filter the sampled points by first checking if there's at least e.g., 50% of white pixels in the image. The code works, but it's far too slow despite hours trying to vectorize and optimize the speed of the operations. Currently, it seems that even with the notebook you shared, the tiling can still at times have issues, giving patches that give little to no useful information.",
          "votes": 2
        },
        {
          "id": 1909729,
          "postDate": "2022-08-22T21:32:10.490Z",
          "content": "<p><a href=\"https://www.kaggle.com/yuqizheng\" target=\"_blank\">@yuqizheng</a> Does CellProfiler also work for non-white background? Since we have some very weird backgrounds in the training data, simple threshold based method may not work ideally.</p>",
          "rawMarkdown": "@yuqizheng Does CellProfiler also work for non-white background? Since we have some very weird backgrounds in the training data, simple threshold based method may not work ideally.",
          "votes": 1
        },
        {
          "id": 1909760,
          "postDate": "2022-08-22T22:44:01.903Z",
          "content": "<p>The base paper for CellProfiler is based an object identifier module.  It will be interesting to see if it can navigate a lot of noise in mayo dataset that could be mistaken for cell shaped features on the fringe of the scan area. Worth a try perhaps - the image above is very impressive.</p>\n<p><a href=\"https://www.kaggle.com/yuqizheng\" target=\"_blank\">@yuqizheng</a> I agree, normal tile creation is very frustratingly slow and not so discriminating amongst tiles. This surely needs an upgrade if downstream classification needs to go up a level. </p>",
          "rawMarkdown": "The base paper for CellProfiler is based an object identifier module.  It will be interesting to see if it can navigate a lot of noise in mayo dataset that could be mistaken for cell shaped features on the fringe of the scan area. Worth a try perhaps - the image above is very impressive.\n\n@yuqizheng I agree, normal tile creation is very frustratingly slow and not so discriminating amongst tiles. This surely needs an upgrade if downstream classification needs to go up a level. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1897204,
      "postDate": "2022-08-13T14:43:10.790Z",
      "content": "<p>I guess that the score is low cause the data is imbalanced. AUC is bad at imbalanced classes.</p>",
      "rawMarkdown": "I guess that the score is low cause the data is imbalanced. AUC is bad at imbalanced classes.",
      "votes": 1,
      "replies": [
        {
          "id": 1897462,
          "postDate": "2022-08-13T18:35:48.503Z",
          "content": "<p>I actually chose AUC based on the following quote from <a href=\"https://www.kaggle.com/abhishek\" target=\"_blank\">@abhishek</a>'s book, Approaching Almost Any Machine Learning Problem,</p>\n<blockquote>\n  <p>As a data doctor would say: this is a classic case of skewed binary classification. Therefore, we choose the evaluation metric to be AUC and go for a stratified k-fold cross-validation scheme.</p>\n</blockquote>\n<p>I'm not sure if you're aware, but the idea (or the one that I had in mind) I think is that since AUC is sensitive to imbalanced data in comparison to something like accuracy. If for example the data is 90% class A and 10% class B then just predicting class A all the time would give a score of 90%. However, since AUC factors in other things like sensitivity and false positives, it makes it more ideal for measuring the ability of a model to perform on imbalanced data.</p>\n<p>If you already knew what I said in the above paragraph you can go ahead and ignore it, didn't mean to lecture you on AUC.</p>",
          "rawMarkdown": "I actually chose AUC based on the following quote from @abhishek's book, Approaching Almost Any Machine Learning Problem,\n\n> As a data doctor would say: this is a classic case of skewed binary classification. Therefore, we choose the evaluation metric to be AUC and go for a stratified k-fold cross-validation scheme.\n\nI'm not sure if you're aware, but the idea (or the one that I had in mind) I think is that since AUC is sensitive to imbalanced data in comparison to something like accuracy. If for example the data is 90% class A and 10% class B then just predicting class A all the time would give a score of 90%. However, since AUC factors in other things like sensitivity and false positives, it makes it more ideal for measuring the ability of a model to perform on imbalanced data.\n\nIf you already knew what I said in the above paragraph you can go ahead and ignore it, didn't mean to lecture you on AUC.",
          "votes": 4
        },
        {
          "id": 1897983,
          "postDate": "2022-08-14T07:11:39.603Z",
          "content": "<p>What about applying F1-score?</p>",
          "rawMarkdown": "What about applying F1-score?",
          "votes": 1
        },
        {
          "id": 1898638,
          "postDate": "2022-08-14T17:34:31.010Z",
          "content": "<p>I haven't tried that method yet, but I'm just starting to get into Kaggle comps. I think it would make sense to create an array of metrics that can be tested epoch-wise and step-wise for something like PyTorch. That way we could gauge, using multiple methods, the effectiveness of a model on a dataset. This could include things like F1, AUC, accuracy, etc. Also, printing all the misclassifications would allow us to examine all the failures (and correct guesses) of the model to try and see why the model is performing well or not.</p>",
          "rawMarkdown": "I haven't tried that method yet, but I'm just starting to get into Kaggle comps. I think it would make sense to create an array of metrics that can be tested epoch-wise and step-wise for something like PyTorch. That way we could gauge, using multiple methods, the effectiveness of a model on a dataset. This could include things like F1, AUC, accuracy, etc. Also, printing all the misclassifications would allow us to examine all the failures (and correct guesses) of the model to try and see why the model is performing well or not.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1904746,
      "postDate": "2022-08-18T12:56:15.103Z",
      "content": "<p>From my experience so far, the data set is extremely difficult to learn. But the proper pre-processing/augmentation helps; I think there is definitely not a \"coin flip problem\" for this competition, as some might have suggested.</p>",
      "rawMarkdown": "From my experience so far, the data set is extremely difficult to learn. But the proper pre-processing/augmentation helps; I think there is definitely not a \"coin flip problem\" for this competition, as some might have suggested.",
      "votes": 2,
      "replies": [
        {
          "id": 1905029,
          "postDate": "2022-08-18T17:36:30.160Z",
          "content": "<p>May you please elaborate on what you mean by proper pre-processing/augmentation? I've been focusing on this data engineering/preprocessing side of things, but I'm bumping into problems. I've been thinking about something along the lines of the following for some basic data preprocessing in accordance with what I've seen in some papers/blogs/Kaggle kernels:</p>\n<ol>\n<li>Load the data</li>\n<li>Remove whitespace with seam carving (resize at this stage?)</li>\n<li>Use Otsu's thresholding to make the image grayscale</li>\n<li>Tile the image</li>\n<li>Find patches with &gt; (50/90/etc.)% of tissue</li>\n</ol>\n<p>When looking at the data, they seem really quite blurry. I feel like the standard 512 pixel image for this type of data is really not enough. My feeling right now is that I need to get the data preprocessing right, and then I can just start iterating through different models.</p>",
          "rawMarkdown": "May you please elaborate on what you mean by proper pre-processing/augmentation? I've been focusing on this data engineering/preprocessing side of things, but I'm bumping into problems. I've been thinking about something along the lines of the following for some basic data preprocessing in accordance with what I've seen in some papers/blogs/Kaggle kernels:\n\n1. Load the data\n2. Remove whitespace with seam carving (resize at this stage?)\n3. Use Otsu's thresholding to make the image grayscale\n4. Tile the image\n5. Find patches with > (50/90/etc.)% of tissue\n\nWhen looking at the data, they seem really quite blurry. I feel like the standard 512 pixel image for this type of data is really not enough. My feeling right now is that I need to get the data preprocessing right, and then I can just start iterating through different models.",
          "votes": 1
        },
        {
          "id": 1905060,
          "postDate": "2022-08-18T17:58:48.870Z",
          "content": "<p>I have also looked at similar approaches to your suggestions, and they also did not work out for me.</p>\n<p>Notice that the data set is strongly imbalanced. A lot of 'advanced' CNNs converge to full classification of the majority class, if you pay attention. I suggest you look into a WeightedRandomClassifier or something similar.</p>",
          "rawMarkdown": "I have also looked at similar approaches to your suggestions, and they also did not work out for me.\n\nNotice that the data set is strongly imbalanced. A lot of 'advanced' CNNs converge to full classification of the majority class, if you pay attention. I suggest you look into a WeightedRandomClassifier or something similar.",
          "votes": 1
        },
        {
          "id": 1905079,
          "postDate": "2022-08-18T18:12:11.073Z",
          "content": "<p>Thank you for the suggestion! May I ask if you have tried tiling to larger images? I noticed people in this comp downsize the images to 512x512 for example, but I've seen that in papers people will have tiles themselves that are 1024x1024. Is it just that the computational constraints make it not possible for these more ideal images?</p>",
          "rawMarkdown": "Thank you for the suggestion! May I ask if you have tried tiling to larger images? I noticed people in this comp downsize the images to 512x512 for example, but I've seen that in papers people will have tiles themselves that are 1024x1024. Is it just that the computational constraints make it not possible for these more ideal images?",
          "votes": 2
        },
        {
          "id": 1905123,
          "postDate": "2022-08-18T18:40:38.097Z",
          "content": "<p>Upscaling to 1024x1024 generally also increases parameters required in the CNN; this is obviously not feasible due to the limit in data size.</p>",
          "rawMarkdown": "Upscaling to 1024x1024 generally also increases parameters required in the CNN; this is obviously not feasible due to the limit in data size.",
          "votes": 2
        },
        {
          "id": 1905128,
          "postDate": "2022-08-18T18:42:10.833Z",
          "content": "<p>I see, thank you again for your insight!</p>",
          "rawMarkdown": "I see, thank you again for your insight!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1900475,
      "postDate": "2022-08-16T04:15:49.060Z",
      "content": "<p>Yes, I really believe that it's not possible to make a accurate model with this data, I even tried a small version of HIPT (<a href=\"https://github.com/mahmoodlab/HIPT\" target=\"_blank\">https://github.com/mahmoodlab/HIPT</a>) and still got a flip of a coin.</p>",
      "rawMarkdown": "Yes, I really believe that it's not possible to make a accurate model with this data, I even tried a small version of HIPT (https://github.com/mahmoodlab/HIPT) and still got a flip of a coin.",
      "votes": 2,
      "replies": [
        {
          "id": 1901573,
          "postDate": "2022-08-16T18:59:10.017Z",
          "content": "<p>Thanks for your input, interesting to hear that your implementation of that model was unsuccessful. It seems like a challenging architecture to build. Also, I would guess that the computational constraints under Kaggle would make such a model even more difficult to develop.</p>",
          "rawMarkdown": "Thanks for your input, interesting to hear that your implementation of that model was unsuccessful. It seems like a challenging architecture to build. Also, I would guess that the computational constraints under Kaggle would make such a model even more difficult to develop.",
          "votes": 1
        },
        {
          "id": 1903178,
          "postDate": "2022-08-17T07:02:28.530Z",
          "content": "<p>I built it on my own pc (poor thing). The difference is that I made it a 2 layer one instead of 3. If you understand the code and how it works it's actually not that hard to reverse engineer.</p>\n<p>I did not use their method of training tho, but I doubt it would make much difference.</p>",
          "rawMarkdown": "I built it on my own pc (poor thing). The difference is that I made it a 2 layer one instead of 3. If you understand the code and how it works it's actually not that hard to reverse engineer.\n\nI did not use their method of training tho, but I doubt it would make much difference.",
          "votes": 1
        },
        {
          "id": 1904268,
          "postDate": "2022-08-18T05:10:26.990Z",
          "content": "<p>I have tried CLAM with the pretrained HIPT model as feature extractor<br>\nMy model overfit with strip_ai data, have you found this issue?</p>",
          "rawMarkdown": "I have tried CLAM with the pretrained HIPT model as feature extractor\nMy model overfit with strip_ai data, have you found this issue?",
          "votes": 1
        },
        {
          "id": 1904287,
          "postDate": "2022-08-18T05:22:22.007Z",
          "content": "<p><a href=\"https://www.kaggle.com/ddt2018\" target=\"_blank\">@ddt2018</a> may you please clarify what you mean by overfitting? How did you come to this conclusion?</p>",
          "rawMarkdown": "@ddt2018 may you please clarify what you mean by overfitting? How did you come to this conclusion?"
        },
        {
          "id": 1904405,
          "postDate": "2022-08-18T06:52:53.697Z",
          "content": "<p>I split the origin training data into train/val/test set, <br>\non the training set, i got high acc, low loss, and the train loss decrease during the training process.<br>\non the val set, I got low auc, high logloss and the val loss decrease at first, then increase during the training process.</p>",
          "rawMarkdown": "I split the origin training data into train/val/test set, \non the training set, i got high acc, low loss, and the train loss decrease during the training process.\non the val set, I got low auc, high logloss and the val loss decrease at first, then increase during the training process.",
          "votes": 2
        },
        {
          "id": 1905442,
          "postDate": "2022-08-19T04:35:25.933Z",
          "content": "<p>I see, thank you for the explanation.</p>",
          "rawMarkdown": "I see, thank you for the explanation."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1907091,
      "author_name": "Satya",
      "author_url": "",
      "post_date": "2022-08-20T12:46:57.577000",
      "content": "<p>Improve Vision Transformers Training by Suppressing Over-smoothing, May be this research paper would help in solving the issue. when we train directly it yields unstable results to take care of this some modification is required in model by incorporating some conv layer …\" <a href=\"https://www.arxiv-vanity.com/papers/2104.12753/\" target=\"_blank\">https://www.arxiv-vanity.com/papers/2104.12753/</a> \". may be it could help :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1907547,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-20T21:57:47.230000",
          "content": "<p>Awesome, thanks for the idea. I will look into this.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1906423,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-08-19T21:37:03.943000",
      "content": "<p>I get the same feeling from nearly all models.. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1906427,
          "author_name": "jcerpent",
          "author_url": "",
          "post_date": "2022-08-19T21:48:26.677000",
          "content": "<p>Have you tried weighted random sampling? </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1906452,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-19T22:36:03.850000",
          "content": "<p><a href=\"https://www.kaggle.com/jcerpentier\" target=\"_blank\">@jcerpentier</a> I did it initially with a resnet18 model, which so far has my best AUC score. The resulting score was a bit lower surprisingly. I had the weighted random sampler as something on my to-do list, but hadn't yet got around to working on until you mentioned it. I plan on trying it later with some more models though.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1907667,
          "author_name": "hjunlee941",
          "author_url": "",
          "post_date": "2022-08-21T00:50:07.387000",
          "content": "<p><a href=\"https://www.kaggle.com/yuqizheng\" target=\"_blank\">@yuqizheng</a> Did you still tile the original images when using this approach?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1907700,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-21T01:58:28.093000",
          "content": "<p><a href=\"https://www.kaggle.com/hjunlee941\" target=\"_blank\">@hjunlee941</a> No, do you have a source for a good approach that utilizes tiling? I'm trying to work on a good pipeline right now for tiles.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1907709,
          "author_name": "hjunlee941",
          "author_url": "",
          "post_date": "2022-08-21T02:19:32.490000",
          "content": "<p><a href=\"https://www.kaggle.com/yuqizheng\" target=\"_blank\">@yuqizheng</a> I literally scratch built the pipeline to use patches. However, I still cannot find any good way to only remove background tiles.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1907721,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-21T02:50:35.807000",
          "content": "<p><a href=\"https://www.kaggle.com/hjunlee941\" target=\"_blank\">@hjunlee941</a> For background removal, a method that I'm using is as follows. Say you're reading patches of a WSI to save memory on your notebook (e.g., OpenSlide's read_region function). So for a given patch you read it of say size 512x512. You can first convert the RGB image to grayscale using Otsu's threshold (there are other choices for thresholding). From the grayscale image, you can count the ratio of black pixels to white pixels. If for example there are at least 50% or 90% of white pixels (indicating tissue regions), then you keep the tile (the untouched version that hasn't been given grayscale). I also noticed this can keep random regions with mostly whitespace, so you would probably also want to add another condition that there's at least 5% of black pixels in the image also.</p>\n<p>I think this method is quite different from what I've seen in many notebooks. I think many people are giving examples of first resizing the WSI image to something like 1024x1024, and then taking patches out of it. This however gives poor resolution patches that lack much information (based on just using my eye). I think the ideal method is to read the region, check if it has enough tissue area, and then resize the patch to something like 512x512.</p>\n<p>There are some other details too, like applying a sliding window with some multiplier (say 5) applied to the 512x512 patch window. This way the sliding window going through the 50k pixel WSI is 2560x2560, which in turn is the patch that is resized to 512x512.</p>\n<p>Sorry if this isn't super clear, I was thinking about sharing a notebook later once I got over the hurdle of ensuring that each image has at least X (e.g., 64) number of patches. When I did it without considering this, I had a minimum of 1 patch and maximum of ~500 patches for a given image.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1908488,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-08-21T18:08:51.880000",
          "content": "<p>I found <a href=\"https://www.kaggle.com/code/yasufuminakama/mayo-train-images-size-1024-n-16-1\" target=\"_blank\">this</a> notebook's tile creation very effective if you are looking at a top N tiles/patches sorted by signal in tile. This method was used effectively in previous competitions as well.</p>\n<p>Along with this I found that using unique value count threshold helped to remove less signal or mostly background tiles in a very simple manner. For e.g.,</p>\n<p>len(np.unique(tile)) &gt; 180 -- worked reasonably well, given tile is (s,s,3) shape. The threshold will probably depend on downscale factor and tile size. I guess this is ok if you are selecting top N tiles, but may be more challenging if you want to process all tiles from the image.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1909681,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-22T20:43:12.307000",
          "content": "<p>Something I've been eager to duplicate is the patching method found in this <a href=\"https://arxiv.org/pdf/1612.07180.pdf\" target=\"_blank\">paper</a>. I've included the image below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182873%2F03207d78b7078cfc5f335abfc03fd727%2FScreen%20Shot%202022-08-22%20at%201.37.06%20PM.png?generation=1661200689201901&amp;alt=media\" alt=\"\"><br>\nThey use a tool called CellProfiler, which is available on GitHub. I was wondering though if there's an ad-hoc way to implement it. Basically, for their breast cancer dataset, it makes sense for them to find the \"densest\" part of the image in terms of the tissue, since this is related to the cancer cells in breast cancer (if I read it correctly). I think if we can do something similar, by basically sampling from around the non-background regions and getting a set of e.g., N=64 patches then it would help us greatly.</p>\n<p>I tried a method to first read patches from the image, use Otsu's threshold to grayscale them. Then, I can determine which areas are tissue versus background. I tried to then randomly sample from these points, and filter the sampled points by first checking if there's at least e.g., 50% of white pixels in the image. The code works, but it's far too slow despite hours trying to vectorize and optimize the speed of the operations. Currently, it seems that even with the notebook you shared, the tiling can still at times have issues, giving patches that give little to no useful information.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1909729,
          "author_name": "hjunlee941",
          "author_url": "",
          "post_date": "2022-08-22T21:32:10.490000",
          "content": "<p><a href=\"https://www.kaggle.com/yuqizheng\" target=\"_blank\">@yuqizheng</a> Does CellProfiler also work for non-white background? Since we have some very weird backgrounds in the training data, simple threshold based method may not work ideally.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1909760,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-08-22T22:44:01.903000",
          "content": "<p>The base paper for CellProfiler is based an object identifier module.  It will be interesting to see if it can navigate a lot of noise in mayo dataset that could be mistaken for cell shaped features on the fringe of the scan area. Worth a try perhaps - the image above is very impressive.</p>\n<p><a href=\"https://www.kaggle.com/yuqizheng\" target=\"_blank\">@yuqizheng</a> I agree, normal tile creation is very frustratingly slow and not so discriminating amongst tiles. This surely needs an upgrade if downstream classification needs to go up a level. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1897204,
      "author_name": "Nikita Glazunov",
      "author_url": "",
      "post_date": "2022-08-13T14:43:10.790000",
      "content": "<p>I guess that the score is low cause the data is imbalanced. AUC is bad at imbalanced classes.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1897462,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-13T18:35:48.503000",
          "content": "<p>I actually chose AUC based on the following quote from <a href=\"https://www.kaggle.com/abhishek\" target=\"_blank\">@abhishek</a>'s book, Approaching Almost Any Machine Learning Problem,</p>\n<blockquote>\n  <p>As a data doctor would say: this is a classic case of skewed binary classification. Therefore, we choose the evaluation metric to be AUC and go for a stratified k-fold cross-validation scheme.</p>\n</blockquote>\n<p>I'm not sure if you're aware, but the idea (or the one that I had in mind) I think is that since AUC is sensitive to imbalanced data in comparison to something like accuracy. If for example the data is 90% class A and 10% class B then just predicting class A all the time would give a score of 90%. However, since AUC factors in other things like sensitivity and false positives, it makes it more ideal for measuring the ability of a model to perform on imbalanced data.</p>\n<p>If you already knew what I said in the above paragraph you can go ahead and ignore it, didn't mean to lecture you on AUC.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1897983,
          "author_name": "Nikita Glazunov",
          "author_url": "",
          "post_date": "2022-08-14T07:11:39.603000",
          "content": "<p>What about applying F1-score?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1898638,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-14T17:34:31.010000",
          "content": "<p>I haven't tried that method yet, but I'm just starting to get into Kaggle comps. I think it would make sense to create an array of metrics that can be tested epoch-wise and step-wise for something like PyTorch. That way we could gauge, using multiple methods, the effectiveness of a model on a dataset. This could include things like F1, AUC, accuracy, etc. Also, printing all the misclassifications would allow us to examine all the failures (and correct guesses) of the model to try and see why the model is performing well or not.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1904746,
      "author_name": "jcerpent",
      "author_url": "",
      "post_date": "2022-08-18T12:56:15.103000",
      "content": "<p>From my experience so far, the data set is extremely difficult to learn. But the proper pre-processing/augmentation helps; I think there is definitely not a \"coin flip problem\" for this competition, as some might have suggested.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1905029,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-18T17:36:30.160000",
          "content": "<p>May you please elaborate on what you mean by proper pre-processing/augmentation? I've been focusing on this data engineering/preprocessing side of things, but I'm bumping into problems. I've been thinking about something along the lines of the following for some basic data preprocessing in accordance with what I've seen in some papers/blogs/Kaggle kernels:</p>\n<ol>\n<li>Load the data</li>\n<li>Remove whitespace with seam carving (resize at this stage?)</li>\n<li>Use Otsu's thresholding to make the image grayscale</li>\n<li>Tile the image</li>\n<li>Find patches with &gt; (50/90/etc.)% of tissue</li>\n</ol>\n<p>When looking at the data, they seem really quite blurry. I feel like the standard 512 pixel image for this type of data is really not enough. My feeling right now is that I need to get the data preprocessing right, and then I can just start iterating through different models.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1905060,
          "author_name": "jcerpent",
          "author_url": "",
          "post_date": "2022-08-18T17:58:48.870000",
          "content": "<p>I have also looked at similar approaches to your suggestions, and they also did not work out for me.</p>\n<p>Notice that the data set is strongly imbalanced. A lot of 'advanced' CNNs converge to full classification of the majority class, if you pay attention. I suggest you look into a WeightedRandomClassifier or something similar.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1905079,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-18T18:12:11.073000",
          "content": "<p>Thank you for the suggestion! May I ask if you have tried tiling to larger images? I noticed people in this comp downsize the images to 512x512 for example, but I've seen that in papers people will have tiles themselves that are 1024x1024. Is it just that the computational constraints make it not possible for these more ideal images?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1905123,
          "author_name": "jcerpent",
          "author_url": "",
          "post_date": "2022-08-18T18:40:38.097000",
          "content": "<p>Upscaling to 1024x1024 generally also increases parameters required in the CNN; this is obviously not feasible due to the limit in data size.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1905128,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-18T18:42:10.833000",
          "content": "<p>I see, thank you again for your insight!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1900475,
      "author_name": "Iuryck Santos",
      "author_url": "",
      "post_date": "2022-08-16T04:15:49.060000",
      "content": "<p>Yes, I really believe that it's not possible to make a accurate model with this data, I even tried a small version of HIPT (<a href=\"https://github.com/mahmoodlab/HIPT\" target=\"_blank\">https://github.com/mahmoodlab/HIPT</a>) and still got a flip of a coin.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1901573,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-16T18:59:10.017000",
          "content": "<p>Thanks for your input, interesting to hear that your implementation of that model was unsuccessful. It seems like a challenging architecture to build. Also, I would guess that the computational constraints under Kaggle would make such a model even more difficult to develop.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1903178,
          "author_name": "Iuryck Santos",
          "author_url": "",
          "post_date": "2022-08-17T07:02:28.530000",
          "content": "<p>I built it on my own pc (poor thing). The difference is that I made it a 2 layer one instead of 3. If you understand the code and how it works it's actually not that hard to reverse engineer.</p>\n<p>I did not use their method of training tho, but I doubt it would make much difference.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1904268,
          "author_name": "ddt0106",
          "author_url": "",
          "post_date": "2022-08-18T05:10:26.990000",
          "content": "<p>I have tried CLAM with the pretrained HIPT model as feature extractor<br>\nMy model overfit with strip_ai data, have you found this issue?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1904287,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-18T05:22:22.007000",
          "content": "<p><a href=\"https://www.kaggle.com/ddt2018\" target=\"_blank\">@ddt2018</a> may you please clarify what you mean by overfitting? How did you come to this conclusion?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1904405,
          "author_name": "ddt0106",
          "author_url": "",
          "post_date": "2022-08-18T06:52:53.697000",
          "content": "<p>I split the origin training data into train/val/test set, <br>\non the training set, i got high acc, low loss, and the train loss decrease during the training process.<br>\non the val set, I got low auc, high logloss and the val loss decrease at first, then increase during the training process.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1905442,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-19T04:35:25.933000",
          "content": "<p>I see, thank you for the explanation.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1895291": "Hi all, I've tried a few basic models like random forest (RF), and general CNN's like alexnet and resnet. I tried the ViT base model via HuggingFace, and the AUC score is 0.51/0.52 after 30 epochs, which is quite a bit lower than even RF with a 5-fold AUC of 0.5662. Is anyone else facing similar issues? It seems to me that the ViT can't learn the problem at all, and is similar to a random guess of 0.5.",
    "1907091": "Improve Vision Transformers Training by Suppressing Over-smoothing, May be this research paper would help in solving the issue. when we train directly it yields unstable results to take care of this some modification is required in model by incorporating some conv layer ...\" https://www.arxiv-vanity.com/papers/2104.12753/ \". may be it could help :)",
    "1906423": "I get the same feeling from nearly all models.. \n\n\n",
    "1897204": "I guess that the score is low cause the data is imbalanced. AUC is bad at imbalanced classes.",
    "1904746": "From my experience so far, the data set is extremely difficult to learn. But the proper pre-processing/augmentation helps; I think there is definitely not a \"coin flip problem\" for this competition, as some might have suggested.",
    "1900475": "Yes, I really believe that it's not possible to make a accurate model with this data, I even tried a small version of HIPT (https://github.com/mahmoodlab/HIPT) and still got a flip of a coin."
  }
}