{
  "id": 460890,
  "title": "How to deal with outliers?",
  "url": "/competitions/UBC-OCEAN/discussion/460890",
  "author_name": "zznznb",
  "post_date": "2023-12-11T16:41:49.079000",
  "votes": 19,
  "comment_count": 7,
  "views": 0,
  "content": "<p>My method is based on MIL and model ensemble. I tried to extract patch features on both the thumbnails and original images, and currently the best scores are 0.51 and 0.54, respectively. This came as a surprise to me, as using higher resolution images only resulted in a slight performance improvement. Because MIL needs to extract fixed features in advance, I did not use any image augmentation. I think I can only make improvements in outlier detection. I tried some methods, such as changing softmax to sigmoid and using threshold, and categorizing inconsistent predictions from multiple models as Other. However, they are not effective. So I want to ask how you deal with outliers? On the other hand, if you know how to make improvements on MIL, welcome your suggestions.</p>",
  "messages": [
    {
      "id": 2557742,
      "postDate": "2023-12-11T16:41:49.080Z",
      "content": "<p>My method is based on MIL and model ensemble. I tried to extract patch features on both the thumbnails and original images, and currently the best scores are 0.51 and 0.54, respectively. This came as a surprise to me, as using higher resolution images only resulted in a slight performance improvement. Because MIL needs to extract fixed features in advance, I did not use any image augmentation. I think I can only make improvements in outlier detection. I tried some methods, such as changing softmax to sigmoid and using threshold, and categorizing inconsistent predictions from multiple models as Other. However, they are not effective. So I want to ask how you deal with outliers? On the other hand, if you know how to make improvements on MIL, welcome your suggestions.</p>",
      "rawMarkdown": "My method is based on MIL and model ensemble. I tried to extract patch features on both the thumbnails and original images, and currently the best scores are 0.51 and 0.54, respectively. This came as a surprise to me, as using higher resolution images only resulted in a slight performance improvement. Because MIL needs to extract fixed features in advance, I did not use any image augmentation. I think I can only make improvements in outlier detection. I tried some methods, such as changing softmax to sigmoid and using threshold, and categorizing inconsistent predictions from multiple models as Other. However, they are not effective. So I want to ask how you deal with outliers? On the other hand, if you know how to make improvements on MIL, welcome your suggestions.",
      "votes": 19
    },
    {
      "id": 2558661,
      "postDate": "2023-12-12T09:05:24.920Z",
      "content": "<p>Hello, I would like to know if your method of feature extraction for thumbnails and WSI is the same.  Do you first segment the image into patches, extract features from each patch, and then concatenate these vectors to save them as a .pt file?</p>",
      "rawMarkdown": "Hello, I would like to know if your method of feature extraction for thumbnails and WSI is the same.  Do you first segment the image into patches, extract features from each patch, and then concatenate these vectors to save them as a .pt file?",
      "votes": 3,
      "replies": [
        {
          "id": 2558678,
          "postDate": "2023-12-12T09:27:54.103Z",
          "content": "<p>Yes, this is a commonly used method in academia to handle WSIs. A .pt file contains an N×D tensor, where N represents the number of patches and D represents the feature dimension.</p>",
          "rawMarkdown": "Yes, this is a commonly used method in academia to handle WSIs. A .pt file contains an N×D tensor, where N represents the number of patches and D represents the feature dimension.",
          "votes": 4,
          "replies": [
            {
              "id": 2558683,
              "postDate": "2023-12-12T09:33:56.343Z",
              "content": "<p>I'm quite curious, you mentioned achieving a 0.51 using the MIL method on thumbnails. However, if you extract features from these thumbnails, it seems impossible to simulate the magnification levels of TMA. During inference, it appears that you can only cut the WSI thumbnails and TMA into patches of the same scale and size.</p>",
              "rawMarkdown": "\nI'm quite curious, you mentioned achieving a 0.51 using the MIL method on thumbnails. However, if you extract features from these thumbnails, it seems impossible to simulate the magnification levels of TMA. During inference, it appears that you can only cut the WSI thumbnails and TMA into patches of the same scale and size.\n",
              "votes": 1
            },
            {
              "id": 2558760,
              "postDate": "2023-12-12T10:44:54.813Z",
              "content": "<p>You're right, I cropped both the WSI thumbnails and TMAs to 224 × 224. This may be a reason for the lower score. By the way, my several scores for this method are 0.48-0.50, and the score for a single model is approximately 0.43. I hope these can provide you with reference.</p>",
              "rawMarkdown": "You're right, I cropped both the WSI thumbnails and TMAs to 224 × 224. This may be a reason for the lower score. By the way, my several scores for this method are 0.48-0.50, and the score for a single model is approximately 0.43. I hope these can provide you with reference.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2557842,
      "postDate": "2023-12-11T17:42:22.690Z",
      "content": "<p>I try set threshold, arcface and get slightly boost. Does yours model ensemble increase score? My ensemble model score equal best single model score :((</p>",
      "rawMarkdown": "I try set threshold, arcface and get slightly boost. Does yours model ensemble increase score? My ensemble model score equal best single model score :((",
      "votes": 4,
      "replies": [
        {
          "id": 2558203,
          "postDate": "2023-12-12T01:40:15.967Z",
          "content": "<p>I think setting a threshold is a method which needs more luck, and a slight boost in public data does not necessarily indicate effectiveness on private data. That's exactly what I'm worried about. On the other hand, model ensemble needs to consider the performance differences of different models and try to remove models with poor performance. For example, I used vit_p16 and vit_p8. Compared to using vit_p16 and resnet50, the LB is 0.05 higher. But as I said, when the LB reaches 0.5, model ensemble will no longer bring significant improvement.</p>",
          "rawMarkdown": "I think setting a threshold is a method which needs more luck, and a slight boost in public data does not necessarily indicate effectiveness on private data. That's exactly what I'm worried about. On the other hand, model ensemble needs to consider the performance differences of different models and try to remove models with poor performance. For example, I used vit_p16 and vit_p8. Compared to using vit_p16 and resnet50, the LB is 0.05 higher. But as I said, when the LB reaches 0.5, model ensemble will no longer bring significant improvement.",
          "votes": 4,
          "replies": [
            {
              "id": 2558224,
              "postDate": "2023-12-12T02:09:51.380Z",
              "content": "<p>Yes. I think another approach: make improvements in TMA, I see my score on only TMA is low. My CV TMA is about 0.8, but  Public TMA is low. Recently, my CV increase but Public not change. Problem can be my cross validation split, training data missing info about source hospital. My CV can be data leak.</p>",
              "rawMarkdown": "Yes. I think another approach: make improvements in TMA, I see my score on only TMA is low. My CV TMA is about 0.8, but  Public TMA is low. Recently, my CV increase but Public not change. Problem can be my cross validation split, training data missing info about source hospital. My CV can be data leak.",
              "votes": 3
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2558661,
      "author_name": "Huang Jin Feng",
      "author_url": "",
      "post_date": "2023-12-12T09:05:24.920000",
      "content": "<p>Hello, I would like to know if your method of feature extraction for thumbnails and WSI is the same.  Do you first segment the image into patches, extract features from each patch, and then concatenate these vectors to save them as a .pt file?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2558678,
          "author_name": "zznznb",
          "author_url": "",
          "post_date": "2023-12-12T09:27:54.103000",
          "content": "<p>Yes, this is a commonly used method in academia to handle WSIs. A .pt file contains an N×D tensor, where N represents the number of patches and D represents the feature dimension.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2558683,
              "author_name": "Huang Jin Feng",
              "author_url": "",
              "post_date": "2023-12-12T09:33:56.343000",
              "content": "<p>I'm quite curious, you mentioned achieving a 0.51 using the MIL method on thumbnails. However, if you extract features from these thumbnails, it seems impossible to simulate the magnification levels of TMA. During inference, it appears that you can only cut the WSI thumbnails and TMA into patches of the same scale and size.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2558760,
              "author_name": "zznznb",
              "author_url": "",
              "post_date": "2023-12-12T10:44:54.813000",
              "content": "<p>You're right, I cropped both the WSI thumbnails and TMAs to 224 × 224. This may be a reason for the lower score. By the way, my several scores for this method are 0.48-0.50, and the score for a single model is approximately 0.43. I hope these can provide you with reference.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2557842,
      "author_name": "Quan Vu",
      "author_url": "",
      "post_date": "2023-12-11T17:42:22.690000",
      "content": "<p>I try set threshold, arcface and get slightly boost. Does yours model ensemble increase score? My ensemble model score equal best single model score :((</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2558203,
          "author_name": "zznznb",
          "author_url": "",
          "post_date": "2023-12-12T01:40:15.967000",
          "content": "<p>I think setting a threshold is a method which needs more luck, and a slight boost in public data does not necessarily indicate effectiveness on private data. That's exactly what I'm worried about. On the other hand, model ensemble needs to consider the performance differences of different models and try to remove models with poor performance. For example, I used vit_p16 and vit_p8. Compared to using vit_p16 and resnet50, the LB is 0.05 higher. But as I said, when the LB reaches 0.5, model ensemble will no longer bring significant improvement.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2558224,
              "author_name": "Quan Vu",
              "author_url": "",
              "post_date": "2023-12-12T02:09:51.380000",
              "content": "<p>Yes. I think another approach: make improvements in TMA, I see my score on only TMA is low. My CV TMA is about 0.8, but  Public TMA is low. Recently, my CV increase but Public not change. Problem can be my cross validation split, training data missing info about source hospital. My CV can be data leak.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2557742": "My method is based on MIL and model ensemble. I tried to extract patch features on both the thumbnails and original images, and currently the best scores are 0.51 and 0.54, respectively. This came as a surprise to me, as using higher resolution images only resulted in a slight performance improvement. Because MIL needs to extract fixed features in advance, I did not use any image augmentation. I think I can only make improvements in outlier detection. I tried some methods, such as changing softmax to sigmoid and using threshold, and categorizing inconsistent predictions from multiple models as Other. However, they are not effective. So I want to ask how you deal with outliers? On the other hand, if you know how to make improvements on MIL, welcome your suggestions.",
    "2558661": "Hello, I would like to know if your method of feature extraction for thumbnails and WSI is the same.  Do you first segment the image into patches, extract features from each patch, and then concatenate these vectors to save them as a .pt file?",
    "2557842": "I try set threshold, arcface and get slightly boost. Does yours model ensemble increase score? My ensemble model score equal best single model score :(("
  }
}