{
  "id": 449209,
  "title": "Validity of image level labels for patch level analysis",
  "url": "/competitions/UBC-OCEAN/discussion/449209",
  "author_name": "KarthiAru",
  "post_date": "2023-10-23T14:17:20.576000",
  "votes": 12,
  "comment_count": 25,
  "views": 0,
  "content": "<p>If I segment the larger WSI images into smaller patches of dimensions 256, 512, or 1024, will the overall image labels remain accurate? I'm looking to confirm that most of the patches aren't merely normal tissue.<br>\nAlso, is it a fair assumption to make that the cancer's grade and stage are consistent?</p>",
  "messages": [
    {
      "id": 2493665,
      "postDate": "2023-10-23T14:17:20.577Z",
      "content": "<p>If I segment the larger WSI images into smaller patches of dimensions 256, 512, or 1024, will the overall image labels remain accurate? I'm looking to confirm that most of the patches aren't merely normal tissue.<br>\nAlso, is it a fair assumption to make that the cancer's grade and stage are consistent?</p>",
      "rawMarkdown": "If I segment the larger WSI images into smaller patches of dimensions 256, 512, or 1024, will the overall image labels remain accurate? I'm looking to confirm that most of the patches aren't merely normal tissue.\nAlso, is it a fair assumption to make that the cancer's grade and stage are consistent?",
      "votes": 12
    },
    {
      "id": 2496582,
      "postDate": "2023-10-24T06:32:15.437Z",
      "content": "<p>I think that for positive WSI, there must be some negative patches. For pathological image classification tasks, multi-instance learning (MIL) is usually used to solve the problem. MIL is used to solve the problem that the positive bag contains at least one positive instance, and all the instances in the negative bag are negative instances.</p>",
      "rawMarkdown": "I think that for positive WSI, there must be some negative patches. For pathological image classification tasks, multi-instance learning (MIL) is usually used to solve the problem. MIL is used to solve the problem that the positive bag contains at least one positive instance, and all the instances in the negative bag are negative instances.",
      "votes": 6,
      "replies": [
        {
          "id": 2505966,
          "postDate": "2023-10-31T02:39:43.020Z",
          "content": "<p>yes, this is true, I am looking into it already… :)</p>",
          "rawMarkdown": "yes, this is true, I am looking into it already... :)"
        },
        {
          "id": 2555475,
          "postDate": "2023-12-10T00:09:53.383Z",
          "content": "<p>Hi, have you employed weakly supervised multiple instance learning? In this environment with memory and inference time constraints, my multiple instance learning has not achieved very good results.</p>",
          "rawMarkdown": "Hi, have you employed weakly supervised multiple instance learning? In this environment with memory and inference time constraints, my multiple instance learning has not achieved very good results.",
          "replies": [
            {
              "id": 2555560,
              "postDate": "2023-12-10T03:16:21.130Z",
              "content": "<p>Yes, I currently use simple backbone and mil. To alleviate memory and inference time constraints, you could try using smaller magnifications (I currently use 224 size patches under x10). My current submission time is about 5 hours.</p>",
              "rawMarkdown": "Yes, I currently use simple backbone and mil. To alleviate memory and inference time constraints, you could try using smaller magnifications (I currently use 224 size patches under x10). My current submission time is about 5 hours.",
              "votes": 2
            },
            {
              "id": 2555611,
              "postDate": "2023-12-10T04:10:39.153Z",
              "content": "<p>Thank you. Under a 10x magnification level for the WSI, there's no need to limit the maximum sampling number, right? Also, did you convert the TMA from a 40x magnification level to a 10x magnification level correspondingly?</p>",
              "rawMarkdown": "Thank you. Under a 10x magnification level for the WSI, there's no need to limit the maximum sampling number, right? Also, did you convert the TMA from a 40x magnification level to a 10x magnification level correspondingly?"
            },
            {
              "id": 2555748,
              "postDate": "2023-12-10T06:38:58.130Z",
              "content": "<p>For inference, using backbone to extract features can be done in chunks, so a large number of patches should not be a big problem. For TMA I downsampled 4x, while WSI only downsampled 2x, in order to make their resolutions the same.</p>",
              "rawMarkdown": "For inference, using backbone to extract features can be done in chunks, so a large number of patches should not be a big problem. For TMA I downsampled 4x, while WSI only downsampled 2x, in order to make their resolutions the same.",
              "votes": 4
            },
            {
              "id": 2555772,
              "postDate": "2023-12-10T07:18:33.390Z",
              "content": "<p>That is good point, but do we know is an image is tma for testing? </p>",
              "rawMarkdown": "That is good point, but do we know is an image is tma for testing? "
            },
            {
              "id": 2555774,
              "postDate": "2023-12-10T07:19:05.600Z",
              "content": "<p>Yes, I understand what you mean.  </p>\n<p>My question is that I also use a backbone (such as ResNet) to extract features from all patches in a WSI (whole slide image) and save them in a temporary storage in .pt format.  However, the process of cutting patches consumes a lot of time, while feature extraction is much faster.  </p>\n<p>To avoid timeouts, I set a maximum sampling number for the patch-cutting function, but this means that the extracted features are not representative of the entire WSI, which isn't ideal for the inference performance of the model I'm training…  </p>\n<p>For example, in the case of a single WSI, do you manage to cut all patches and extract their features in just 5 hours?</p>",
              "rawMarkdown": "Yes, I understand what you mean.  \n\nMy question is that I also use a backbone (such as ResNet) to extract features from all patches in a WSI (whole slide image) and save them in a temporary storage in .pt format.  However, the process of cutting patches consumes a lot of time, while feature extraction is much faster.  \n\nTo avoid timeouts, I set a maximum sampling number for the patch-cutting function, but this means that the extracted features are not representative of the entire WSI, which isn't ideal for the inference performance of the model I'm training...  \n\nFor example, in the case of a single WSI, do you manage to cut all patches and extract their features in just 5 hours?",
              "votes": 1
            },
            {
              "id": 2555775,
              "postDate": "2023-12-10T07:27:28.810Z",
              "content": "<p>Yes the same, I had to set limit 50 patches per WSI </p>",
              "rawMarkdown": "Yes the same, I had to set limit 50 patches per WSI ",
              "votes": 1
            },
            {
              "id": 2555781,
              "postDate": "2023-12-10T07:35:36.473Z",
              "content": "<p>I use <code>is_tma = image.shape[0] &lt; 5000 and image.shape[1] &lt; 5000</code> to check it. There may be problems with this method, but I don't have a better solution at the moment.</p>",
              "rawMarkdown": "I use `is_tma = image.shape[0] < 5000 and image.shape[1] < 5000` to check it. There may be problems with this method, but I don't have a better solution at the moment.",
              "votes": 4
            },
            {
              "id": 2555782,
              "postDate": "2023-12-10T07:39:49.790Z",
              "content": "<p>Thanks,I also determine TMA and WSI based on their size. My question above was about whether you use all patches of each WSI during inference.</p>",
              "rawMarkdown": "Thanks,I also determine TMA and WSI based on their size. My question above was about whether you use all patches of each WSI during inference.",
              "votes": 1
            },
            {
              "id": 2555785,
              "postDate": "2023-12-10T07:46:22.360Z",
              "content": "<p>When I read the entire wsi, I will first downsample it by 2x, then use <code>cv2.findContours</code>, etc. to remove duplicate slices, and then use a sliding window to determine whether each patch needs to be retained. I think the main time-consuming part of this entire pipeline is reading wsi from the hard disk to the memory, so you can try to use <code>num_workers &gt; 1</code> of <code>DataLoader</code> to read in multiple processes to speed up io. My total submission time should be a little more than 6 hours now.</p>",
              "rawMarkdown": "When I read the entire wsi, I will first downsample it by 2x, then use `cv2.findContours`, etc. to remove duplicate slices, and then use a sliding window to determine whether each patch needs to be retained. I think the main time-consuming part of this entire pipeline is reading wsi from the hard disk to the memory, so you can try to use `num_workers > 1` of `DataLoader` to read in multiple processes to speed up io. My total submission time should be a little more than 6 hours now.",
              "votes": 2
            },
            {
              "id": 2555787,
              "postDate": "2023-12-10T07:47:50.537Z",
              "content": "<p>Currently, only patches with tissue areas greater than 0.25 are retained for inference.</p>",
              "rawMarkdown": "Currently, only patches with tissue areas greater than 0.25 are retained for inference.",
              "votes": 2
            },
            {
              "id": 2555788,
              "postDate": "2023-12-10T07:48:10.737Z",
              "content": "<p>Thank you very much! I will try it out right now.</p>",
              "rawMarkdown": "Thank you very much! I will try it out right now."
            },
            {
              "id": 2555955,
              "postDate": "2023-12-10T10:39:44.020Z",
              "content": "<p>Do you run the contour finder on patch or thumbnail? </p>",
              "rawMarkdown": "Do you run the contour finder on patch or thumbnail? "
            },
            {
              "id": 2555964,
              "postDate": "2023-12-10T10:46:12.103Z",
              "content": "<p>My understanding is to find the foreground contours on a WSI that has been scaled down by 2x, filter out the background, obtain the minimum bounding rectangle of these contours, and then use a sliding window within this rectangle to cut patches.</p>",
              "rawMarkdown": "My understanding is to find the foreground contours on a WSI that has been scaled down by 2x, filter out the background, obtain the minimum bounding rectangle of these contours, and then use a sliding window within this rectangle to cut patches.",
              "votes": 1
            },
            {
              "id": 2558722,
              "postDate": "2023-12-12T10:03:06.990Z",
              "content": "<p>Hello, I cut the picture to patches according to your way, but it still ran out of time. May I ask how you finished cutting the patch within 5 hours? Do you have any tips? What is the key to saving time? Look forward to your reply</p>",
              "rawMarkdown": "Hello, I cut the picture to patches according to your way, but it still ran out of time. May I ask how you finished cutting the patch within 5 hours? Do you have any tips? What is the key to saving time? Look forward to your reply",
              "votes": 1
            },
            {
              "id": 2558913,
              "postDate": "2023-12-12T12:46:07.323Z",
              "content": "<p>My total pipeline is above. I think <code>num_workers==2</code> should be the key to saving time.</p>",
              "rawMarkdown": "My total pipeline is above. I think `num_workers==2` should be the key to saving time."
            }
          ]
        },
        {
          "id": 2560245,
          "postDate": "2023-12-13T11:51:05.557Z",
          "content": "<p>Hi, what's your best local CV on single folder? I got huge gap between local CV and LB, CV is about 0.9, and LB is about 0.46. I'm curious if you've considered fine-tuning the backbone to achieve better results.</p>",
          "rawMarkdown": "Hi, what's your best local CV on single folder? I got huge gap between local CV and LB, CV is about 0.9, and LB is about 0.46. I'm curious if you've considered fine-tuning the backbone to achieve better results.",
          "votes": 1,
          "replies": [
            {
              "id": 2560334,
              "postDate": "2023-12-13T13:46:15.040Z",
              "content": "<p>Single-fold results may be unreliable. My current 5-fold CV result is around 0.854 and the highest single fold result is 0.918. I have tried end-to-end training to fine-tune the backbone before, but due to insufficient gpu memory, it has not been successful yet.</p>",
              "rawMarkdown": "Single-fold results may be unreliable. My current 5-fold CV result is around 0.854 and the highest single fold result is 0.918. I have tried end-to-end training to fine-tune the backbone before, but due to insufficient gpu memory, it has not been successful yet.",
              "votes": 2
            },
            {
              "id": 2560410,
              "postDate": "2023-12-13T15:44:20Z",
              "content": "<p>That's weird. My OOF score is 85.69 (TMA: 88, WSI: 85.31) and LB score is 0.46 too.</p>",
              "rawMarkdown": "That's weird. My OOF score is 85.69 (TMA: 88, WSI: 85.31) and LB score is 0.46 too.",
              "votes": 1
            },
            {
              "id": 2564413,
              "postDate": "2023-12-17T05:08:24.327Z",
              "content": "<blockquote>\n  <p>That's weird. My OOF score is 85.69 (TMA: 88, WSI: 85.31) and LB score is 0.46 too.</p>\n</blockquote>\n<p>The same phenomenon here. 5-fold CV around 0.9 but the LB score is 0.49.</p>",
              "rawMarkdown": "> That's weird. My OOF score is 85.69 (TMA: 88, WSI: 85.31) and LB score is 0.46 too.\n\nThe same phenomenon here. 5-fold CV around 0.9 but the LB score is 0.49."
            }
          ]
        }
      ]
    },
    {
      "id": 2496366,
      "postDate": "2023-10-24T02:26:24.567Z",
      "content": "<p>A more effective approach, akin to Google Recaptcha's image verification grid, might involve annotating WSI images at the patch level, thereby inherently adjusting for variations in the cancer's grade and stage.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F973418%2Fb09dffbb39c678affc47f86e2580c3b8%2FFireShot%20Capture%20001%20-%20Google%20reCAPTCHA%20-%20developers.nopecha.com.png?generation=1698114180923084&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "A more effective approach, akin to Google Recaptcha's image verification grid, might involve annotating WSI images at the patch level, thereby inherently adjusting for variations in the cancer's grade and stage.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F973418%2Fb09dffbb39c678affc47f86e2580c3b8%2FFireShot%20Capture%20001%20-%20Google%20reCAPTCHA%20-%20developers.nopecha.com.png?generation=1698114180923084&alt=media)",
      "votes": 4
    },
    {
      "id": 2494039,
      "postDate": "2023-10-23T17:27:16.630Z",
      "content": "<p>I'm not sure but I think the answer is, it depends. A domain expert should verify whether the patch has characteristics of its class.</p>",
      "rawMarkdown": "I'm not sure but I think the answer is, it depends. A domain expert should verify whether the patch has characteristics of its class.",
      "votes": 2
    },
    {
      "id": 2499407,
      "postDate": "2023-10-26T02:55:01Z",
      "content": "<p><a href=\"https://www.kaggle.com/dhinkris/crop-duplicate-wsi-images\" target=\"_blank\">https://www.kaggle.com/dhinkris/crop-duplicate-wsi-images</a></p>",
      "rawMarkdown": "https://www.kaggle.com/dhinkris/crop-duplicate-wsi-images"
    }
  ],
  "comments": [
    {
      "id": 2496582,
      "author_name": "m1dsolo",
      "author_url": "",
      "post_date": "2023-10-24T06:32:15.437000",
      "content": "<p>I think that for positive WSI, there must be some negative patches. For pathological image classification tasks, multi-instance learning (MIL) is usually used to solve the problem. MIL is used to solve the problem that the positive bag contains at least one positive instance, and all the instances in the negative bag are negative instances.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2505966,
          "author_name": "Jirka",
          "author_url": "",
          "post_date": "2023-10-31T02:39:43.020000",
          "content": "<p>yes, this is true, I am looking into it already… :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2555475,
          "author_name": "Huang Jin Feng",
          "author_url": "",
          "post_date": "2023-12-10T00:09:53.383000",
          "content": "<p>Hi, have you employed weakly supervised multiple instance learning? In this environment with memory and inference time constraints, my multiple instance learning has not achieved very good results.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2555560,
              "author_name": "m1dsolo",
              "author_url": "",
              "post_date": "2023-12-10T03:16:21.130000",
              "content": "<p>Yes, I currently use simple backbone and mil. To alleviate memory and inference time constraints, you could try using smaller magnifications (I currently use 224 size patches under x10). My current submission time is about 5 hours.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2555611,
              "author_name": "Huang Jin Feng",
              "author_url": "",
              "post_date": "2023-12-10T04:10:39.153000",
              "content": "<p>Thank you. Under a 10x magnification level for the WSI, there's no need to limit the maximum sampling number, right? Also, did you convert the TMA from a 40x magnification level to a 10x magnification level correspondingly?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2555748,
              "author_name": "m1dsolo",
              "author_url": "",
              "post_date": "2023-12-10T06:38:58.130000",
              "content": "<p>For inference, using backbone to extract features can be done in chunks, so a large number of patches should not be a big problem. For TMA I downsampled 4x, while WSI only downsampled 2x, in order to make their resolutions the same.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2555772,
              "author_name": "Jirka",
              "author_url": "",
              "post_date": "2023-12-10T07:18:33.390000",
              "content": "<p>That is good point, but do we know is an image is tma for testing? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2555774,
              "author_name": "Huang Jin Feng",
              "author_url": "",
              "post_date": "2023-12-10T07:19:05.600000",
              "content": "<p>Yes, I understand what you mean.  </p>\n<p>My question is that I also use a backbone (such as ResNet) to extract features from all patches in a WSI (whole slide image) and save them in a temporary storage in .pt format.  However, the process of cutting patches consumes a lot of time, while feature extraction is much faster.  </p>\n<p>To avoid timeouts, I set a maximum sampling number for the patch-cutting function, but this means that the extracted features are not representative of the entire WSI, which isn't ideal for the inference performance of the model I'm training…  </p>\n<p>For example, in the case of a single WSI, do you manage to cut all patches and extract their features in just 5 hours?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2555775,
              "author_name": "Jirka",
              "author_url": "",
              "post_date": "2023-12-10T07:27:28.810000",
              "content": "<p>Yes the same, I had to set limit 50 patches per WSI </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2555781,
              "author_name": "m1dsolo",
              "author_url": "",
              "post_date": "2023-12-10T07:35:36.473000",
              "content": "<p>I use <code>is_tma = image.shape[0] &lt; 5000 and image.shape[1] &lt; 5000</code> to check it. There may be problems with this method, but I don't have a better solution at the moment.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2555782,
              "author_name": "Huang Jin Feng",
              "author_url": "",
              "post_date": "2023-12-10T07:39:49.790000",
              "content": "<p>Thanks,I also determine TMA and WSI based on their size. My question above was about whether you use all patches of each WSI during inference.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2555785,
              "author_name": "m1dsolo",
              "author_url": "",
              "post_date": "2023-12-10T07:46:22.360000",
              "content": "<p>When I read the entire wsi, I will first downsample it by 2x, then use <code>cv2.findContours</code>, etc. to remove duplicate slices, and then use a sliding window to determine whether each patch needs to be retained. I think the main time-consuming part of this entire pipeline is reading wsi from the hard disk to the memory, so you can try to use <code>num_workers &gt; 1</code> of <code>DataLoader</code> to read in multiple processes to speed up io. My total submission time should be a little more than 6 hours now.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2555787,
              "author_name": "m1dsolo",
              "author_url": "",
              "post_date": "2023-12-10T07:47:50.537000",
              "content": "<p>Currently, only patches with tissue areas greater than 0.25 are retained for inference.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2555788,
              "author_name": "Huang Jin Feng",
              "author_url": "",
              "post_date": "2023-12-10T07:48:10.737000",
              "content": "<p>Thank you very much! I will try it out right now.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2555955,
              "author_name": "Jirka",
              "author_url": "",
              "post_date": "2023-12-10T10:39:44.020000",
              "content": "<p>Do you run the contour finder on patch or thumbnail? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2555964,
              "author_name": "Huang Jin Feng",
              "author_url": "",
              "post_date": "2023-12-10T10:46:12.103000",
              "content": "<p>My understanding is to find the foreground contours on a WSI that has been scaled down by 2x, filter out the background, obtain the minimum bounding rectangle of these contours, and then use a sliding window within this rectangle to cut patches.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2558722,
              "author_name": "Lingxi",
              "author_url": "",
              "post_date": "2023-12-12T10:03:06.990000",
              "content": "<p>Hello, I cut the picture to patches according to your way, but it still ran out of time. May I ask how you finished cutting the patch within 5 hours? Do you have any tips? What is the key to saving time? Look forward to your reply</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2558913,
              "author_name": "m1dsolo",
              "author_url": "",
              "post_date": "2023-12-12T12:46:07.323000",
              "content": "<p>My total pipeline is above. I think <code>num_workers==2</code> should be the key to saving time.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2560245,
          "author_name": "OverWhelmingFIT",
          "author_url": "",
          "post_date": "2023-12-13T11:51:05.557000",
          "content": "<p>Hi, what's your best local CV on single folder? I got huge gap between local CV and LB, CV is about 0.9, and LB is about 0.46. I'm curious if you've considered fine-tuning the backbone to achieve better results.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2560334,
              "author_name": "m1dsolo",
              "author_url": "",
              "post_date": "2023-12-13T13:46:15.040000",
              "content": "<p>Single-fold results may be unreliable. My current 5-fold CV result is around 0.854 and the highest single fold result is 0.918. I have tried end-to-end training to fine-tune the backbone before, but due to insufficient gpu memory, it has not been successful yet.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2560410,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2023-12-13T15:44:20",
              "content": "<p>That's weird. My OOF score is 85.69 (TMA: 88, WSI: 85.31) and LB score is 0.46 too.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2564413,
              "author_name": "AnnieGo",
              "author_url": "",
              "post_date": "2023-12-17T05:08:24.327000",
              "content": "<blockquote>\n  <p>That's weird. My OOF score is 85.69 (TMA: 88, WSI: 85.31) and LB score is 0.46 too.</p>\n</blockquote>\n<p>The same phenomenon here. 5-fold CV around 0.9 but the LB score is 0.49.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2496366,
      "author_name": "KarthiAru",
      "author_url": "",
      "post_date": "2023-10-24T02:26:24.567000",
      "content": "<p>A more effective approach, akin to Google Recaptcha's image verification grid, might involve annotating WSI images at the patch level, thereby inherently adjusting for variations in the cancer's grade and stage.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F973418%2Fb09dffbb39c678affc47f86e2580c3b8%2FFireShot%20Capture%20001%20-%20Google%20reCAPTCHA%20-%20developers.nopecha.com.png?generation=1698114180923084&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2494039,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2023-10-23T17:27:16.630000",
      "content": "<p>I'm not sure but I think the answer is, it depends. A domain expert should verify whether the patch has characteristics of its class.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2499407,
      "author_name": "dhinesh",
      "author_url": "",
      "post_date": "2023-10-26T02:55:01",
      "content": "<p><a href=\"https://www.kaggle.com/dhinkris/crop-duplicate-wsi-images\" target=\"_blank\">https://www.kaggle.com/dhinkris/crop-duplicate-wsi-images</a></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2493665": "If I segment the larger WSI images into smaller patches of dimensions 256, 512, or 1024, will the overall image labels remain accurate? I'm looking to confirm that most of the patches aren't merely normal tissue.\nAlso, is it a fair assumption to make that the cancer's grade and stage are consistent?",
    "2496582": "I think that for positive WSI, there must be some negative patches. For pathological image classification tasks, multi-instance learning (MIL) is usually used to solve the problem. MIL is used to solve the problem that the positive bag contains at least one positive instance, and all the instances in the negative bag are negative instances.",
    "2496366": "A more effective approach, akin to Google Recaptcha's image verification grid, might involve annotating WSI images at the patch level, thereby inherently adjusting for variations in the cancer's grade and stage.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F973418%2Fb09dffbb39c678affc47f86e2580c3b8%2FFireShot%20Capture%20001%20-%20Google%20reCAPTCHA%20-%20developers.nopecha.com.png?generation=1698114180923084&alt=media)",
    "2494039": "I'm not sure but I think the answer is, it depends. A domain expert should verify whether the patch has characteristics of its class.",
    "2499407": "https://www.kaggle.com/dhinkris/crop-duplicate-wsi-images"
  }
}