{
  "id": 453096,
  "title": "How to correctly convert from 40x to 20x magnification",
  "url": "/competitions/UBC-OCEAN/discussion/453096",
  "author_name": "Patchef",
  "post_date": "2023-11-04T20:37:53.209000",
  "votes": 9,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n<p>I have a quick question, or rather, I need someone to confirm the following for me because I'm not very familiar with the physics of microscope images…</p>\n<p>So, during training, we have images at 20x magnification, and during testing, we have them at 40x. I believe most people will likely crop patches from the original images and then resize those patches to the model's input size.</p>\n<p>To ensure that the objects in the 40x test images are the same size as in the 20x training images, one should resize the images. From what I currently understand, during the inference of the 40x images, you should first resize the image by a factor of 20/40=0.5 (i.e., shrink it), then crop the patches, and finally resize the patches to the model's input size again… Is that correct? I would greatly appreciate it if someone could confirm this or correct my thinking.</p>\n<p>PS: Of course, you could also resize the 20x images in training initially by a factor of 2.</p>",
  "messages": [
    {
      "id": 2512794,
      "postDate": "2023-11-04T20:37:53.210Z",
      "content": "<p>Hello everyone,</p>\n<p>I have a quick question, or rather, I need someone to confirm the following for me because I'm not very familiar with the physics of microscope images…</p>\n<p>So, during training, we have images at 20x magnification, and during testing, we have them at 40x. I believe most people will likely crop patches from the original images and then resize those patches to the model's input size.</p>\n<p>To ensure that the objects in the 40x test images are the same size as in the 20x training images, one should resize the images. From what I currently understand, during the inference of the 40x images, you should first resize the image by a factor of 20/40=0.5 (i.e., shrink it), then crop the patches, and finally resize the patches to the model's input size again… Is that correct? I would greatly appreciate it if someone could confirm this or correct my thinking.</p>\n<p>PS: Of course, you could also resize the 20x images in training initially by a factor of 2.</p>",
      "rawMarkdown": "Hello everyone,\n\nI have a quick question, or rather, I need someone to confirm the following for me because I'm not very familiar with the physics of microscope images...\n\nSo, during training, we have images at 20x magnification, and during testing, we have them at 40x. I believe most people will likely crop patches from the original images and then resize those patches to the model's input size.\n\nTo ensure that the objects in the 40x test images are the same size as in the 20x training images, one should resize the images. From what I currently understand, during the inference of the 40x images, you should first resize the image by a factor of 20/40=0.5 (i.e., shrink it), then crop the patches, and finally resize the patches to the model's input size again... Is that correct? I would greatly appreciate it if someone could confirm this or correct my thinking.\n\nPS: Of course, you could also resize the 20x images in training initially by a factor of 2.",
      "votes": 9
    },
    {
      "id": 2518046,
      "postDate": "2023-11-09T02:14:13.360Z",
      "content": "<p>My approach is the same as yours: <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/451902\" target=\"_blank\">https://www.kaggle.com/competitions/UBC-OCEAN/discussion/451902</a></p>",
      "rawMarkdown": "My approach is the same as yours: https://www.kaggle.com/competitions/UBC-OCEAN/discussion/451902",
      "votes": 1
    },
    {
      "id": 2520602,
      "postDate": "2023-11-11T01:35:31.693Z",
      "content": "<p>all WSIs are 20x, and TMAs are 40x.<br>\nboth train and test set have 20x and 40x images.<br>\ntrain has few TMAs, thus few 40x images, on the other hand test has more TMAs<br>\nHow you treat them is up to you: you either zoom into WSI images with 2x or zoom out TMAs witha factor of 2. Then again maybe you want to train with a smaller size than WSIs…<br>\nButtomline: your models need to see consistent-resized images in both train and test. <br>\nThen you just 'predict = classify' and that is it..you are done<br>\nso no more resizing (to original) :-)</p>",
      "rawMarkdown": "all WSIs are 20x, and TMAs are 40x.\nboth train and test set have 20x and 40x images.\ntrain has few TMAs, thus few 40x images, on the other hand test has more TMAs\nHow you treat them is up to you: you either zoom into WSI images with 2x or zoom out TMAs witha factor of 2. Then again maybe you want to train with a smaller size than WSIs...\nButtomline: your models need to see consistent-resized images in both train and test. \nThen you just 'predict = classify' and that is it..you are done\nso no more resizing (to original) :-)\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 2520867,
          "postDate": "2023-11-11T08:26:43.367Z",
          "content": "<blockquote>\n  <p>Buttomline: your models need to see consistent-resized images in both train and test.</p>\n</blockquote>\n<p>That isn't necessarily correct. Model can learn scale invariance if you feed images without normalizing the scale.</p>",
          "rawMarkdown": ">Buttomline: your models need to see consistent-resized images in both train and test.\n\nThat isn't necessarily correct. Model can learn scale invariance if you feed images without normalizing the scale.",
          "votes": 1,
          "replies": [
            {
              "id": 2520892,
              "postDate": "2023-11-11T09:06:38.177Z",
              "content": "<p>fair point, thank you. </p>\n<p>But I would still create a scale-consistent dataset.</p>\n<p>…and learning 'scale invariance ',, I would leave up to the image augmentation during training (as in Resize or better: RandomResizedCrop)… at least for the baseline,</p>\n<p>because the scene can get pretty complicated with other/advanced techniques for scale invariance, doesn't it (for a semi-hallucinated quicklist of techniques for 'scale invariance in training' refer to chatgpt :-)</p>\n<p>Good thing all can work out of a 'consistently scaled' dataset so you dont have to change your dataset but your preprocessing/model etc.</p>\n<p>what am I missing?</p>",
              "rawMarkdown": "fair point, thank you. \n\nBut I would still create a scale-consistent dataset.\n\n...and learning 'scale invariance ',, I would leave up to the image augmentation during training (as in Resize or better: RandomResizedCrop)... at least for the baseline,\n\nbecause the scene can get pretty complicated with other/advanced techniques for scale invariance, doesn't it (for a semi-hallucinated quicklist of techniques for 'scale invariance in training' refer to chatgpt :-)\n\nGood thing all can work out of a 'consistently scaled' dataset so you dont have to change your dataset but your preprocessing/model etc.\n\nwhat am I missing?\n\n"
            },
            {
              "id": 2521750,
              "postDate": "2023-11-12T03:23:22.823Z",
              "content": "<p>Thank you. All the images in the test set have thumbnails. We cannot determine whether an image is a WSI (Whole Slide Image) or a TMA (Tissue MicroArray) based on the directory. How should we standardize the magnification in the test set? If we magnify the WSIs, the TMAs will be correspondingly magnified as well.</p>",
              "rawMarkdown": "Thank you. All the images in the test set have thumbnails. We cannot determine whether an image is a WSI (Whole Slide Image) or a TMA (Tissue MicroArray) based on the directory. How should we standardize the magnification in the test set? If we magnify the WSIs, the TMAs will be correspondingly magnified as well.\n\n\n\n\n\n",
              "votes": 2
            },
            {
              "id": 2521751,
              "postDate": "2023-11-12T03:25:01.397Z",
              "content": "<p>Perhaps after cutting the patches, we can also determine based on the number of patches.</p>",
              "rawMarkdown": "Perhaps after cutting the patches, we can also determine based on the number of patches.\n\n\n\n\n\n",
              "votes": 1
            },
            {
              "id": 2521801,
              "postDate": "2023-11-12T05:05:25.923Z",
              "content": "<p>from 'Data' tab of the competition: &gt;[train/test]_thumbnails A folder containing smaller .png copies of the whole slide images. <strong>Thumbnails are not provided for TMAs.</strong></p>",
              "rawMarkdown": "from 'Data' tab of the competition: >[train/test]_thumbnails A folder containing smaller .png copies of the whole slide images. **Thumbnails are not provided for TMAs.**"
            },
            {
              "id": 2521998,
              "postDate": "2023-11-12T09:07:21.197Z",
              "content": "<p>I have tried submitting by only obtaining the paths of the thumbnails and found that the score is normal. The score is consistent with the method of obtaining paths considering the TMA directory source. Other top-ranking participants have also tested this and found that TMA should have thumbnails in the test dataset.</p>",
              "rawMarkdown": "I have tried submitting by only obtaining the paths of the thumbnails and found that the score is normal. The score is consistent with the method of obtaining paths considering the TMA directory source. Other top-ranking participants have also tested this and found that TMA should have thumbnails in the test dataset.",
              "votes": 1
            },
            {
              "id": 2522488,
              "postDate": "2023-11-12T17:47:00.577Z",
              "content": "<p>Yes, you are right.<br>\nI have seen it: [<a href=\"https://www.kaggle.com/code/yukkyo/probing-all-test-sample-have-thumbnail/notebook\" target=\"_blank\">https://www.kaggle.com/code/yukkyo/probing-all-test-sample-have-thumbnail/notebook</a>]<br>\nOn the other hand, I think there are ways to findout if a .png file is tma or not.<br>\nFor example, you can create a simple classifier that can predict is_TMA, based on tiles from WSI's and TMAs.<br>\nAssuming your final subtype classifier also works off the tiles (or their patchwork) you will already have the tiles so the overhead of this extra classifier will only be the classication, which you can also do in a downscaled images to save time. </p>\n<p>But probably a more empirical way is to exploit current image size stats, which you can find in this very useful notebook and alike: [<a href=\"https://www.kaggle.com/code/gunesevitan/ubc-ocean-eda\" target=\"_blank\">https://www.kaggle.com/code/gunesevitan/ubc-ocean-eda</a>]<br>\nNamely: that TMA images have</p>\n<ol>\n<li>a square shape </li>\n<li>a two option fixed width/height</li>\n<li>few times smaller size than minimum WSI height/width.</li>\n</ol>\n<p>and, WSI thumbnails,</p>\n<ol>\n<li>have a fixed width of 3000 px.</li>\n</ol>\n<p>Good thing is that once you make few submissions based on above assumptions, you can have a feeling on if they are valid or not, unless there is a big statistical discrepancy between public and private test data, that I doubt,, but you never know:-)</p>",
              "rawMarkdown": "Yes, you are right.\nI have seen it: [https://www.kaggle.com/code/yukkyo/probing-all-test-sample-have-thumbnail/notebook]\nOn the other hand, I think there are ways to findout if a .png file is tma or not.\nFor example, you can create a simple classifier that can predict is_TMA, based on tiles from WSI's and TMAs.\nAssuming your final subtype classifier also works off the tiles (or their patchwork) you will already have the tiles so the overhead of this extra classifier will only be the classication, which you can also do in a downscaled images to save time. \n\nBut probably a more empirical way is to exploit current image size stats, which you can find in this very useful notebook and alike: [https://www.kaggle.com/code/gunesevitan/ubc-ocean-eda]\nNamely: that TMA images have\n1. a square shape \n2. a two option fixed width/height\n3. few times smaller size than minimum WSI height/width.\n\nand, WSI thumbnails,\n1. have a fixed width of 3000 px.\n\nGood thing is that once you make few submissions based on above assumptions, you can have a feeling on if they are valid or not, unless there is a big statistical discrepancy between public and private test data, that I doubt,, but you never know:-)\n",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2512933,
      "postDate": "2023-11-05T03:19:32.630Z",
      "content": "<p>You're right in your approach. Doubling the size of the 20x images seems to be the optimal strategy. However, the real bottleneck, as you've identified, is the time it takes to read each image. Utilizing the entire Whole Slide Image (WSI) directly isn't feasible. To address this, I've developed a method that extracts patches from the images, as detailed in my Kaggle notebook here: <a href=\"https://www.kaggle.com/code/dhinkris/top-10-informative-patches-based-on-std-dev\" target=\"_blank\">Top 10 Informative Patches Based on Std Dev</a>. Despite this, the process remains quite slow and often causes the notebook to crash. I would greatly appreciate any insights or suggestions on this matter.</p>",
      "rawMarkdown": "You're right in your approach. Doubling the size of the 20x images seems to be the optimal strategy. However, the real bottleneck, as you've identified, is the time it takes to read each image. Utilizing the entire Whole Slide Image (WSI) directly isn't feasible. To address this, I've developed a method that extracts patches from the images, as detailed in my Kaggle notebook here: [Top 10 Informative Patches Based on Std Dev](https://www.kaggle.com/code/dhinkris/top-10-informative-patches-based-on-std-dev). Despite this, the process remains quite slow and often causes the notebook to crash. I would greatly appreciate any insights or suggestions on this matter.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2518046,
      "author_name": "Time Master",
      "author_url": "",
      "post_date": "2023-11-09T02:14:13.360000",
      "content": "<p>My approach is the same as yours: <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/451902\" target=\"_blank\">https://www.kaggle.com/competitions/UBC-OCEAN/discussion/451902</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2520602,
      "author_name": "GUNER",
      "author_url": "",
      "post_date": "2023-11-11T01:35:31.693000",
      "content": "<p>all WSIs are 20x, and TMAs are 40x.<br>\nboth train and test set have 20x and 40x images.<br>\ntrain has few TMAs, thus few 40x images, on the other hand test has more TMAs<br>\nHow you treat them is up to you: you either zoom into WSI images with 2x or zoom out TMAs witha factor of 2. Then again maybe you want to train with a smaller size than WSIs…<br>\nButtomline: your models need to see consistent-resized images in both train and test. <br>\nThen you just 'predict = classify' and that is it..you are done<br>\nso no more resizing (to original) :-)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2520867,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2023-11-11T08:26:43.367000",
          "content": "<blockquote>\n  <p>Buttomline: your models need to see consistent-resized images in both train and test.</p>\n</blockquote>\n<p>That isn't necessarily correct. Model can learn scale invariance if you feed images without normalizing the scale.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2520892,
              "author_name": "GUNER",
              "author_url": "",
              "post_date": "2023-11-11T09:06:38.177000",
              "content": "<p>fair point, thank you. </p>\n<p>But I would still create a scale-consistent dataset.</p>\n<p>…and learning 'scale invariance ',, I would leave up to the image augmentation during training (as in Resize or better: RandomResizedCrop)… at least for the baseline,</p>\n<p>because the scene can get pretty complicated with other/advanced techniques for scale invariance, doesn't it (for a semi-hallucinated quicklist of techniques for 'scale invariance in training' refer to chatgpt :-)</p>\n<p>Good thing all can work out of a 'consistently scaled' dataset so you dont have to change your dataset but your preprocessing/model etc.</p>\n<p>what am I missing?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2521750,
              "author_name": "Huang Jin Feng",
              "author_url": "",
              "post_date": "2023-11-12T03:23:22.823000",
              "content": "<p>Thank you. All the images in the test set have thumbnails. We cannot determine whether an image is a WSI (Whole Slide Image) or a TMA (Tissue MicroArray) based on the directory. How should we standardize the magnification in the test set? If we magnify the WSIs, the TMAs will be correspondingly magnified as well.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2521751,
              "author_name": "Huang Jin Feng",
              "author_url": "",
              "post_date": "2023-11-12T03:25:01.397000",
              "content": "<p>Perhaps after cutting the patches, we can also determine based on the number of patches.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2521801,
              "author_name": "GUNER",
              "author_url": "",
              "post_date": "2023-11-12T05:05:25.923000",
              "content": "<p>from 'Data' tab of the competition: &gt;[train/test]_thumbnails A folder containing smaller .png copies of the whole slide images. <strong>Thumbnails are not provided for TMAs.</strong></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2521998,
              "author_name": "Huang Jin Feng",
              "author_url": "",
              "post_date": "2023-11-12T09:07:21.197000",
              "content": "<p>I have tried submitting by only obtaining the paths of the thumbnails and found that the score is normal. The score is consistent with the method of obtaining paths considering the TMA directory source. Other top-ranking participants have also tested this and found that TMA should have thumbnails in the test dataset.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2522488,
              "author_name": "GUNER",
              "author_url": "",
              "post_date": "2023-11-12T17:47:00.577000",
              "content": "<p>Yes, you are right.<br>\nI have seen it: [<a href=\"https://www.kaggle.com/code/yukkyo/probing-all-test-sample-have-thumbnail/notebook\" target=\"_blank\">https://www.kaggle.com/code/yukkyo/probing-all-test-sample-have-thumbnail/notebook</a>]<br>\nOn the other hand, I think there are ways to findout if a .png file is tma or not.<br>\nFor example, you can create a simple classifier that can predict is_TMA, based on tiles from WSI's and TMAs.<br>\nAssuming your final subtype classifier also works off the tiles (or their patchwork) you will already have the tiles so the overhead of this extra classifier will only be the classication, which you can also do in a downscaled images to save time. </p>\n<p>But probably a more empirical way is to exploit current image size stats, which you can find in this very useful notebook and alike: [<a href=\"https://www.kaggle.com/code/gunesevitan/ubc-ocean-eda\" target=\"_blank\">https://www.kaggle.com/code/gunesevitan/ubc-ocean-eda</a>]<br>\nNamely: that TMA images have</p>\n<ol>\n<li>a square shape </li>\n<li>a two option fixed width/height</li>\n<li>few times smaller size than minimum WSI height/width.</li>\n</ol>\n<p>and, WSI thumbnails,</p>\n<ol>\n<li>have a fixed width of 3000 px.</li>\n</ol>\n<p>Good thing is that once you make few submissions based on above assumptions, you can have a feeling on if they are valid or not, unless there is a big statistical discrepancy between public and private test data, that I doubt,, but you never know:-)</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2512933,
      "author_name": "dhinesh",
      "author_url": "",
      "post_date": "2023-11-05T03:19:32.630000",
      "content": "<p>You're right in your approach. Doubling the size of the 20x images seems to be the optimal strategy. However, the real bottleneck, as you've identified, is the time it takes to read each image. Utilizing the entire Whole Slide Image (WSI) directly isn't feasible. To address this, I've developed a method that extracts patches from the images, as detailed in my Kaggle notebook here: <a href=\"https://www.kaggle.com/code/dhinkris/top-10-informative-patches-based-on-std-dev\" target=\"_blank\">Top 10 Informative Patches Based on Std Dev</a>. Despite this, the process remains quite slow and often causes the notebook to crash. I would greatly appreciate any insights or suggestions on this matter.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2512794": "Hello everyone,\n\nI have a quick question, or rather, I need someone to confirm the following for me because I'm not very familiar with the physics of microscope images...\n\nSo, during training, we have images at 20x magnification, and during testing, we have them at 40x. I believe most people will likely crop patches from the original images and then resize those patches to the model's input size.\n\nTo ensure that the objects in the 40x test images are the same size as in the 20x training images, one should resize the images. From what I currently understand, during the inference of the 40x images, you should first resize the image by a factor of 20/40=0.5 (i.e., shrink it), then crop the patches, and finally resize the patches to the model's input size again... Is that correct? I would greatly appreciate it if someone could confirm this or correct my thinking.\n\nPS: Of course, you could also resize the 20x images in training initially by a factor of 2.",
    "2518046": "My approach is the same as yours: https://www.kaggle.com/competitions/UBC-OCEAN/discussion/451902",
    "2520602": "all WSIs are 20x, and TMAs are 40x.\nboth train and test set have 20x and 40x images.\ntrain has few TMAs, thus few 40x images, on the other hand test has more TMAs\nHow you treat them is up to you: you either zoom into WSI images with 2x or zoom out TMAs witha factor of 2. Then again maybe you want to train with a smaller size than WSIs...\nButtomline: your models need to see consistent-resized images in both train and test. \nThen you just 'predict = classify' and that is it..you are done\nso no more resizing (to original) :-)\n\n",
    "2512933": "You're right in your approach. Doubling the size of the 20x images seems to be the optimal strategy. However, the real bottleneck, as you've identified, is the time it takes to read each image. Utilizing the entire Whole Slide Image (WSI) directly isn't feasible. To address this, I've developed a method that extracts patches from the images, as detailed in my Kaggle notebook here: [Top 10 Informative Patches Based on Std Dev](https://www.kaggle.com/code/dhinkris/top-10-informative-patches-based-on-std-dev). Despite this, the process remains quite slow and often causes the notebook to crash. I would greatly appreciate any insights or suggestions on this matter."
  }
}