{
  "id": 446688,
  "title": "Why WSIs are uploaded as PNGs?",
  "url": "/competitions/UBC-OCEAN/discussion/446688",
  "author_name": "Eren Tekin",
  "post_date": "2023-10-12T16:45:51.016000",
  "votes": 25,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I am dealing with WSIs everyday in my work. I don't know how the WSIs scanned but probably they were in some kind of pyramidal image format (e.g tiff, svs) and later converted to png from these formats. If you can upload the original WSIs (maybe you can convert them to tiff using vips?) so that we can work on these slides without getting any memory errors.</p>",
  "messages": [
    {
      "id": 2479491,
      "postDate": "2023-10-12T16:45:51.017Z",
      "content": "<p>I am dealing with WSIs everyday in my work. I don't know how the WSIs scanned but probably they were in some kind of pyramidal image format (e.g tiff, svs) and later converted to png from these formats. If you can upload the original WSIs (maybe you can convert them to tiff using vips?) so that we can work on these slides without getting any memory errors.</p>",
      "rawMarkdown": "I am dealing with WSIs everyday in my work. I don't know how the WSIs scanned but probably they were in some kind of pyramidal image format (e.g tiff, svs) and later converted to png from these formats. If you can upload the original WSIs (maybe you can convert them to tiff using vips?) so that we can work on these slides without getting any memory errors.",
      "votes": 25
    },
    {
      "id": 2479540,
      "postDate": "2023-10-12T17:15:50.923Z",
      "content": "<p>It's a totally fair question! The short answer is that it's my fault and that I ran out of time. The longer answer is that the raw data was indeed a variety of .tiff and .svs formats from the various imaging machines. From a competitions perspective, we essentially can't ever use raw medical images. They must be rebuilt prior to release in order to remove both personal information and (often) the actual class labels from the image metadata. I actually reprocessed the entire dataset to tiled tiff with two different libraries, including pyvips, but caught fatal flaws in the results in each case. Frankly, I ran out of time to try again before the competition launch date. </p>\n<p>The current setup is definitely suboptimal but we're hoping to provide an update early next week that will improve the situation.</p>",
      "rawMarkdown": "It's a totally fair question! The short answer is that it's my fault and that I ran out of time. The longer answer is that the raw data was indeed a variety of .tiff and .svs formats from the various imaging machines. From a competitions perspective, we essentially can't ever use raw medical images. They must be rebuilt prior to release in order to remove both personal information and (often) the actual class labels from the image metadata. I actually reprocessed the entire dataset to tiled tiff with two different libraries, including pyvips, but caught fatal flaws in the results in each case. Frankly, I ran out of time to try again before the competition launch date. \n\nThe current setup is definitely suboptimal but we're hoping to provide an update early next week that will improve the situation.",
      "votes": 17,
      "replies": [
        {
          "id": 2479563,
          "postDate": "2023-10-12T17:37:12.203Z",
          "content": "<p>Of course all medical data should be deidentified before released into the public, I totally understand that! Thank you for the update.</p>",
          "rawMarkdown": "Of course all medical data should be deidentified before released into the public, I totally understand that! Thank you for the update.",
          "votes": 1
        },
        {
          "id": 2482861,
          "postDate": "2023-10-15T10:29:31.600Z",
          "content": "<p>Are there any plans to follow up with raw competition data (i.e. svs or tiff format)?</p>",
          "rawMarkdown": "Are there any plans to follow up with raw competition data (i.e. svs or tiff format)?",
          "votes": 2
        },
        {
          "id": 2489028,
          "postDate": "2023-10-19T16:29:21.340Z",
          "content": "<p>Any updates on if the dataset is going to change <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>? </p>\n<p>Just wondering if I should start downloading the current dataset or wait for an update :)</p>",
          "rawMarkdown": "Any updates on if the dataset is going to change @sohier? \n\nJust wondering if I should start downloading the current dataset or wait for an update :)",
          "votes": 1
        },
        {
          "id": 2489759,
          "postDate": "2023-10-20T07:24:28.360Z",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Yes, me too.<br>\nPlease let me know when you update dataset.</p>",
          "rawMarkdown": "@sohier Yes, me too.\nPlease let me know when you update dataset.",
          "votes": 1
        },
        {
          "id": 2501120,
          "postDate": "2023-10-27T08:57:55.917Z",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> any update on the format issue or any timeline you can provide ? tiff or svs would be much better (and more standard too). High resolution is key for good performances in digital pathology, and keeping the png format would bias the submissions toward suboptimal solutions :/ </p>",
          "rawMarkdown": "@sohier any update on the format issue or any timeline you can provide ? tiff or svs would be much better (and more standard too). High resolution is key for good performances in digital pathology, and keeping the png format would bias the submissions toward suboptimal solutions :/ ",
          "votes": 5,
          "replies": [
            {
              "id": 2501133,
              "postDate": "2023-10-27T09:03:00.313Z",
              "content": "<p>+1! I totally agree with <a href=\"https://www.kaggle.com/simjeg\" target=\"_blank\">@simjeg</a>. By design, one cannot only decode <em>parts</em> of a large PNG image. Instead, one has to decode <em>the whole image</em> at once. This creates a huge overhead, likely preventing most participants from making good use of full resolution images. TIFF/SVS would be definitely better!</p>",
              "rawMarkdown": "+1! I totally agree with @simjeg. By design, one cannot only decode *parts* of a large PNG image. Instead, one has to decode *the whole image* at once. This creates a huge overhead, likely preventing most participants from making good use of full resolution images. TIFF/SVS would be definitely better!",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2482640,
      "postDate": "2023-10-15T07:43:30.417Z",
      "content": "<p>Does using raw tiff files have any advantage over lossless pngs or is it for using 3rd party software like QuPath?</p>",
      "rawMarkdown": "Does using raw tiff files have any advantage over lossless pngs or is it for using 3rd party software like QuPath?",
      "votes": 2,
      "replies": [
        {
          "id": 2482811,
          "postDate": "2023-10-15T09:49:21.983Z",
          "content": "<p>Yes, big images (100kx50k) is hard to deal with. Libraries like OpenSlide make it easier to read these images into tiles, they don't load all the image into the memory, so you don't encounter memory errors too. Also another advantage of raw files is that you have a pyramidal image (first layer 100kx50k second layer 50kx25k so 2x downsample) you can read your tiles in any downsample you want and create your model accordingly. You can do this with PNG's too of course, but it's easier doing it like this. </p>",
          "rawMarkdown": "Yes, big images (100kx50k) is hard to deal with. Libraries like OpenSlide make it easier to read these images into tiles, they don't load all the image into the memory, so you don't encounter memory errors too. Also another advantage of raw files is that you have a pyramidal image (first layer 100kx50k second layer 50kx25k so 2x downsample) you can read your tiles in any downsample you want and create your model accordingly. You can do this with PNG's too of course, but it's easier doing it like this. ",
          "votes": 14
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2479540,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2023-10-12T17:15:50.923000",
      "content": "<p>It's a totally fair question! The short answer is that it's my fault and that I ran out of time. The longer answer is that the raw data was indeed a variety of .tiff and .svs formats from the various imaging machines. From a competitions perspective, we essentially can't ever use raw medical images. They must be rebuilt prior to release in order to remove both personal information and (often) the actual class labels from the image metadata. I actually reprocessed the entire dataset to tiled tiff with two different libraries, including pyvips, but caught fatal flaws in the results in each case. Frankly, I ran out of time to try again before the competition launch date. </p>\n<p>The current setup is definitely suboptimal but we're hoping to provide an update early next week that will improve the situation.</p>",
      "votes": 17,
      "replies": [
        {
          "id": 2479563,
          "author_name": "Eren Tekin",
          "author_url": "",
          "post_date": "2023-10-12T17:37:12.203000",
          "content": "<p>Of course all medical data should be deidentified before released into the public, I totally understand that! Thank you for the update.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2482861,
          "author_name": "阳光开朗大男孩",
          "author_url": "",
          "post_date": "2023-10-15T10:29:31.600000",
          "content": "<p>Are there any plans to follow up with raw competition data (i.e. svs or tiff format)?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2489028,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2023-10-19T16:29:21.340000",
          "content": "<p>Any updates on if the dataset is going to change <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>? </p>\n<p>Just wondering if I should start downloading the current dataset or wait for an update :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2489759,
          "author_name": "Tensor Titan",
          "author_url": "",
          "post_date": "2023-10-20T07:24:28.360000",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Yes, me too.<br>\nPlease let me know when you update dataset.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2501120,
          "author_name": "simjeg",
          "author_url": "",
          "post_date": "2023-10-27T08:57:55.917000",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> any update on the format issue or any timeline you can provide ? tiff or svs would be much better (and more standard too). High resolution is key for good performances in digital pathology, and keeping the png format would bias the submissions toward suboptimal solutions :/ </p>",
          "votes": 5,
          "replies": [
            {
              "id": 2501133,
              "author_name": "jibounet",
              "author_url": "",
              "post_date": "2023-10-27T09:03:00.313000",
              "content": "<p>+1! I totally agree with <a href=\"https://www.kaggle.com/simjeg\" target=\"_blank\">@simjeg</a>. By design, one cannot only decode <em>parts</em> of a large PNG image. Instead, one has to decode <em>the whole image</em> at once. This creates a huge overhead, likely preventing most participants from making good use of full resolution images. TIFF/SVS would be definitely better!</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2482640,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2023-10-15T07:43:30.417000",
      "content": "<p>Does using raw tiff files have any advantage over lossless pngs or is it for using 3rd party software like QuPath?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2482811,
          "author_name": "Eren Tekin",
          "author_url": "",
          "post_date": "2023-10-15T09:49:21.983000",
          "content": "<p>Yes, big images (100kx50k) is hard to deal with. Libraries like OpenSlide make it easier to read these images into tiles, they don't load all the image into the memory, so you don't encounter memory errors too. Also another advantage of raw files is that you have a pyramidal image (first layer 100kx50k second layer 50kx25k so 2x downsample) you can read your tiles in any downsample you want and create your model accordingly. You can do this with PNG's too of course, but it's easier doing it like this. </p>",
          "votes": 14,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2479491": "I am dealing with WSIs everyday in my work. I don't know how the WSIs scanned but probably they were in some kind of pyramidal image format (e.g tiff, svs) and later converted to png from these formats. If you can upload the original WSIs (maybe you can convert them to tiff using vips?) so that we can work on these slides without getting any memory errors.",
    "2479540": "It's a totally fair question! The short answer is that it's my fault and that I ran out of time. The longer answer is that the raw data was indeed a variety of .tiff and .svs formats from the various imaging machines. From a competitions perspective, we essentially can't ever use raw medical images. They must be rebuilt prior to release in order to remove both personal information and (often) the actual class labels from the image metadata. I actually reprocessed the entire dataset to tiled tiff with two different libraries, including pyvips, but caught fatal flaws in the results in each case. Frankly, I ran out of time to try again before the competition launch date. \n\nThe current setup is definitely suboptimal but we're hoping to provide an update early next week that will improve the situation.",
    "2482640": "Does using raw tiff files have any advantage over lossless pngs or is it for using 3rd party software like QuPath?"
  }
}