{
  "id": 452027,
  "title": "Why image format of PNG must be changed or competition solution quality will suffer",
  "url": "/competitions/UBC-OCEAN/discussion/452027",
  "author_name": "David Austin",
  "post_date": "2023-10-31T14:58:13.822000",
  "votes": 45,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Two of the three quality issues <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447886\" target=\"_blank\">previously reported</a> have now been resolved which are steps in the right direction.  However moving away from the PNG format doesn't seem to be getting any traction so I'm hoping to use this thread to convince <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <a href=\"https://www.kaggle.com/masadia\" target=\"_blank\">@masadia</a> that the solution quality will suffer if you don't move to another format.</p>\n<p>Background: PNG files use the Deflate compression algorithm which relies on interdependencies between data elements making it necessary to fully decode the entire image before extracting specific regions or crops.  <a href=\"https://www.kaggle.com/simjeg\" target=\"_blank\">@simjeg</a> correctly pointed out the problem <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/451892#2506246\" target=\"_blank\">here</a>.</p>\n<p>Through some probing I've found that the mean image size of the WSI's in test is between 20-23k sq pixels.  The time to fully decode and open the mean size png using PIL (which uses zlib and is faster than libpng based tools like pyvips for png) is 47s.  Through previous probing I found there's approximately 900 WSI images in test, so the time JUST to decode the WSI's is 47*900/3600 = <strong>11.75 hours if decoding was to be done sequentially out of a 12 hour budget</strong>.  Now there are ways to use multiprocessing and other tricks to bring this down some, but then we quickly start hitting memory bottlenecks due to the size of some of the larger WSI's.</p>\n<p>Do we really want the majority of the compute of this competition to be spent decoding PNG's rather than solving the real problem of cancer subtype classification? To me the reason given here is a [really bad justification](<a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447886#2501775\" target=\"_blank\">https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447886#2501775</a> for not doing the conversion.  Let's find the problem and fix and so this competition can use state of the art medical image processing and not make it about who has the best png engineering skills.</p>",
  "messages": [
    {
      "id": 2506812,
      "postDate": "2023-10-31T14:58:13.823Z",
      "content": "<p>Two of the three quality issues <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447886\" target=\"_blank\">previously reported</a> have now been resolved which are steps in the right direction.  However moving away from the PNG format doesn't seem to be getting any traction so I'm hoping to use this thread to convince <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <a href=\"https://www.kaggle.com/masadia\" target=\"_blank\">@masadia</a> that the solution quality will suffer if you don't move to another format.</p>\n<p>Background: PNG files use the Deflate compression algorithm which relies on interdependencies between data elements making it necessary to fully decode the entire image before extracting specific regions or crops.  <a href=\"https://www.kaggle.com/simjeg\" target=\"_blank\">@simjeg</a> correctly pointed out the problem <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/451892#2506246\" target=\"_blank\">here</a>.</p>\n<p>Through some probing I've found that the mean image size of the WSI's in test is between 20-23k sq pixels.  The time to fully decode and open the mean size png using PIL (which uses zlib and is faster than libpng based tools like pyvips for png) is 47s.  Through previous probing I found there's approximately 900 WSI images in test, so the time JUST to decode the WSI's is 47*900/3600 = <strong>11.75 hours if decoding was to be done sequentially out of a 12 hour budget</strong>.  Now there are ways to use multiprocessing and other tricks to bring this down some, but then we quickly start hitting memory bottlenecks due to the size of some of the larger WSI's.</p>\n<p>Do we really want the majority of the compute of this competition to be spent decoding PNG's rather than solving the real problem of cancer subtype classification? To me the reason given here is a [really bad justification](<a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447886#2501775\" target=\"_blank\">https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447886#2501775</a> for not doing the conversion.  Let's find the problem and fix and so this competition can use state of the art medical image processing and not make it about who has the best png engineering skills.</p>",
      "rawMarkdown": "Two of the three quality issues [previously reported](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447886) have now been resolved which are steps in the right direction.  However moving away from the PNG format doesn't seem to be getting any traction so I'm hoping to use this thread to convince @sohier @masadia that the solution quality will suffer if you don't move to another format.\n\nBackground: PNG files use the Deflate compression algorithm which relies on interdependencies between data elements making it necessary to fully decode the entire image before extracting specific regions or crops.  @simjeg correctly pointed out the problem [here](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/451892#2506246).\n\nThrough some probing I've found that the mean image size of the WSI's in test is between 20-23k sq pixels.  The time to fully decode and open the mean size png using PIL (which uses zlib and is faster than libpng based tools like pyvips for png) is 47s.  Through previous probing I found there's approximately 900 WSI images in test, so the time JUST to decode the WSI's is 47*900/3600 = **11.75 hours if decoding was to be done sequentially out of a 12 hour budget**.  Now there are ways to use multiprocessing and other tricks to bring this down some, but then we quickly start hitting memory bottlenecks due to the size of some of the larger WSI's.\n\nDo we really want the majority of the compute of this competition to be spent decoding PNG's rather than solving the real problem of cancer subtype classification? To me the reason given here is a [really bad justification](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447886#2501775 for not doing the conversion.  Let's find the problem and fix and so this competition can use state of the art medical image processing and not make it about who has the best png engineering skills.",
      "votes": 45
    },
    {
      "id": 2506926,
      "postDate": "2023-10-31T16:24:11.030Z",
      "content": "<p>I totally agree with you that the tiff format is the better choice for this competition, if it was used from the beginning. But I think many people have built their pipelines based on the png format, including me. We have spend a lot of time to deal with the problem that you mentioned. I think it is not fair to change the format in the middle of the competition. That means our efforts are wasted.<br>\nI don't know the final decision of the organizer, but this is my opinion.</p>\n<p>PS.<br>\nThere are 4 CPU cores and 29GB of RAM. As my experience, it can be done in less 6 hours with a simple model. You can use 3 cores to decode the images and 1 core + GPU to do the prediction. It is a little bit tricky, but not impossible.</p>",
      "rawMarkdown": "I totally agree with you that the tiff format is the better choice for this competition, if it was used from the beginning. But I think many people have built their pipelines based on the png format, including me. We have spend a lot of time to deal with the problem that you mentioned. I think it is not fair to change the format in the middle of the competition. That means our efforts are wasted.\nI don't know the final decision of the organizer, but this is my opinion.\n\nPS.\nThere are 4 CPU cores and 29GB of RAM. As my experience, it can be done in less 6 hours with a simple model. You can use 3 cores to decode the images and 1 core + GPU to do the prediction. It is a little bit tricky, but not impossible.",
      "votes": 11,
      "replies": [
        {
          "id": 2506943,
          "postDate": "2023-10-31T16:38:13.693Z",
          "content": "<p>Yes, that's a downside of a format change for sure.  But if the format moved to tiff you could still just decode the entire image into an array and then be able to maintain your existing pipeline from there, no?</p>",
          "rawMarkdown": "Yes, that's a downside of a format change for sure.  But if the format moved to tiff you could still just decode the entire image into an array and then be able to maintain your existing pipeline from there, no?",
          "votes": 3,
          "replies": [
            {
              "id": 2506958,
              "postDate": "2023-10-31T16:54:48.677Z",
              "content": "<p>Yes, but it doesn't means no cost. I don't know if the vips can deal with the tiff format just like png. <br>\nThe difference is, someone want to save their time they have spent, and someone want to save their time they will spend. I think the organizer should consider both of them.</p>",
              "rawMarkdown": "Yes, but it doesn't means no cost. I don't know if the vips can deal with the tiff format just like png. \nThe difference is, someone want to save their time they have spent, and someone want to save their time they will spend. I think the organizer should consider both of them.",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2506979,
      "postDate": "2023-10-31T17:08:48.553Z",
      "content": "<p>Maybe keeping the png and adding the tiff are the best solution.😁</p>",
      "rawMarkdown": "Maybe keeping the png and adding the tiff are the best solution.😁",
      "votes": 6
    },
    {
      "id": 2522645,
      "postDate": "2023-11-12T22:14:19.640Z",
      "content": "<p>I'm convinced, sounds like adding the TIFF files is necessary for this competition to really be meaningful.</p>",
      "rawMarkdown": "I'm convinced, sounds like adding the TIFF files is necessary for this competition to really be meaningful.",
      "votes": 3
    },
    {
      "id": 2508180,
      "postDate": "2023-11-01T14:26:46.103Z",
      "content": "<p>I agree as well. For the first 10 images of the training set, the time to convert a WSI png to tif using pyvips ranges from 23 to 150, with an average of 65 seconds…. thus it would require over 16 hours just for conversion alone. <br>\nHow can we apply and build upon the current methods to classify and find the \"outlier\" class when most of the time is spent on file conversion and resource management?</p>",
      "rawMarkdown": "I agree as well. For the first 10 images of the training set, the time to convert a WSI png to tif using pyvips ranges from 23 to 150, with an average of 65 seconds.... thus it would require over 16 hours just for conversion alone. \nHow can we apply and build upon the current methods to classify and find the \"outlier\" class when most of the time is spent on file conversion and resource management?",
      "votes": 4
    },
    {
      "id": 2524973,
      "postDate": "2023-11-14T16:47:58.533Z",
      "content": "<p>Could this be the reason why my submission fails?</p>",
      "rawMarkdown": "Could this be the reason why my submission fails?\n  "
    },
    {
      "id": 2509145,
      "postDate": "2023-11-02T07:49:34.147Z",
      "content": "<p>Yes, reading the image takes too much memory and time. It takes me 80 seconds to read a 2GB image, while the inference time for this image is less than 20 seconds.</p>",
      "rawMarkdown": "Yes, reading the image takes too much memory and time. It takes me 80 seconds to read a 2GB image, while the inference time for this image is less than 20 seconds."
    }
  ],
  "comments": [
    {
      "id": 2506926,
      "author_name": "Johnny Lee",
      "author_url": "",
      "post_date": "2023-10-31T16:24:11.030000",
      "content": "<p>I totally agree with you that the tiff format is the better choice for this competition, if it was used from the beginning. But I think many people have built their pipelines based on the png format, including me. We have spend a lot of time to deal with the problem that you mentioned. I think it is not fair to change the format in the middle of the competition. That means our efforts are wasted.<br>\nI don't know the final decision of the organizer, but this is my opinion.</p>\n<p>PS.<br>\nThere are 4 CPU cores and 29GB of RAM. As my experience, it can be done in less 6 hours with a simple model. You can use 3 cores to decode the images and 1 core + GPU to do the prediction. It is a little bit tricky, but not impossible.</p>",
      "votes": 11,
      "replies": [
        {
          "id": 2506943,
          "author_name": "David Austin",
          "author_url": "",
          "post_date": "2023-10-31T16:38:13.693000",
          "content": "<p>Yes, that's a downside of a format change for sure.  But if the format moved to tiff you could still just decode the entire image into an array and then be able to maintain your existing pipeline from there, no?</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2506958,
              "author_name": "Johnny Lee",
              "author_url": "",
              "post_date": "2023-10-31T16:54:48.677000",
              "content": "<p>Yes, but it doesn't means no cost. I don't know if the vips can deal with the tiff format just like png. <br>\nThe difference is, someone want to save their time they have spent, and someone want to save their time they will spend. I think the organizer should consider both of them.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2506979,
      "author_name": "Johnny Lee",
      "author_url": "",
      "post_date": "2023-10-31T17:08:48.553000",
      "content": "<p>Maybe keeping the png and adding the tiff are the best solution.😁</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2522645,
      "author_name": "Joe Norton",
      "author_url": "",
      "post_date": "2023-11-12T22:14:19.640000",
      "content": "<p>I'm convinced, sounds like adding the TIFF files is necessary for this competition to really be meaningful.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2508180,
      "author_name": "Noli Alonso",
      "author_url": "",
      "post_date": "2023-11-01T14:26:46.103000",
      "content": "<p>I agree as well. For the first 10 images of the training set, the time to convert a WSI png to tif using pyvips ranges from 23 to 150, with an average of 65 seconds…. thus it would require over 16 hours just for conversion alone. <br>\nHow can we apply and build upon the current methods to classify and find the \"outlier\" class when most of the time is spent on file conversion and resource management?</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2524973,
      "author_name": "habeeb hassan",
      "author_url": "",
      "post_date": "2023-11-14T16:47:58.533000",
      "content": "<p>Could this be the reason why my submission fails?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2509145,
      "author_name": "WangXuC",
      "author_url": "",
      "post_date": "2023-11-02T07:49:34.147000",
      "content": "<p>Yes, reading the image takes too much memory and time. It takes me 80 seconds to read a 2GB image, while the inference time for this image is less than 20 seconds.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2506812": "Two of the three quality issues [previously reported](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447886) have now been resolved which are steps in the right direction.  However moving away from the PNG format doesn't seem to be getting any traction so I'm hoping to use this thread to convince @sohier @masadia that the solution quality will suffer if you don't move to another format.\n\nBackground: PNG files use the Deflate compression algorithm which relies on interdependencies between data elements making it necessary to fully decode the entire image before extracting specific regions or crops.  @simjeg correctly pointed out the problem [here](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/451892#2506246).\n\nThrough some probing I've found that the mean image size of the WSI's in test is between 20-23k sq pixels.  The time to fully decode and open the mean size png using PIL (which uses zlib and is faster than libpng based tools like pyvips for png) is 47s.  Through previous probing I found there's approximately 900 WSI images in test, so the time JUST to decode the WSI's is 47*900/3600 = **11.75 hours if decoding was to be done sequentially out of a 12 hour budget**.  Now there are ways to use multiprocessing and other tricks to bring this down some, but then we quickly start hitting memory bottlenecks due to the size of some of the larger WSI's.\n\nDo we really want the majority of the compute of this competition to be spent decoding PNG's rather than solving the real problem of cancer subtype classification? To me the reason given here is a [really bad justification](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447886#2501775 for not doing the conversion.  Let's find the problem and fix and so this competition can use state of the art medical image processing and not make it about who has the best png engineering skills.",
    "2506926": "I totally agree with you that the tiff format is the better choice for this competition, if it was used from the beginning. But I think many people have built their pipelines based on the png format, including me. We have spend a lot of time to deal with the problem that you mentioned. I think it is not fair to change the format in the middle of the competition. That means our efforts are wasted.\nI don't know the final decision of the organizer, but this is my opinion.\n\nPS.\nThere are 4 CPU cores and 29GB of RAM. As my experience, it can be done in less 6 hours with a simple model. You can use 3 cores to decode the images and 1 core + GPU to do the prediction. It is a little bit tricky, but not impossible.",
    "2506979": "Maybe keeping the png and adding the tiff are the best solution.😁",
    "2522645": "I'm convinced, sounds like adding the TIFF files is necessary for this competition to really be meaningful.",
    "2508180": "I agree as well. For the first 10 images of the training set, the time to convert a WSI png to tif using pyvips ranges from 23 to 150, with an average of 65 seconds.... thus it would require over 16 hours just for conversion alone. \nHow can we apply and build upon the current methods to classify and find the \"outlier\" class when most of the time is spent on file conversion and resource management?",
    "2524973": "Could this be the reason why my submission fails?\n  ",
    "2509145": "Yes, reading the image takes too much memory and time. It takes me 80 seconds to read a 2GB image, while the inference time for this image is less than 20 seconds."
  }
}