{
  "id": 447886,
  "title": "Please fix the many quality issues with the competition data and leaderboard",
  "url": "/competitions/UBC-OCEAN/discussion/447886",
  "author_name": "David Austin",
  "post_date": "2023-10-17T16:39:40.595000",
  "votes": 57,
  "comment_count": 25,
  "views": 0,
  "content": "<p>It's always exciting to participate in medical image competitions, especially with the unique aspect of out-of-distribution test data we don't always see in most comps. However it's concerning to see several ongoing issues despite being brought to the attention of the Kaggle staff. I understand there's extensive preparation that goes into running these competitions, but for participants who invest significant effort it's also important know that any raised concerns are addressed promptly. Here are the critical issues that need resolution:</p>\n<ol>\n<li><p><a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/446760\" target=\"_blank\">LB metric has an error</a>.  This was highlighted 5 days ago with no response from kaggle to even say they're looking into it.</p></li>\n<li><p>There are many quality issues with the images, as <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447063\" target=\"_blank\">originally highlighted </a>by <a href=\"https://www.kaggle.com/yukkyo\" target=\"_blank\">@yukkyo</a>.  I've identified 46 images with clear problems, many of them looked to be flipped relative to where the mask was supposed to be selected, I'm sharing some examples here.  Not knowing any better I assume these issues are present in test as well.  We all expect a bit of noise in the data but 8% of the images like this is too much.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F888191%2F29480ebbf511c3d7f86552d1e1772565%2Fcombined_image.jpg?generation=1697560264510346&amp;alt=media\" alt=\"\"></p></li>\n<li><p>PNG was a bad choice for image format and <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/446688#2479540\" target=\"_blank\">acknowledged as such</a>.  If the data format is going to change it needs to happen quickly, having to re-download over 700GB of data is not a good practice.</p></li>\n</ol>\n<p>Please let us know what's being done about these issues so we know how we want to proceed. <a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a> <a href=\"https://www.kaggle.com/homesmac\" target=\"_blank\">@homesmac</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
  "messages": [
    {
      "id": 2486047,
      "postDate": "2023-10-17T16:39:40.597Z",
      "content": "<p>It's always exciting to participate in medical image competitions, especially with the unique aspect of out-of-distribution test data we don't always see in most comps. However it's concerning to see several ongoing issues despite being brought to the attention of the Kaggle staff. I understand there's extensive preparation that goes into running these competitions, but for participants who invest significant effort it's also important know that any raised concerns are addressed promptly. Here are the critical issues that need resolution:</p>\n<ol>\n<li><p><a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/446760\" target=\"_blank\">LB metric has an error</a>.  This was highlighted 5 days ago with no response from kaggle to even say they're looking into it.</p></li>\n<li><p>There are many quality issues with the images, as <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447063\" target=\"_blank\">originally highlighted </a>by <a href=\"https://www.kaggle.com/yukkyo\" target=\"_blank\">@yukkyo</a>.  I've identified 46 images with clear problems, many of them looked to be flipped relative to where the mask was supposed to be selected, I'm sharing some examples here.  Not knowing any better I assume these issues are present in test as well.  We all expect a bit of noise in the data but 8% of the images like this is too much.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F888191%2F29480ebbf511c3d7f86552d1e1772565%2Fcombined_image.jpg?generation=1697560264510346&amp;alt=media\" alt=\"\"></p></li>\n<li><p>PNG was a bad choice for image format and <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/446688#2479540\" target=\"_blank\">acknowledged as such</a>.  If the data format is going to change it needs to happen quickly, having to re-download over 700GB of data is not a good practice.</p></li>\n</ol>\n<p>Please let us know what's being done about these issues so we know how we want to proceed. <a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a> <a href=\"https://www.kaggle.com/homesmac\" target=\"_blank\">@homesmac</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
      "rawMarkdown": "It's always exciting to participate in medical image competitions, especially with the unique aspect of out-of-distribution test data we don't always see in most comps. However it's concerning to see several ongoing issues despite being brought to the attention of the Kaggle staff. I understand there's extensive preparation that goes into running these competitions, but for participants who invest significant effort it's also important know that any raised concerns are addressed promptly. Here are the critical issues that need resolution:\n\n\n1. [LB metric has an error](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/446760).  This was highlighted 5 days ago with no response from kaggle to even say they're looking into it.\n\n2. There are many quality issues with the images, as [originally highlighted ](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447063)by @yukkyo.  I've identified 46 images with clear problems, many of them looked to be flipped relative to where the mask was supposed to be selected, I'm sharing some examples here.  Not knowing any better I assume these issues are present in test as well.  We all expect a bit of noise in the data but 8% of the images like this is too much.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F888191%2F29480ebbf511c3d7f86552d1e1772565%2Fcombined_image.jpg?generation=1697560264510346&alt=media)\n\n3. PNG was a bad choice for image format and [acknowledged as such](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/446688#2479540).  If the data format is going to change it needs to happen quickly, having to re-download over 700GB of data is not a good practice.\n\nPlease let us know what's being done about these issues so we know how we want to proceed. @ashleychow @homesmac @sohier ",
      "votes": 56
    },
    {
      "id": 2486488,
      "postDate": "2023-10-18T01:08:23.600Z",
      "content": "<p>Hello,</p>\n<p>We want to express our gratitude for raising this important issue.</p>\n<p>It appears that a displacement of the masks occurred during the Kaggle production stage. Our intention was to offer masks that exclusively encompass the tissue while eliminating the background, thus simplifying the task for participants.<br>\nWe are working with the Kaggle team to resolve this issue as soon as possible.</p>\n<p>Regards,<br>\nMaryam</p>",
      "rawMarkdown": "Hello,\n\nWe want to express our gratitude for raising this important issue.\n\nIt appears that a displacement of the masks occurred during the Kaggle production stage. Our intention was to offer masks that exclusively encompass the tissue while eliminating the background, thus simplifying the task for participants.\nWe are working with the Kaggle team to resolve this issue as soon as possible.\n\nRegards,\nMaryam",
      "votes": 16,
      "replies": [
        {
          "id": 2489761,
          "postDate": "2023-10-20T07:25:39.997Z",
          "content": "<p>When will be the data update? It takes 4 days for me to download the entire dataset so I have to wait for it and do it once. </p>",
          "rawMarkdown": "When will be the data update? It takes 4 days for me to download the entire dataset so I have to wait for it and do it once. ",
          "votes": 12,
          "replies": [
            {
              "id": 2496056,
              "postDate": "2023-10-23T18:02:03.133Z",
              "content": "<p><a href=\"https://www.kaggle.com/masadia\" target=\"_blank\">@masadia</a>, do you have any updates on this? </p>\n<p>Excited to start the competition, but I have the same concerns as <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>. </p>",
              "rawMarkdown": "@masadia, do you have any updates on this? \n\nExcited to start the competition, but I have the same concerns as @gunesevitan. ",
              "votes": 2
            },
            {
              "id": 2496594,
              "postDate": "2023-10-24T06:41:45.063Z",
              "content": "<p>There is no update yet. You can view the thumbnails in <code>train_thumbnails</code> in the data, such as <code>11263_thumbnail.png</code> to check whether the competition organizer has fixed the error.</p>",
              "rawMarkdown": "There is no update yet. You can view the thumbnails in `train_thumbnails` in the data, such as `11263_thumbnail.png` to check whether the competition organizer has fixed the error.",
              "votes": 1
            },
            {
              "id": 2496630,
              "postDate": "2023-10-24T07:02:32.620Z",
              "content": "<p>I really don't think the data update should take this long. I hope competition deadline will be extended. It is needed this time unlike the last RSNA competition :D</p>",
              "rawMarkdown": "I really don't think the data update should take this long. I hope competition deadline will be extended. It is needed this time unlike the last RSNA competition :D",
              "votes": 5
            },
            {
              "id": 2496913,
              "postDate": "2023-10-24T10:29:56.373Z",
              "content": "<p>Why do you think so? We have more than 2 months to go, seems like quite a lot of time without any extensions </p>",
              "rawMarkdown": "Why do you think so? We have more than 2 months to go, seems like quite a lot of time without any extensions "
            },
            {
              "id": 2496922,
              "postDate": "2023-10-24T10:36:28.230Z",
              "content": "<p>Kaggle allocated 3 months for this competition but lots of people haven't started yet since they're waiting for the update. That's a loss for organizers and Kaggle should compensate that by extending the deadline.</p>",
              "rawMarkdown": "Kaggle allocated 3 months for this competition but lots of people haven't started yet since they're waiting for the update. That's a loss for organizers and Kaggle should compensate that by extending the deadline.",
              "votes": 4
            },
            {
              "id": 2496930,
              "postDate": "2023-10-24T10:44:03.737Z",
              "content": "<p>There is some truth behind that. However, I personally started. I realize that some of the data is broken, but so what? Implement the pipeline, find some tricks. When data is fixed - re-run your experiments :) The time with broken data is not lost </p>",
              "rawMarkdown": "There is some truth behind that. However, I personally started. I realize that some of the data is broken, but so what? Implement the pipeline, find some tricks. When data is fixed - re-run your experiments :) The time with broken data is not lost ",
              "votes": 5
            }
          ]
        }
      ]
    },
    {
      "id": 2486129,
      "postDate": "2023-10-17T17:47:34.970Z",
      "content": "<p>I fully agree with you. Additionally I am also concerned about the following points</p>\n<ol>\n<li>The small number of training TMAs</li>\n</ol>\n<p>We are given 25 TMAs for training. However, the <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/data\" target=\"_blank\">Dataset Description</a> says that majority of the tests (roughly 2000) are TMAs. Is 25 TMAs really a sufficient number? Especially in medical imaging competitions, even different data sources (e.g., hospitals) can be a challenge. With this in mind, I am wondering if the small number of TMAs given for training is an appropriate design.</p>\n<p>Of course, this is a matter of ingenuity on the part of the participants, and may be a challenge for the host. Personally, however, I would like Kaggle staff to reconsider whether this number is appropriate.</p>",
      "rawMarkdown": "I fully agree with you. Additionally I am also concerned about the following points\n\n1. The small number of training TMAs\n\nWe are given 25 TMAs for training. However, the [Dataset Description](https://www.kaggle.com/competitions/UBC-OCEAN/data) says that majority of the tests (roughly 2000) are TMAs. Is 25 TMAs really a sufficient number? Especially in medical imaging competitions, even different data sources (e.g., hospitals) can be a challenge. With this in mind, I am wondering if the small number of TMAs given for training is an appropriate design.\n\nOf course, this is a matter of ingenuity on the part of the participants, and may be a challenge for the host. Personally, however, I would like Kaggle staff to reconsider whether this number is appropriate.",
      "votes": 8,
      "replies": [
        {
          "id": 2486217,
          "postDate": "2023-10-17T18:56:55.030Z",
          "content": "<p>The difference in TMA/WSI distribution between train and test is indeed steep but at least it's a conscious design decision that intends to challenge models to handle different distributions.  If it makes you feel (slightly) better, through some probing I found there's ~1100 TMA's in test so \"majority\" is something like 55%.</p>",
          "rawMarkdown": "The difference in TMA/WSI distribution between train and test is indeed steep but at least it's a conscious design decision that intends to challenge models to handle different distributions.  If it makes you feel (slightly) better, through some probing I found there's ~1100 TMA's in test so \"majority\" is something like 55%.",
          "votes": 23,
          "replies": [
            {
              "id": 2499096,
              "postDate": "2023-10-25T18:19:05.870Z",
              "content": "<p>What do you mean by \"some probing\"? Sounds interesting 🤔</p>",
              "rawMarkdown": "What do you mean by \"some probing\"? Sounds interesting 🤔"
            },
            {
              "id": 2499233,
              "postDate": "2023-10-25T21:28:37.503Z",
              "content": "<p>Here's an example of a <a href=\"https://www.kaggle.com/code/yukkyo/probing-all-test-sample-have-thumbnail\" target=\"_blank\">notebook</a> to probe the LB from <a href=\"https://www.kaggle.com/yukkyo\" target=\"_blank\">@yukkyo</a>.  The basic idea is you can single bit probe a condition in the test set by having the code return one of two known states.  </p>",
              "rawMarkdown": "Here's an example of a [notebook](https://www.kaggle.com/code/yukkyo/probing-all-test-sample-have-thumbnail) to probe the LB from @yukkyo.  The basic idea is you can single bit probe a condition in the test set by having the code return one of two known states.  ",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2498381,
      "postDate": "2023-10-25T09:21:01.197Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/masadia\" target=\"_blank\">@masadia</a>, by when can we expect an update (if there is going to be one)?</p>\n<p><a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a> <a href=\"https://www.kaggle.com/homesmac\" target=\"_blank\">@homesmac</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a></p>",
      "rawMarkdown": "Hey @masadia, by when can we expect an update (if there is going to be one)?\n\n@ashleychow @homesmac @sohier",
      "votes": 3,
      "replies": [
        {
          "id": 2499005,
          "postDate": "2023-10-25T16:55:27.550Z",
          "content": "<p>I'm reprocessing images with masking problems now. Given how long it takes to process a 700 GB data bundle, the corrections won't be available for at least another day or two. Early next week is most likely. I'll provide a list of the updated images to make it possible to avoid downloading the entire data bundle again. </p>\n<p>The good news is that I believe we've identified the root cause and the problems were limited to a specific subset of the training data (i.e. the test set won't change at all). <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> seems to have found nearly all of them.</p>",
          "rawMarkdown": "I'm reprocessing images with masking problems now. Given how long it takes to process a 700 GB data bundle, the corrections won't be available for at least another day or two. Early next week is most likely. I'll provide a list of the updated images to make it possible to avoid downloading the entire data bundle again. \n\nThe good news is that I believe we've identified the root cause and the problems were limited to a specific subset of the training data (i.e. the test set won't change at all). @tivfrvqhs5 seems to have found nearly all of them.",
          "votes": 8,
          "replies": [
            {
              "id": 2499057,
              "postDate": "2023-10-25T17:45:42.027Z",
              "content": "<p>Does that mean the train and test images are going to remain in png format?</p>",
              "rawMarkdown": "Does that mean the train and test images are going to remain in png format?",
              "votes": 1
            },
            {
              "id": 2499079,
              "postDate": "2023-10-25T17:57:13.460Z",
              "content": "<p>Yes.                </p>",
              "rawMarkdown": "Yes.                ",
              "votes": -2
            },
            {
              "id": 2500566,
              "postDate": "2023-10-26T19:20:09.940Z",
              "content": "<p>Does this mean the data will stay in png format for the remainder of the challenge, or will it be converted to e.g. tif at a later date? I'd like to argue that magnification differences between the slides can throw off models considerably, considering that most networks are not scale invariant. Having them in a format where the levels are stored, along with perhaps the magnification metadata would be helpful.</p>",
              "rawMarkdown": "Does this mean the data will stay in png format for the remainder of the challenge, or will it be converted to e.g. tif at a later date? I'd like to argue that magnification differences between the slides can throw off models considerably, considering that most networks are not scale invariant. Having them in a format where the levels are stored, along with perhaps the magnification metadata would be helpful.",
              "votes": 3
            },
            {
              "id": 2501145,
              "postDate": "2023-10-27T09:09:04.797Z",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> If you can change the image format in favor of TIFF or SVS, please do it. Unless mistaken, one has to decode the <strong>complete</strong> PNG image at once. For large images (&gt;2GB), this decoding is really time-consuming (&gt;1min per image). I am convinced that this overhead prevents participants from making good use of full resolution images in this competition. </p>\n<p>EDIT: I've been working for several years in the field of computational pathology. IMHO, PNG is a highly unusual choice of image format for digital pathology. Image formats such as SVS are much more suited to digital pathology. If needed, <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>, I can provide some help! Do not hesitate to get in touch!</p>",
              "rawMarkdown": "@sohier If you can change the image format in favor of TIFF or SVS, please do it. Unless mistaken, one has to decode the **complete** PNG image at once. For large images (>2GB), this decoding is really time-consuming (>1min per image). I am convinced that this overhead prevents participants from making good use of full resolution images in this competition. \n\nEDIT: I've been working for several years in the field of computational pathology. IMHO, PNG is a highly unusual choice of image format for digital pathology. Image formats such as SVS are much more suited to digital pathology. If needed, @sohier, I can provide some help! Do not hesitate to get in touch!",
              "votes": 11
            },
            {
              "id": 2501191,
              "postDate": "2023-10-27T09:54:28.957Z",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> For info, see below a list of previous competitions involving digital pathology. None used PNG images.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/histopathologic-cancer-detection/data\" target=\"_blank\">https://www.kaggle.com/competitions/histopathologic-cancer-detection/data</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/hubmap-kidney-segmentation/data\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-kidney-segmentation/data</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/prostate-cancer-grade-assessment/data\" target=\"_blank\">https://www.kaggle.com/competitions/prostate-cancer-grade-assessment/data</a></li>\n<li><a href=\"https://grand-challenge.org/challenges/?search=&amp;modalities=11&amp;educational=unknown&amp;status=&amp;submit=Apply+Filters\" target=\"_blank\">https://grand-challenge.org/challenges/?search=&amp;modalities=11&amp;educational=unknown&amp;status=&amp;submit=Apply+Filters</a></li>\n</ul>",
              "rawMarkdown": "@sohier For info, see below a list of previous competitions involving digital pathology. None used PNG images.\n- https://www.kaggle.com/competitions/histopathologic-cancer-detection/data\n- https://www.kaggle.com/competitions/hubmap-kidney-segmentation/data\n- https://www.kaggle.com/competitions/prostate-cancer-grade-assessment/data\n- https://grand-challenge.org/challenges/?search=&modalities=11&educational=unknown&status=&submit=Apply+Filters",
              "votes": 6
            },
            {
              "id": 2501435,
              "postDate": "2023-10-27T12:47:27.793Z",
              "content": "<p>Also worth mentioning, it's two lines of code to do the conversion.  No good reason not to do it.</p>\n<pre><code> png  pnglist:\n    image = pyvips.Image.new_from_file(f)\n    image.tiffsave(f, =, =, =)\n</code></pre>",
              "rawMarkdown": "Also worth mentioning, it's two lines of code to do the conversion.  No good reason not to do it.\n\n```\nfor png in pnglist:\n    image = pyvips.Image.new_from_file(f\"{png}\")\n    image.tiffsave(f\"{png[:-4]}.tif\", tile=True, pyramid=True, bigtiff=True)\n```",
              "votes": 8
            },
            {
              "id": 2501475,
              "postDate": "2023-10-27T13:28:20.287Z",
              "content": "<p>Exactly! 💯</p>",
              "rawMarkdown": "Exactly! 💯"
            },
            {
              "id": 2501775,
              "postDate": "2023-10-27T16:42:18.100Z",
              "content": "<p><a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> From memory, when I went that route a few months ago a small portion of the images ended up with very odd color artifacts. Think green slide backgrounds.</p>",
              "rawMarkdown": "@tivfrvqhs5 From memory, when I went that route a few months ago a small portion of the images ended up with very odd color artifacts. Think green slide backgrounds."
            },
            {
              "id": 2501784,
              "postDate": "2023-10-27T16:44:33.960Z",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> If you're not going to change the image format, please increase the maximum duration of notebooks to 16H or 18H.  Eventually, we - the participants - are ought to deliver the best models for ovarian cancer subtyping, for the good of the patients! I strongly believe that this will only be possible if participants can thoroughly analyze the full resolution images (which is highly unlikely given the 12H limit).</p>",
              "rawMarkdown": "@sohier If you're not going to change the image format, please increase the maximum duration of notebooks to 16H or 18H.  Eventually, we - the participants - are ought to deliver the best models for ovarian cancer subtyping, for the good of the patients! I strongly believe that this will only be possible if participants can thoroughly analyze the full resolution images (which is highly unlikely given the 12H limit).",
              "votes": 8
            },
            {
              "id": 2501830,
              "postDate": "2023-10-27T17:34:29.037Z",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> If you can reference a train png image that had green artifacts after conversion to tiff, I'm pretty sure the community could help figure it out</p>",
              "rawMarkdown": "@sohier If you can reference a train png image that had green artifacts after conversion to tiff, I'm pretty sure the community could help figure it out",
              "votes": 5
            },
            {
              "id": 2503560,
              "postDate": "2023-10-29T07:51:52.770Z",
              "content": "<p><a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a> <a href=\"https://www.kaggle.com/jonathanmcwilliams\" target=\"_blank\">@jonathanmcwilliams</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Any update/decision on this?</p>",
              "rawMarkdown": "@ashleychow @jonathanmcwilliams @sohier Any update/decision on this?"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2486488,
      "author_name": "masadia",
      "author_url": "",
      "post_date": "2023-10-18T01:08:23.600000",
      "content": "<p>Hello,</p>\n<p>We want to express our gratitude for raising this important issue.</p>\n<p>It appears that a displacement of the masks occurred during the Kaggle production stage. Our intention was to offer masks that exclusively encompass the tissue while eliminating the background, thus simplifying the task for participants.<br>\nWe are working with the Kaggle team to resolve this issue as soon as possible.</p>\n<p>Regards,<br>\nMaryam</p>",
      "votes": 16,
      "replies": [
        {
          "id": 2489761,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2023-10-20T07:25:39.997000",
          "content": "<p>When will be the data update? It takes 4 days for me to download the entire dataset so I have to wait for it and do it once. </p>",
          "votes": 12,
          "replies": [
            {
              "id": 2496056,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2023-10-23T18:02:03.133000",
              "content": "<p><a href=\"https://www.kaggle.com/masadia\" target=\"_blank\">@masadia</a>, do you have any updates on this? </p>\n<p>Excited to start the competition, but I have the same concerns as <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2496594,
              "author_name": "m1dsolo",
              "author_url": "",
              "post_date": "2023-10-24T06:41:45.063000",
              "content": "<p>There is no update yet. You can view the thumbnails in <code>train_thumbnails</code> in the data, such as <code>11263_thumbnail.png</code> to check whether the competition organizer has fixed the error.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2496630,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2023-10-24T07:02:32.620000",
              "content": "<p>I really don't think the data update should take this long. I hope competition deadline will be extended. It is needed this time unlike the last RSNA competition :D</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2496913,
              "author_name": "Ivan Panshin",
              "author_url": "",
              "post_date": "2023-10-24T10:29:56.373000",
              "content": "<p>Why do you think so? We have more than 2 months to go, seems like quite a lot of time without any extensions </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2496922,
              "author_name": "Gunes Evitan",
              "author_url": "",
              "post_date": "2023-10-24T10:36:28.230000",
              "content": "<p>Kaggle allocated 3 months for this competition but lots of people haven't started yet since they're waiting for the update. That's a loss for organizers and Kaggle should compensate that by extending the deadline.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2496930,
              "author_name": "Ivan Panshin",
              "author_url": "",
              "post_date": "2023-10-24T10:44:03.737000",
              "content": "<p>There is some truth behind that. However, I personally started. I realize that some of the data is broken, but so what? Implement the pipeline, find some tricks. When data is fixed - re-run your experiments :) The time with broken data is not lost </p>",
              "votes": 5,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2486129,
      "author_name": "fam_taro",
      "author_url": "",
      "post_date": "2023-10-17T17:47:34.970000",
      "content": "<p>I fully agree with you. Additionally I am also concerned about the following points</p>\n<ol>\n<li>The small number of training TMAs</li>\n</ol>\n<p>We are given 25 TMAs for training. However, the <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/data\" target=\"_blank\">Dataset Description</a> says that majority of the tests (roughly 2000) are TMAs. Is 25 TMAs really a sufficient number? Especially in medical imaging competitions, even different data sources (e.g., hospitals) can be a challenge. With this in mind, I am wondering if the small number of TMAs given for training is an appropriate design.</p>\n<p>Of course, this is a matter of ingenuity on the part of the participants, and may be a challenge for the host. Personally, however, I would like Kaggle staff to reconsider whether this number is appropriate.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2486217,
          "author_name": "David Austin",
          "author_url": "",
          "post_date": "2023-10-17T18:56:55.030000",
          "content": "<p>The difference in TMA/WSI distribution between train and test is indeed steep but at least it's a conscious design decision that intends to challenge models to handle different distributions.  If it makes you feel (slightly) better, through some probing I found there's ~1100 TMA's in test so \"majority\" is something like 55%.</p>",
          "votes": 23,
          "replies": [
            {
              "id": 2499096,
              "author_name": "Patchef",
              "author_url": "",
              "post_date": "2023-10-25T18:19:05.870000",
              "content": "<p>What do you mean by \"some probing\"? Sounds interesting 🤔</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2499233,
              "author_name": "David Austin",
              "author_url": "",
              "post_date": "2023-10-25T21:28:37.503000",
              "content": "<p>Here's an example of a <a href=\"https://www.kaggle.com/code/yukkyo/probing-all-test-sample-have-thumbnail\" target=\"_blank\">notebook</a> to probe the LB from <a href=\"https://www.kaggle.com/yukkyo\" target=\"_blank\">@yukkyo</a>.  The basic idea is you can single bit probe a condition in the test set by having the code return one of two known states.  </p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2498381,
      "author_name": "04RR",
      "author_url": "",
      "post_date": "2023-10-25T09:21:01.197000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/masadia\" target=\"_blank\">@masadia</a>, by when can we expect an update (if there is going to be one)?</p>\n<p><a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a> <a href=\"https://www.kaggle.com/homesmac\" target=\"_blank\">@homesmac</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 2499005,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2023-10-25T16:55:27.550000",
          "content": "<p>I'm reprocessing images with masking problems now. Given how long it takes to process a 700 GB data bundle, the corrections won't be available for at least another day or two. Early next week is most likely. I'll provide a list of the updated images to make it possible to avoid downloading the entire data bundle again. </p>\n<p>The good news is that I believe we've identified the root cause and the problems were limited to a specific subset of the training data (i.e. the test set won't change at all). <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> seems to have found nearly all of them.</p>",
          "votes": 8,
          "replies": [
            {
              "id": 2499057,
              "author_name": "David Austin",
              "author_url": "",
              "post_date": "2023-10-25T17:45:42.027000",
              "content": "<p>Does that mean the train and test images are going to remain in png format?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2499079,
              "author_name": "Sohier Dane",
              "author_url": "",
              "post_date": "2023-10-25T17:57:13.460000",
              "content": "<p>Yes.                </p>",
              "votes": -2,
              "replies": []
            },
            {
              "id": 2500566,
              "author_name": "Stephan",
              "author_url": "",
              "post_date": "2023-10-26T19:20:09.940000",
              "content": "<p>Does this mean the data will stay in png format for the remainder of the challenge, or will it be converted to e.g. tif at a later date? I'd like to argue that magnification differences between the slides can throw off models considerably, considering that most networks are not scale invariant. Having them in a format where the levels are stored, along with perhaps the magnification metadata would be helpful.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2501145,
              "author_name": "jibounet",
              "author_url": "",
              "post_date": "2023-10-27T09:09:04.797000",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> If you can change the image format in favor of TIFF or SVS, please do it. Unless mistaken, one has to decode the <strong>complete</strong> PNG image at once. For large images (&gt;2GB), this decoding is really time-consuming (&gt;1min per image). I am convinced that this overhead prevents participants from making good use of full resolution images in this competition. </p>\n<p>EDIT: I've been working for several years in the field of computational pathology. IMHO, PNG is a highly unusual choice of image format for digital pathology. Image formats such as SVS are much more suited to digital pathology. If needed, <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>, I can provide some help! Do not hesitate to get in touch!</p>",
              "votes": 11,
              "replies": []
            },
            {
              "id": 2501191,
              "author_name": "jibounet",
              "author_url": "",
              "post_date": "2023-10-27T09:54:28.957000",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> For info, see below a list of previous competitions involving digital pathology. None used PNG images.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/histopathologic-cancer-detection/data\" target=\"_blank\">https://www.kaggle.com/competitions/histopathologic-cancer-detection/data</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/hubmap-kidney-segmentation/data\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-kidney-segmentation/data</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/prostate-cancer-grade-assessment/data\" target=\"_blank\">https://www.kaggle.com/competitions/prostate-cancer-grade-assessment/data</a></li>\n<li><a href=\"https://grand-challenge.org/challenges/?search=&amp;modalities=11&amp;educational=unknown&amp;status=&amp;submit=Apply+Filters\" target=\"_blank\">https://grand-challenge.org/challenges/?search=&amp;modalities=11&amp;educational=unknown&amp;status=&amp;submit=Apply+Filters</a></li>\n</ul>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2501435,
              "author_name": "David Austin",
              "author_url": "",
              "post_date": "2023-10-27T12:47:27.793000",
              "content": "<p>Also worth mentioning, it's two lines of code to do the conversion.  No good reason not to do it.</p>\n<pre><code> png  pnglist:\n    image = pyvips.Image.new_from_file(f)\n    image.tiffsave(f, =, =, =)\n</code></pre>",
              "votes": 8,
              "replies": []
            },
            {
              "id": 2501475,
              "author_name": "jibounet",
              "author_url": "",
              "post_date": "2023-10-27T13:28:20.287000",
              "content": "<p>Exactly! 💯</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2501775,
              "author_name": "Sohier Dane",
              "author_url": "",
              "post_date": "2023-10-27T16:42:18.100000",
              "content": "<p><a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> From memory, when I went that route a few months ago a small portion of the images ended up with very odd color artifacts. Think green slide backgrounds.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2501784,
              "author_name": "jibounet",
              "author_url": "",
              "post_date": "2023-10-27T16:44:33.960000",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> If you're not going to change the image format, please increase the maximum duration of notebooks to 16H or 18H.  Eventually, we - the participants - are ought to deliver the best models for ovarian cancer subtyping, for the good of the patients! I strongly believe that this will only be possible if participants can thoroughly analyze the full resolution images (which is highly unlikely given the 12H limit).</p>",
              "votes": 8,
              "replies": []
            },
            {
              "id": 2501830,
              "author_name": "David Austin",
              "author_url": "",
              "post_date": "2023-10-27T17:34:29.037000",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> If you can reference a train png image that had green artifacts after conversion to tiff, I'm pretty sure the community could help figure it out</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2503560,
              "author_name": "jibounet",
              "author_url": "",
              "post_date": "2023-10-29T07:51:52.770000",
              "content": "<p><a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a> <a href=\"https://www.kaggle.com/jonathanmcwilliams\" target=\"_blank\">@jonathanmcwilliams</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Any update/decision on this?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2486047": "It's always exciting to participate in medical image competitions, especially with the unique aspect of out-of-distribution test data we don't always see in most comps. However it's concerning to see several ongoing issues despite being brought to the attention of the Kaggle staff. I understand there's extensive preparation that goes into running these competitions, but for participants who invest significant effort it's also important know that any raised concerns are addressed promptly. Here are the critical issues that need resolution:\n\n\n1. [LB metric has an error](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/446760).  This was highlighted 5 days ago with no response from kaggle to even say they're looking into it.\n\n2. There are many quality issues with the images, as [originally highlighted ](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447063)by @yukkyo.  I've identified 46 images with clear problems, many of them looked to be flipped relative to where the mask was supposed to be selected, I'm sharing some examples here.  Not knowing any better I assume these issues are present in test as well.  We all expect a bit of noise in the data but 8% of the images like this is too much.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F888191%2F29480ebbf511c3d7f86552d1e1772565%2Fcombined_image.jpg?generation=1697560264510346&alt=media)\n\n3. PNG was a bad choice for image format and [acknowledged as such](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/446688#2479540).  If the data format is going to change it needs to happen quickly, having to re-download over 700GB of data is not a good practice.\n\nPlease let us know what's being done about these issues so we know how we want to proceed. @ashleychow @homesmac @sohier ",
    "2486488": "Hello,\n\nWe want to express our gratitude for raising this important issue.\n\nIt appears that a displacement of the masks occurred during the Kaggle production stage. Our intention was to offer masks that exclusively encompass the tissue while eliminating the background, thus simplifying the task for participants.\nWe are working with the Kaggle team to resolve this issue as soon as possible.\n\nRegards,\nMaryam",
    "2486129": "I fully agree with you. Additionally I am also concerned about the following points\n\n1. The small number of training TMAs\n\nWe are given 25 TMAs for training. However, the [Dataset Description](https://www.kaggle.com/competitions/UBC-OCEAN/data) says that majority of the tests (roughly 2000) are TMAs. Is 25 TMAs really a sufficient number? Especially in medical imaging competitions, even different data sources (e.g., hospitals) can be a challenge. With this in mind, I am wondering if the small number of TMAs given for training is an appropriate design.\n\nOf course, this is a matter of ingenuity on the part of the participants, and may be a challenge for the host. Personally, however, I would like Kaggle staff to reconsider whether this number is appropriate.",
    "2498381": "Hey @masadia, by when can we expect an update (if there is going to be one)?\n\n@ashleychow @homesmac @sohier"
  }
}