{
  "id": 451892,
  "title": "Data updated to fix flawed masks",
  "url": "/competitions/UBC-OCEAN/discussion/451892",
  "author_name": "Sohier Dane",
  "post_date": "2023-10-30T23:41:27.994000",
  "votes": 33,
  "comment_count": 31,
  "views": 0,
  "content": "<p>I began uploading a data patch for the <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447063\" target=\"_blank\">improperly masked images discussed here</a> earlier today. With any luck the patch should be available a few hours from now (4:30 PST on October 30th). </p>\n<p>49 images have been updated, all of them whole slide images in the train set. A complete list of the update files will be provided in a new file <code>updated_image_ids.json</code>. I would encourage you to download that file manually and then use it to download just the subset of the data that has been corrected. I need to drop offline soon but will be available tomorrow to respond to any questions you might have.</p>\n<p>Apologies for the inconvenience and thank you for your patience.</p>\n<p>Edit: I highly recommend taking a look <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/451892#2506858\" target=\"_blank\">the script </a><a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> provided for downloading only the relevant images using <a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">the Kaggle API</a>.</p>",
  "messages": [
    {
      "id": 2505886,
      "postDate": "2023-10-30T23:41:27.993Z",
      "content": "<p>I began uploading a data patch for the <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447063\" target=\"_blank\">improperly masked images discussed here</a> earlier today. With any luck the patch should be available a few hours from now (4:30 PST on October 30th). </p>\n<p>49 images have been updated, all of them whole slide images in the train set. A complete list of the update files will be provided in a new file <code>updated_image_ids.json</code>. I would encourage you to download that file manually and then use it to download just the subset of the data that has been corrected. I need to drop offline soon but will be available tomorrow to respond to any questions you might have.</p>\n<p>Apologies for the inconvenience and thank you for your patience.</p>\n<p>Edit: I highly recommend taking a look <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/451892#2506858\" target=\"_blank\">the script </a><a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> provided for downloading only the relevant images using <a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">the Kaggle API</a>.</p>",
      "rawMarkdown": "I began uploading a data patch for the [improperly masked images discussed here](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447063) earlier today. With any luck the patch should be available a few hours from now (4:30 PST on October 30th). \n\n49 images have been updated, all of them whole slide images in the train set. A complete list of the update files will be provided in a new file `updated_image_ids.json`. I would encourage you to download that file manually and then use it to download just the subset of the data that has been corrected. I need to drop offline soon but will be available tomorrow to respond to any questions you might have.\n\nApologies for the inconvenience and thank you for your patience.\n\nEdit: I highly recommend taking a look [the script @tivfrvqhs5 provided](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/451892#2506858) for downloading only the relevant images using [the Kaggle API](https://github.com/Kaggle/kaggle-api).",
      "votes": 33
    },
    {
      "id": 2506858,
      "postDate": "2023-10-31T15:39:18.223Z",
      "content": "<p>to download the updated images using the json file</p>\n<pre><code> json\n subprocess\n\n (, )  json_file:\n    file_list = json.load(json_file)\n\ncompetition_name = \n\n file_name  file_list:\n    command = \n    :\n        subprocess.run(command, shell=, check=)\n        ()\n     subprocess.CalledProcessError  e:\n        ()\n\n()\n</code></pre>",
      "rawMarkdown": "to download the updated images using the json file\n\n```\nimport json\nimport subprocess\n\nwith open('updated_image_ids.json', 'r') as json_file:\n    file_list = json.load(json_file)\n\ncompetition_name = 'UBC-OCEAN'\n\nfor file_name in file_list:\n    command = f'kaggle competitions download -c {competition_name} -f train_images/{file_name}.png'\n    try:\n        subprocess.run(command, shell=True, check=True)\n        print(f'Successfully downloaded {file_name}.png')\n    except subprocess.CalledProcessError as e:\n        print(f'Error downloading {file_name}.png: {e}')\n\nprint('Download process completed.')\n```",
      "votes": 11,
      "replies": [
        {
          "id": 2507316,
          "postDate": "2023-10-31T22:27:26.547Z",
          "content": "<p>just a note: If you use the thumbnails, you should update them as well… you might forget later</p>",
          "rawMarkdown": "just a note: If you use the thumbnails, you should update them as well... you might forget later"
        }
      ]
    },
    {
      "id": 2506304,
      "postDate": "2023-10-31T07:55:08.203Z",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <br>\nThank you for correcting the data deficiencies.<br>\nI have checked train_thumbnails and the mask issue seems to be resolved.</p>\n<p>However, do you have any plans to correct the sample below(32035), which seems to have the wrong aspect ratio? Or can we assume that the data is correct?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1670024%2Fdfea63e7b9b509c5e2b4bf26c6649809%2F32035_thumbnail_resized.png?generation=1698738894407653&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "@sohier \nThank you for correcting the data deficiencies.\nI have checked train_thumbnails and the mask issue seems to be resolved.\n\nHowever, do you have any plans to correct the sample below(32035), which seems to have the wrong aspect ratio? Or can we assume that the data is correct?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1670024%2Fdfea63e7b9b509c5e2b4bf26c6649809%2F32035_thumbnail_resized.png?generation=1698738894407653&alt=media)",
      "votes": 9,
      "replies": [
        {
          "id": 2506338,
          "postDate": "2023-10-31T08:22:40.503Z",
          "content": "<p>it looks weird but maybe its actually normal, because I checked train and it's the same</p>",
          "rawMarkdown": "it looks weird but maybe its actually normal, because I checked train and it's the same",
          "votes": 1
        },
        {
          "id": 2507319,
          "postDate": "2023-10-31T22:35:55.077Z",
          "content": "<p>From memory, that was a file where I had to take some unusual steps to get any image out at all. I don't expect I'll be able to dedicate time to cleaning it up further, if that would even be possible.</p>",
          "rawMarkdown": "From memory, that was a file where I had to take some unusual steps to get any image out at all. I don't expect I'll be able to dedicate time to cleaning it up further, if that would even be possible.",
          "votes": 3,
          "replies": [
            {
              "id": 2507406,
              "postDate": "2023-11-01T01:58:32.143Z",
              "content": "<p>Thank you for your response.<br>\nIt is very helpful to have confirmation that this is incorrect data and that it will be difficult to correct during the convention.</p>",
              "rawMarkdown": "Thank you for your response.\nIt is very helpful to have confirmation that this is incorrect data and that it will be difficult to correct during the convention.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2506246,
      "postDate": "2023-10-31T07:18:46.907Z",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> do you have any update regarding changing the image format to a standard one (tiff or svs) ? None of previous competitions in digital pathology, both on Kaggle and grand-challenge.org, use PNG format.</p>\n<p>Indeed, as indicated in other discussions, PNG does not allow to leverage high resolution as you have to decode the full image to access the best resolution which is not feasible given the submission time. Instead, tiff and svs allow to directly access the highest resolution on local tiles.</p>",
      "rawMarkdown": "@sohier do you have any update regarding changing the image format to a standard one (tiff or svs) ? None of previous competitions in digital pathology, both on Kaggle and grand-challenge.org, use PNG format.\n\nIndeed, as indicated in other discussions, PNG does not allow to leverage high resolution as you have to decode the full image to access the best resolution which is not feasible given the submission time. Instead, tiff and svs allow to directly access the highest resolution on local tiles.\n\n",
      "votes": 7,
      "replies": [
        {
          "id": 2506257,
          "postDate": "2023-10-31T07:26:03.077Z",
          "content": "<p>I think it was answered in this discussion <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447886\" target=\"_blank\">https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447886</a> that images will stay in PNG format, unfortunately.</p>",
          "rawMarkdown": "I think it was answered in this discussion https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447886 that images will stay in PNG format, unfortunately.",
          "votes": 1,
          "replies": [
            {
              "id": 2506279,
              "postDate": "2023-10-31T07:41:20.023Z",
              "content": "<p>Indeed, but <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> also wrote <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/446688#2479540\" target=\"_blank\">here</a> on this topic:</p>\n<blockquote>\n  <p>The current setup is definitely suboptimal but we're hoping to provide an update early next week that will improve the situation.</p>\n</blockquote>\n<p>So my understanding was that the issue was currently under review. I hope it will be fixed soon 🤞And as <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> and <a href=\"https://www.kaggle.com/jbschiratti\" target=\"_blank\">@jbschiratti</a> mentionned, the Kaggle community can help to fix it, it would be a pity for this competition to be a png engineering challenge </p>",
              "rawMarkdown": "Indeed, but @sohier also wrote [here](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/446688#2479540) on this topic:\n\n> The current setup is definitely suboptimal but we're hoping to provide an update early next week that will improve the situation.\n\nSo my understanding was that the issue was currently under review. I hope it will be fixed soon 🤞And as @tivfrvqhs5 and @jbschiratti mentionned, the Kaggle community can help to fix it, it would be a pity for this competition to be a png engineering challenge ",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2507437,
      "postDate": "2023-11-01T02:41:01.313Z",
      "content": "<p>Please update image sizes in train.csv file. <br>\nIt doesn't match with new updated wsi images. </p>",
      "rawMarkdown": "Please update image sizes in train.csv file. \nIt doesn't match with new updated wsi images. ",
      "votes": 4
    },
    {
      "id": 2509953,
      "postDate": "2023-11-02T16:44:02.910Z",
      "content": "<p>If you upload the image size, so it would be better, rest you are doing very great.</p>",
      "rawMarkdown": "If you upload the image size, so it would be better, rest you are doing very great.",
      "votes": 1
    },
    {
      "id": 2506937,
      "postDate": "2023-10-31T16:33:43.550Z",
      "content": "<p>Thank you for the update. Can you provide better clarity on what is needed to be done. Are the updated images part of the dataset already? Or, is there something we need to do on our part?</p>",
      "rawMarkdown": "Thank you for the update. Can you provide better clarity on what is needed to be done. Are the updated images part of the dataset already? Or, is there something we need to do on our part?",
      "votes": 1
    },
    {
      "id": 2509351,
      "postDate": "2023-11-02T10:33:53.860Z",
      "content": "<p>hello <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>, have you also updated the corresponding thumbnails of these ids ? I've re-ran my pre-processing pipeline and noticed that they didn't match.</p>",
      "rawMarkdown": "hello @sohier, have you also updated the corresponding thumbnails of these ids ? I've re-ran my pre-processing pipeline and noticed that they didn't match.",
      "votes": 2,
      "replies": [
        {
          "id": 2509580,
          "postDate": "2023-11-02T13:07:03.123Z",
          "content": "<p>I guess they updated them as well… I have downloaded the updated thumbnails yesterday and they match the updated images</p>",
          "rawMarkdown": "I guess they updated them as well... I have downloaded the updated thumbnails yesterday and they match the updated images"
        }
      ]
    },
    {
      "id": 2506310,
      "postDate": "2023-10-31T08:01:24.813Z",
      "content": "<p>I have download updated_image_ids.json manually,  how to download just the subset of the data? I cannot download  image manually at the Data page by name  because it doesn't show all images below the online Data Explorer.</p>",
      "rawMarkdown": "I have download updated_image_ids.json manually,  how to download just the subset of the data? I cannot download  image manually at the Data page by name  because it doesn't show all images below the online Data Explorer.",
      "votes": 2,
      "replies": [
        {
          "id": 2506365,
          "postDate": "2023-10-31T08:53:59.337Z",
          "content": "<p>under explorer, <strong>click on the train folder</strong>, you see a screen full of images on the left, scroll down and click show more, you have to download one by one(I think)</p>",
          "rawMarkdown": "under explorer, **click on the train folder**, you see a screen full of images on the left, scroll down and click show more, you have to download one by one(I think)",
          "votes": 1,
          "replies": [
            {
              "id": 2506417,
              "postDate": "2023-10-31T09:49:40.090Z",
              "content": "<p>It's weird that I cannot see full images, and I cannot scroll.</p>",
              "rawMarkdown": "It's weird that I cannot see full images, and I cannot scroll.",
              "votes": 1
            },
            {
              "id": 2506531,
              "postDate": "2023-10-31T11:10:59.893Z",
              "content": "<p>I think it puts too much pressure on the servers, I can't check them now either</p>",
              "rawMarkdown": "I think it puts too much pressure on the servers, I can't check them now either",
              "votes": 1
            }
          ]
        },
        {
          "id": 2506401,
          "postDate": "2023-10-31T09:38:18.667Z",
          "content": "<p>You may download the subset of the data with kaggle api, for example:<br>\n<code>kaggle competitions download -c UBC-OCEAN -f train_images/2906.png</code></p>",
          "rawMarkdown": "You may download the subset of the data with kaggle api, for example:\n`kaggle competitions download -c UBC-OCEAN -f train_images/2906.png`",
          "votes": 6,
          "replies": [
            {
              "id": 2506419,
              "postDate": "2023-10-31T09:50:24.640Z",
              "content": "<p>Thx! help a lot</p>",
              "rawMarkdown": "Thx! help a lot"
            }
          ]
        }
      ]
    },
    {
      "id": 2542789,
      "postDate": "2023-11-29T14:13:04.950Z",
      "content": "<p>There are files ([44232, 50962, 42549, 36063] - id) that are not in train.csv. Where can I get these label information?</p>",
      "rawMarkdown": "There are files ([44232, 50962, 42549, 36063] - id) that are not in train.csv. Where can I get these label information?",
      "replies": [
        {
          "id": 2542980,
          "postDate": "2023-11-29T17:18:54.473Z",
          "content": "<p>well, they are in train.csv </p>",
          "rawMarkdown": "well, they are in train.csv ",
          "votes": 2,
          "replies": [
            {
              "id": 2543467,
              "postDate": "2023-11-30T05:49:35.507Z",
              "content": "<p>I was mistaken. It was an ID that was deleted from train.csv in the past because the resource was bad. Thank you for answer!</p>",
              "rawMarkdown": "I was mistaken. It was an ID that was deleted from train.csv in the past because the resource was bad. Thank you for answer!",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2540923,
      "postDate": "2023-11-28T05:25:59.947Z",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Please pin this post</p>",
      "rawMarkdown": "@sohier Please pin this post"
    },
    {
      "id": 2540921,
      "postDate": "2023-11-28T05:22:51.820Z",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>, can you pin this discussion for late comers please.</p>",
      "rawMarkdown": "@sohier, can you pin this discussion for late comers please."
    },
    {
      "id": 2511122,
      "postDate": "2023-11-03T13:44:56.137Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>,</p>\n<p>It looks like something is not quite right with the thumbnails for the 49 updated images: they don't have the same dimensions in pixels as before. Specifically, they don't have the same <strong>height</strong> in pixels as before. For example, the thumbnail for image_id 44530 was measuring 2777 x 3000 (h, w) before the update; it now measures as 2247 x 3000. Likewise for image_id 34822: it used to be 2038 x 3000; it is now 1677 x 3000. It appears that you have used different scale for x (width) and y (height) when resizing the updated WSis to thumbnails. This makes using those updated thumbnails useless in my pipeline as they don't conform to the (correct) proportionally resized thumbnails for the rest of the dataset.</p>\n<p>When can we expect this to be fixed?</p>",
      "rawMarkdown": "Hi @sohier,\n\nIt looks like something is not quite right with the thumbnails for the 49 updated images: they don't have the same dimensions in pixels as before. Specifically, they don't have the same **height** in pixels as before. For example, the thumbnail for image_id 44530 was measuring 2777 x 3000 (h, w) before the update; it now measures as 2247 x 3000. Likewise for image_id 34822: it used to be 2038 x 3000; it is now 1677 x 3000. It appears that you have used different scale for x (width) and y (height) when resizing the updated WSis to thumbnails. This makes using those updated thumbnails useless in my pipeline as they don't conform to the (correct) proportionally resized thumbnails for the rest of the dataset.\n\nWhen can we expect this to be fixed?",
      "replies": [
        {
          "id": 2511139,
          "postDate": "2023-11-03T13:52:27.670Z",
          "content": "<p>I don't think it's a mistake at all. The thumbnails are always resized to have the longer side at 3000x, while maintaining the aspect ratios. So, if the WSI images have changed in size, the size of the thumbnails will logically change as well.</p>",
          "rawMarkdown": "I don't think it's a mistake at all. The thumbnails are always resized to have the longer side at 3000x, while maintaining the aspect ratios. So, if the WSI images have changed in size, the size of the thumbnails will logically change as well.",
          "votes": 3,
          "replies": [
            {
              "id": 2511306,
              "postDate": "2023-11-03T15:15:21.883Z",
              "content": "<p>Then, the size values in the train.csv file need to be updated, as they still reflect the \"old\" values. </p>",
              "rawMarkdown": "Then, the size values in the train.csv file need to be updated, as they still reflect the \"old\" values. ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2506128,
      "postDate": "2023-10-31T05:49:00.713Z",
      "content": "<p>I'm looking at this notebook:<br>\n<a href=\"https://www.kaggle.com/code/gunesevitan/ubc-ocean-eda\" target=\"_blank\">https://www.kaggle.com/code/gunesevitan/ubc-ocean-eda</a><br>\nif you decided to do something about duplicates:</p>\n<p>duplicates:<br>\n281,706,1252,1295,1660,1666,1943,2391,2706,3055,3092,3098,3264,3672,4827,5251,5264,5307,5852,<br>\n5992,6175,6449,6793,6843,7955,8130,8985,9341,10548,10642,11559,12222,12442,13526,13987,14039,<br>\n14312,14542,15221,15293,15470,15486,15912,16042,16494,17067,17416,18014,18196,18810,18896,18914,<br>\n19157,19512,20670,20858,21232,21260,21910,22290,22425,22654,22924,23523,24023,24759,25256,25561,<br>\n25792,26190,26603,26644,26862,27851,27950,28121,28519,28562,28603,29240,29331,29915,30508,30515,<br>\n30539,30738,30792,30868,31333,31473,32112,32596,32636,34277,34845,35592,35652,35792,35953,36008,<br>\n36204,37190,38041,38097,38118,38585,38687,38959,39144,39208,39297,39365,40129,40503,41361,41801,<br>\n42296,43280,43796,43875,43998,44804,44962,44976,45104,45254,45578,45725,45990,46543,46736,<br>\n46769,46793,47020,47911,47984,48502,48861,48973,50048,50246,50712,51021,51346,51832,52259,52375,<br>\n52420,52461,52612,52752,52931,53059,53900,54473,54506,54825,54990,55279,55281,56500,56843,57100,<br>\n58974,59900,60287,60685,60928,61033,61100,61689,62828,63367,63429,64188,64824,64950,65300</p>\n<p>(I removed the ones you said were fixed, I haven't checked them out yet though)</p>\n<p>has marker on it:<br>\n36513(also duplicate)<br>\nthis seems a bit weird, Im not sure:<br>\n32192,39252,39258,39872,47431</p>\n<p>this notebook does help with duplicates:<br>\n<a href=\"https://www.kaggle.com/code/dhinkris/crop-duplicate-wsi-images/notebook\" target=\"_blank\">https://www.kaggle.com/code/dhinkris/crop-duplicate-wsi-images/notebook</a></p>\n<p>did thumbnails also change?<br>\ntest set has these same issues?</p>",
      "rawMarkdown": "I'm looking at this notebook:\nhttps://www.kaggle.com/code/gunesevitan/ubc-ocean-eda\nif you decided to do something about duplicates:\n\nduplicates:\n281,706,1252,1295,1660,1666,1943,2391,2706,3055,3092,3098,3264,3672,4827,5251,5264,5307,5852,\n5992,6175,6449,6793,6843,7955,8130,8985,9341,10548,10642,11559,12222,12442,13526,13987,14039,\n14312,14542,15221,15293,15470,15486,15912,16042,16494,17067,17416,18014,18196,18810,18896,18914,\n19157,19512,20670,20858,21232,21260,21910,22290,22425,22654,22924,23523,24023,24759,25256,25561,\n25792,26190,26603,26644,26862,27851,27950,28121,28519,28562,28603,29240,29331,29915,30508,30515,\n30539,30738,30792,30868,31333,31473,32112,32596,32636,34277,34845,35592,35652,35792,35953,36008,\n36204,37190,38041,38097,38118,38585,38687,38959,39144,39208,39297,39365,40129,40503,41361,41801,\n42296,43280,43796,43875,43998,44804,44962,44976,45104,45254,45578,45725,45990,46543,46736,\n46769,46793,47020,47911,47984,48502,48861,48973,50048,50246,50712,51021,51346,51832,52259,52375,\n52420,52461,52612,52752,52931,53059,53900,54473,54506,54825,54990,55279,55281,56500,56843,57100,\n58974,59900,60287,60685,60928,61033,61100,61689,62828,63367,63429,64188,64824,64950,65300\n\n(I removed the ones you said were fixed, I haven't checked them out yet though)\n\nhas marker on it:\n36513(also duplicate)\nthis seems a bit weird, Im not sure:\n32192,39252,39258,39872,47431\n\nthis notebook does help with duplicates:\nhttps://www.kaggle.com/code/dhinkris/crop-duplicate-wsi-images/notebook\n\ndid thumbnails also change?\ntest set has these same issues?"
    },
    {
      "id": 2508816,
      "postDate": "2023-11-02T03:03:20.927Z",
      "content": "<p>Great thank you</p>",
      "rawMarkdown": "Great thank you"
    },
    {
      "id": 2506114,
      "postDate": "2023-10-31T05:36:03.523Z",
      "content": "<p>thank you!</p>",
      "rawMarkdown": "thank you!"
    }
  ],
  "comments": [
    {
      "id": 2506858,
      "author_name": "David Austin",
      "author_url": "",
      "post_date": "2023-10-31T15:39:18.223000",
      "content": "<p>to download the updated images using the json file</p>\n<pre><code> json\n subprocess\n\n (, )  json_file:\n    file_list = json.load(json_file)\n\ncompetition_name = \n\n file_name  file_list:\n    command = \n    :\n        subprocess.run(command, shell=, check=)\n        ()\n     subprocess.CalledProcessError  e:\n        ()\n\n()\n</code></pre>",
      "votes": 11,
      "replies": [
        {
          "id": 2507316,
          "author_name": "Patchef",
          "author_url": "",
          "post_date": "2023-10-31T22:27:26.547000",
          "content": "<p>just a note: If you use the thumbnails, you should update them as well… you might forget later</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2506304,
      "author_name": "fam_taro",
      "author_url": "",
      "post_date": "2023-10-31T07:55:08.203000",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <br>\nThank you for correcting the data deficiencies.<br>\nI have checked train_thumbnails and the mask issue seems to be resolved.</p>\n<p>However, do you have any plans to correct the sample below(32035), which seems to have the wrong aspect ratio? Or can we assume that the data is correct?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1670024%2Fdfea63e7b9b509c5e2b4bf26c6649809%2F32035_thumbnail_resized.png?generation=1698738894407653&amp;alt=media\" alt=\"\"></p>",
      "votes": 9,
      "replies": [
        {
          "id": 2506338,
          "author_name": "Parham Gousheh",
          "author_url": "",
          "post_date": "2023-10-31T08:22:40.503000",
          "content": "<p>it looks weird but maybe its actually normal, because I checked train and it's the same</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2507319,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2023-10-31T22:35:55.077000",
          "content": "<p>From memory, that was a file where I had to take some unusual steps to get any image out at all. I don't expect I'll be able to dedicate time to cleaning it up further, if that would even be possible.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2507406,
              "author_name": "fam_taro",
              "author_url": "",
              "post_date": "2023-11-01T01:58:32.143000",
              "content": "<p>Thank you for your response.<br>\nIt is very helpful to have confirmation that this is incorrect data and that it will be difficult to correct during the convention.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2506246,
      "author_name": "simjeg",
      "author_url": "",
      "post_date": "2023-10-31T07:18:46.907000",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> do you have any update regarding changing the image format to a standard one (tiff or svs) ? None of previous competitions in digital pathology, both on Kaggle and grand-challenge.org, use PNG format.</p>\n<p>Indeed, as indicated in other discussions, PNG does not allow to leverage high resolution as you have to decode the full image to access the best resolution which is not feasible given the submission time. Instead, tiff and svs allow to directly access the highest resolution on local tiles.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2506257,
          "author_name": "hyc",
          "author_url": "",
          "post_date": "2023-10-31T07:26:03.077000",
          "content": "<p>I think it was answered in this discussion <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447886\" target=\"_blank\">https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447886</a> that images will stay in PNG format, unfortunately.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2506279,
              "author_name": "simjeg",
              "author_url": "",
              "post_date": "2023-10-31T07:41:20.023000",
              "content": "<p>Indeed, but <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> also wrote <a href=\"https://www.kaggle.com/competitions/UBC-OCEAN/discussion/446688#2479540\" target=\"_blank\">here</a> on this topic:</p>\n<blockquote>\n  <p>The current setup is definitely suboptimal but we're hoping to provide an update early next week that will improve the situation.</p>\n</blockquote>\n<p>So my understanding was that the issue was currently under review. I hope it will be fixed soon 🤞And as <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> and <a href=\"https://www.kaggle.com/jbschiratti\" target=\"_blank\">@jbschiratti</a> mentionned, the Kaggle community can help to fix it, it would be a pity for this competition to be a png engineering challenge </p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2507437,
      "author_name": "Longyi Kim",
      "author_url": "",
      "post_date": "2023-11-01T02:41:01.313000",
      "content": "<p>Please update image sizes in train.csv file. <br>\nIt doesn't match with new updated wsi images. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2509953,
      "author_name": "VIKRAM MISHRA",
      "author_url": "",
      "post_date": "2023-11-02T16:44:02.910000",
      "content": "<p>If you upload the image size, so it would be better, rest you are doing very great.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2506937,
      "author_name": "William Green",
      "author_url": "",
      "post_date": "2023-10-31T16:33:43.550000",
      "content": "<p>Thank you for the update. Can you provide better clarity on what is needed to be done. Are the updated images part of the dataset already? Or, is there something we need to do on our part?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2509351,
      "author_name": "NguyenThanhNhan",
      "author_url": "",
      "post_date": "2023-11-02T10:33:53.860000",
      "content": "<p>hello <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>, have you also updated the corresponding thumbnails of these ids ? I've re-ran my pre-processing pipeline and noticed that they didn't match.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2509580,
          "author_name": "Patchef",
          "author_url": "",
          "post_date": "2023-11-02T13:07:03.123000",
          "content": "<p>I guess they updated them as well… I have downloaded the updated thumbnails yesterday and they match the updated images</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2506310,
      "author_name": "Mr.Fire",
      "author_url": "",
      "post_date": "2023-10-31T08:01:24.813000",
      "content": "<p>I have download updated_image_ids.json manually,  how to download just the subset of the data? I cannot download  image manually at the Data page by name  because it doesn't show all images below the online Data Explorer.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2506365,
          "author_name": "Parham Gousheh",
          "author_url": "",
          "post_date": "2023-10-31T08:53:59.337000",
          "content": "<p>under explorer, <strong>click on the train folder</strong>, you see a screen full of images on the left, scroll down and click show more, you have to download one by one(I think)</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2506417,
              "author_name": "Mr.Fire",
              "author_url": "",
              "post_date": "2023-10-31T09:49:40.090000",
              "content": "<p>It's weird that I cannot see full images, and I cannot scroll.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2506531,
              "author_name": "Parham Gousheh",
              "author_url": "",
              "post_date": "2023-10-31T11:10:59.893000",
              "content": "<p>I think it puts too much pressure on the servers, I can't check them now either</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2506401,
          "author_name": "sheep",
          "author_url": "",
          "post_date": "2023-10-31T09:38:18.667000",
          "content": "<p>You may download the subset of the data with kaggle api, for example:<br>\n<code>kaggle competitions download -c UBC-OCEAN -f train_images/2906.png</code></p>",
          "votes": 6,
          "replies": [
            {
              "id": 2506419,
              "author_name": "Mr.Fire",
              "author_url": "",
              "post_date": "2023-10-31T09:50:24.640000",
              "content": "<p>Thx! help a lot</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2542789,
      "author_name": "Trappist",
      "author_url": "",
      "post_date": "2023-11-29T14:13:04.950000",
      "content": "<p>There are files ([44232, 50962, 42549, 36063] - id) that are not in train.csv. Where can I get these label information?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2542980,
          "author_name": "Patchef",
          "author_url": "",
          "post_date": "2023-11-29T17:18:54.473000",
          "content": "<p>well, they are in train.csv </p>",
          "votes": 2,
          "replies": [
            {
              "id": 2543467,
              "author_name": "Trappist",
              "author_url": "",
              "post_date": "2023-11-30T05:49:35.507000",
              "content": "<p>I was mistaken. It was an ID that was deleted from train.csv in the past because the resource was bad. Thank you for answer!</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2540923,
      "author_name": "KarthiAru",
      "author_url": "",
      "post_date": "2023-11-28T05:25:59.947000",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Please pin this post</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2540921,
      "author_name": "GUNER",
      "author_url": "",
      "post_date": "2023-11-28T05:22:51.820000",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>, can you pin this discussion for late comers please.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2511122,
      "author_name": "Patrick Robitaille",
      "author_url": "",
      "post_date": "2023-11-03T13:44:56.137000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>,</p>\n<p>It looks like something is not quite right with the thumbnails for the 49 updated images: they don't have the same dimensions in pixels as before. Specifically, they don't have the same <strong>height</strong> in pixels as before. For example, the thumbnail for image_id 44530 was measuring 2777 x 3000 (h, w) before the update; it now measures as 2247 x 3000. Likewise for image_id 34822: it used to be 2038 x 3000; it is now 1677 x 3000. It appears that you have used different scale for x (width) and y (height) when resizing the updated WSis to thumbnails. This makes using those updated thumbnails useless in my pipeline as they don't conform to the (correct) proportionally resized thumbnails for the rest of the dataset.</p>\n<p>When can we expect this to be fixed?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2511139,
          "author_name": "Patchef",
          "author_url": "",
          "post_date": "2023-11-03T13:52:27.670000",
          "content": "<p>I don't think it's a mistake at all. The thumbnails are always resized to have the longer side at 3000x, while maintaining the aspect ratios. So, if the WSI images have changed in size, the size of the thumbnails will logically change as well.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2511306,
              "author_name": "Patrick Robitaille",
              "author_url": "",
              "post_date": "2023-11-03T15:15:21.883000",
              "content": "<p>Then, the size values in the train.csv file need to be updated, as they still reflect the \"old\" values. </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2506128,
      "author_name": "Parham Gousheh",
      "author_url": "",
      "post_date": "2023-10-31T05:49:00.713000",
      "content": "<p>I'm looking at this notebook:<br>\n<a href=\"https://www.kaggle.com/code/gunesevitan/ubc-ocean-eda\" target=\"_blank\">https://www.kaggle.com/code/gunesevitan/ubc-ocean-eda</a><br>\nif you decided to do something about duplicates:</p>\n<p>duplicates:<br>\n281,706,1252,1295,1660,1666,1943,2391,2706,3055,3092,3098,3264,3672,4827,5251,5264,5307,5852,<br>\n5992,6175,6449,6793,6843,7955,8130,8985,9341,10548,10642,11559,12222,12442,13526,13987,14039,<br>\n14312,14542,15221,15293,15470,15486,15912,16042,16494,17067,17416,18014,18196,18810,18896,18914,<br>\n19157,19512,20670,20858,21232,21260,21910,22290,22425,22654,22924,23523,24023,24759,25256,25561,<br>\n25792,26190,26603,26644,26862,27851,27950,28121,28519,28562,28603,29240,29331,29915,30508,30515,<br>\n30539,30738,30792,30868,31333,31473,32112,32596,32636,34277,34845,35592,35652,35792,35953,36008,<br>\n36204,37190,38041,38097,38118,38585,38687,38959,39144,39208,39297,39365,40129,40503,41361,41801,<br>\n42296,43280,43796,43875,43998,44804,44962,44976,45104,45254,45578,45725,45990,46543,46736,<br>\n46769,46793,47020,47911,47984,48502,48861,48973,50048,50246,50712,51021,51346,51832,52259,52375,<br>\n52420,52461,52612,52752,52931,53059,53900,54473,54506,54825,54990,55279,55281,56500,56843,57100,<br>\n58974,59900,60287,60685,60928,61033,61100,61689,62828,63367,63429,64188,64824,64950,65300</p>\n<p>(I removed the ones you said were fixed, I haven't checked them out yet though)</p>\n<p>has marker on it:<br>\n36513(also duplicate)<br>\nthis seems a bit weird, Im not sure:<br>\n32192,39252,39258,39872,47431</p>\n<p>this notebook does help with duplicates:<br>\n<a href=\"https://www.kaggle.com/code/dhinkris/crop-duplicate-wsi-images/notebook\" target=\"_blank\">https://www.kaggle.com/code/dhinkris/crop-duplicate-wsi-images/notebook</a></p>\n<p>did thumbnails also change?<br>\ntest set has these same issues?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2508816,
      "author_name": "Pablo Florindo",
      "author_url": "",
      "post_date": "2023-11-02T03:03:20.927000",
      "content": "<p>Great thank you</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2506114,
      "author_name": "04RR",
      "author_url": "",
      "post_date": "2023-10-31T05:36:03.523000",
      "content": "<p>thank you!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2505886": "I began uploading a data patch for the [improperly masked images discussed here](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/447063) earlier today. With any luck the patch should be available a few hours from now (4:30 PST on October 30th). \n\n49 images have been updated, all of them whole slide images in the train set. A complete list of the update files will be provided in a new file `updated_image_ids.json`. I would encourage you to download that file manually and then use it to download just the subset of the data that has been corrected. I need to drop offline soon but will be available tomorrow to respond to any questions you might have.\n\nApologies for the inconvenience and thank you for your patience.\n\nEdit: I highly recommend taking a look [the script @tivfrvqhs5 provided](https://www.kaggle.com/competitions/UBC-OCEAN/discussion/451892#2506858) for downloading only the relevant images using [the Kaggle API](https://github.com/Kaggle/kaggle-api).",
    "2506858": "to download the updated images using the json file\n\n```\nimport json\nimport subprocess\n\nwith open('updated_image_ids.json', 'r') as json_file:\n    file_list = json.load(json_file)\n\ncompetition_name = 'UBC-OCEAN'\n\nfor file_name in file_list:\n    command = f'kaggle competitions download -c {competition_name} -f train_images/{file_name}.png'\n    try:\n        subprocess.run(command, shell=True, check=True)\n        print(f'Successfully downloaded {file_name}.png')\n    except subprocess.CalledProcessError as e:\n        print(f'Error downloading {file_name}.png: {e}')\n\nprint('Download process completed.')\n```",
    "2506304": "@sohier \nThank you for correcting the data deficiencies.\nI have checked train_thumbnails and the mask issue seems to be resolved.\n\nHowever, do you have any plans to correct the sample below(32035), which seems to have the wrong aspect ratio? Or can we assume that the data is correct?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1670024%2Fdfea63e7b9b509c5e2b4bf26c6649809%2F32035_thumbnail_resized.png?generation=1698738894407653&alt=media)",
    "2506246": "@sohier do you have any update regarding changing the image format to a standard one (tiff or svs) ? None of previous competitions in digital pathology, both on Kaggle and grand-challenge.org, use PNG format.\n\nIndeed, as indicated in other discussions, PNG does not allow to leverage high resolution as you have to decode the full image to access the best resolution which is not feasible given the submission time. Instead, tiff and svs allow to directly access the highest resolution on local tiles.\n\n",
    "2507437": "Please update image sizes in train.csv file. \nIt doesn't match with new updated wsi images. ",
    "2509953": "If you upload the image size, so it would be better, rest you are doing very great.",
    "2506937": "Thank you for the update. Can you provide better clarity on what is needed to be done. Are the updated images part of the dataset already? Or, is there something we need to do on our part?",
    "2509351": "hello @sohier, have you also updated the corresponding thumbnails of these ids ? I've re-ran my pre-processing pipeline and noticed that they didn't match.",
    "2506310": "I have download updated_image_ids.json manually,  how to download just the subset of the data? I cannot download  image manually at the Data page by name  because it doesn't show all images below the online Data Explorer.",
    "2542789": "There are files ([44232, 50962, 42549, 36063] - id) that are not in train.csv. Where can I get these label information?",
    "2540923": "@sohier Please pin this post",
    "2540921": "@sohier, can you pin this discussion for late comers please.",
    "2511122": "Hi @sohier,\n\nIt looks like something is not quite right with the thumbnails for the 49 updated images: they don't have the same dimensions in pixels as before. Specifically, they don't have the same **height** in pixels as before. For example, the thumbnail for image_id 44530 was measuring 2777 x 3000 (h, w) before the update; it now measures as 2247 x 3000. Likewise for image_id 34822: it used to be 2038 x 3000; it is now 1677 x 3000. It appears that you have used different scale for x (width) and y (height) when resizing the updated WSis to thumbnails. This makes using those updated thumbnails useless in my pipeline as they don't conform to the (correct) proportionally resized thumbnails for the rest of the dataset.\n\nWhen can we expect this to be fixed?",
    "2506128": "I'm looking at this notebook:\nhttps://www.kaggle.com/code/gunesevitan/ubc-ocean-eda\nif you decided to do something about duplicates:\n\nduplicates:\n281,706,1252,1295,1660,1666,1943,2391,2706,3055,3092,3098,3264,3672,4827,5251,5264,5307,5852,\n5992,6175,6449,6793,6843,7955,8130,8985,9341,10548,10642,11559,12222,12442,13526,13987,14039,\n14312,14542,15221,15293,15470,15486,15912,16042,16494,17067,17416,18014,18196,18810,18896,18914,\n19157,19512,20670,20858,21232,21260,21910,22290,22425,22654,22924,23523,24023,24759,25256,25561,\n25792,26190,26603,26644,26862,27851,27950,28121,28519,28562,28603,29240,29331,29915,30508,30515,\n30539,30738,30792,30868,31333,31473,32112,32596,32636,34277,34845,35592,35652,35792,35953,36008,\n36204,37190,38041,38097,38118,38585,38687,38959,39144,39208,39297,39365,40129,40503,41361,41801,\n42296,43280,43796,43875,43998,44804,44962,44976,45104,45254,45578,45725,45990,46543,46736,\n46769,46793,47020,47911,47984,48502,48861,48973,50048,50246,50712,51021,51346,51832,52259,52375,\n52420,52461,52612,52752,52931,53059,53900,54473,54506,54825,54990,55279,55281,56500,56843,57100,\n58974,59900,60287,60685,60928,61033,61100,61689,62828,63367,63429,64188,64824,64950,65300\n\n(I removed the ones you said were fixed, I haven't checked them out yet though)\n\nhas marker on it:\n36513(also duplicate)\nthis seems a bit weird, Im not sure:\n32192,39252,39258,39872,47431\n\nthis notebook does help with duplicates:\nhttps://www.kaggle.com/code/dhinkris/crop-duplicate-wsi-images/notebook\n\ndid thumbnails also change?\ntest set has these same issues?",
    "2508816": "Great thank you",
    "2506114": "thank you!"
  }
}