{
  "id": 446104,
  "title": "Tiled Images dataset (256x256)",
  "url": "/competitions/UBC-OCEAN/discussion/446104",
  "author_name": "pjmathematician",
  "post_date": "2023-10-10T09:36:55.234000",
  "votes": 85,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Hello! Based on this amazing notebook : <a href=\"https://www.kaggle.com/code/dheerajmpai/ucb-ocean-process-large-images-golang-50x-fast\" target=\"_blank\">https://www.kaggle.com/code/dheerajmpai/ucb-ocean-process-large-images-golang-50x-fast</a> <br>\nI have resized and tiled the images, with removal of <em>almost black</em> images, using the following notebooks:</p>\n<p><a href=\"https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker\" target=\"_blank\">https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker</a><br>\n<a href=\"https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p2\" target=\"_blank\">https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p2</a><br>\n<a href=\"https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p3\" target=\"_blank\">https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p3</a><br>\n<a href=\"https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p4\" target=\"_blank\">https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p4</a><br>\n<a href=\"https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p5\" target=\"_blank\">https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p5</a></p>\n<p>The datasets can be found here:<br>\n<a href=\"https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-1\" target=\"_blank\">https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-1</a><br>\n<a href=\"https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-2\" target=\"_blank\">https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-2</a><br>\n<a href=\"https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-3\" target=\"_blank\">https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-3</a><br>\n<a href=\"https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-4\" target=\"_blank\">https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-4</a><br>\n<a href=\"https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-5\" target=\"_blank\">https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-5</a></p>\n<p>An example notebook to load the datasets is:<br>\n<a href=\"https://www.kaggle.com/pjmathematician/ucbo-256-tiles-loading\" target=\"_blank\">https://www.kaggle.com/pjmathematician/ucbo-256-tiles-loading</a></p>\n<p>I will create 512x512 datasets as well very soon!!<br>\nThank you for this amazing competition, looking forward to participating in it!</p>",
  "messages": [
    {
      "id": 2476032,
      "postDate": "2023-10-10T09:36:55.233Z",
      "content": "<p>Hello! Based on this amazing notebook : <a href=\"https://www.kaggle.com/code/dheerajmpai/ucb-ocean-process-large-images-golang-50x-fast\" target=\"_blank\">https://www.kaggle.com/code/dheerajmpai/ucb-ocean-process-large-images-golang-50x-fast</a> <br>\nI have resized and tiled the images, with removal of <em>almost black</em> images, using the following notebooks:</p>\n<p><a href=\"https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker\" target=\"_blank\">https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker</a><br>\n<a href=\"https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p2\" target=\"_blank\">https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p2</a><br>\n<a href=\"https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p3\" target=\"_blank\">https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p3</a><br>\n<a href=\"https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p4\" target=\"_blank\">https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p4</a><br>\n<a href=\"https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p5\" target=\"_blank\">https://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p5</a></p>\n<p>The datasets can be found here:<br>\n<a href=\"https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-1\" target=\"_blank\">https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-1</a><br>\n<a href=\"https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-2\" target=\"_blank\">https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-2</a><br>\n<a href=\"https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-3\" target=\"_blank\">https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-3</a><br>\n<a href=\"https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-4\" target=\"_blank\">https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-4</a><br>\n<a href=\"https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-5\" target=\"_blank\">https://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-5</a></p>\n<p>An example notebook to load the datasets is:<br>\n<a href=\"https://www.kaggle.com/pjmathematician/ucbo-256-tiles-loading\" target=\"_blank\">https://www.kaggle.com/pjmathematician/ucbo-256-tiles-loading</a></p>\n<p>I will create 512x512 datasets as well very soon!!<br>\nThank you for this amazing competition, looking forward to participating in it!</p>",
      "rawMarkdown": "Hello! Based on this amazing notebook : https://www.kaggle.com/code/dheerajmpai/ucb-ocean-process-large-images-golang-50x-fast \nI have resized and tiled the images, with removal of *almost black* images, using the following notebooks:\n\nhttps://www.kaggle.com/code/pjmathematician/ucbo-tilemaker\nhttps://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p2\nhttps://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p3\nhttps://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p4\nhttps://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p5\n\nThe datasets can be found here:\nhttps://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-1\nhttps://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-2\nhttps://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-3\nhttps://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-4\nhttps://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-5\n\nAn example notebook to load the datasets is:\nhttps://www.kaggle.com/pjmathematician/ucbo-256-tiles-loading\n\nI will create 512x512 datasets as well very soon!!\nThank you for this amazing competition, looking forward to participating in it!\n\n",
      "votes": 84
    },
    {
      "id": 2478643,
      "postDate": "2023-10-12T06:04:39.883Z",
      "content": "<p>This is very helpful. Thank you!<br>\nBut there may be a little issue, upon browsing the training set thumbnails and full-size images, there is a lot of data cleaning that needs to be done. There are areas with colored markers, processing artifacts like scratches, holes, folds, dust, uneven staining etc., blurry areas, older overly pink slides, tumors located in other parts of the body instead of the ovary, poorly aligned annotations (some are flipped or inverted or translated horizontally), some have a single tiny tile with the tumor (the rest is normal tissue), and one slide is vertically stretched. Also, many slides have 2 or more slices of the same tumor, that are essentially the same. It would be a truly laborious task to clean it up so that these tiles can be optimal.</p>",
      "rawMarkdown": "This is very helpful. Thank you!\nBut there may be a little issue, upon browsing the training set thumbnails and full-size images, there is a lot of data cleaning that needs to be done. There are areas with colored markers, processing artifacts like scratches, holes, folds, dust, uneven staining etc., blurry areas, older overly pink slides, tumors located in other parts of the body instead of the ovary, poorly aligned annotations (some are flipped or inverted or translated horizontally), some have a single tiny tile with the tumor (the rest is normal tissue), and one slide is vertically stretched. Also, many slides have 2 or more slices of the same tumor, that are essentially the same. It would be a truly laborious task to clean it up so that these tiles can be optimal.",
      "votes": 10,
      "replies": [
        {
          "id": 2478699,
          "postDate": "2023-10-12T06:53:19.073Z",
          "content": "<p>good point. Currently what i am doing is, I selected the middle K tiles from every image for training. And changing the position/K according to the CV score. Using every tile would be very inefficient as there are a total of ~500k tiles.</p>",
          "rawMarkdown": "good point. Currently what i am doing is, I selected the middle K tiles from every image for training. And changing the position/K according to the CV score. Using every tile would be very inefficient as there are a total of ~500k tiles.",
          "votes": 10,
          "replies": [
            {
              "id": 2554266,
              "postDate": "2023-12-09T01:38:39.317Z",
              "content": "<p>Ah! good idea. Thanks for sharing!</p>",
              "rawMarkdown": "Ah! good idea. Thanks for sharing!"
            }
          ]
        }
      ]
    },
    {
      "id": 2499693,
      "postDate": "2023-10-26T07:23:08.463Z",
      "content": "<p>I have here also 512*512 with scale 0.25<br>\n<a href=\"https://www.kaggle.com/datasets/jirkaborovec/tiles-of-cancer-2048px-scale-0-25\" target=\"_blank\">https://www.kaggle.com/datasets/jirkaborovec/tiles-of-cancer-2048px-scale-0-25</a></p>",
      "rawMarkdown": "I have here also 512*512 with scale 0.25\nhttps://www.kaggle.com/datasets/jirkaborovec/tiles-of-cancer-2048px-scale-0-25",
      "votes": 4
    },
    {
      "id": 2491605,
      "postDate": "2023-10-21T19:06:32.057Z",
      "content": "<p>Thanks for the datasets! </p>\n<p>I had to convert everything myself. You know it's gonna be a fun competition when you have to buy another 4TB disk to play around with a single dataset</p>",
      "rawMarkdown": "Thanks for the datasets! \n\nI had to convert everything myself. You know it's gonna be a fun competition when you have to buy another 4TB disk to play around with a single dataset",
      "votes": 4
    },
    {
      "id": 2570054,
      "postDate": "2023-12-21T20:36:48.097Z",
      "content": "<p>Hello, thank you for your efforts. I'm interested in understanding whether it's necessary to implement tile processing on the test set as well.</p>",
      "rawMarkdown": "\nHello, thank you for your efforts. I'm interested in understanding whether it's necessary to implement tile processing on the test set as well.",
      "votes": 1
    },
    {
      "id": 2491559,
      "postDate": "2023-10-21T17:52:28.403Z",
      "content": "<p>Good work . Keep it up</p>",
      "rawMarkdown": "Good work . Keep it up",
      "votes": 2
    },
    {
      "id": 2665170,
      "postDate": "2024-02-23T13:35:46.167Z",
      "content": "<p><a href=\"https://www.kaggle.com/pjmathematician\" target=\"_blank\">@pjmathematician</a> The datasets you provided have been incredibly helpful. Currently, I'm using this code to generate my patches. Could you please explain how you converted the outputs into datasets? I'm encountering an error when trying to create a dataset directly from the output. Any guidance would be appreciated!</p>",
      "rawMarkdown": "@pjmathematician The datasets you provided have been incredibly helpful. Currently, I'm using this code to generate my patches. Could you please explain how you converted the outputs into datasets? I'm encountering an error when trying to create a dataset directly from the output. Any guidance would be appreciated!",
      "replies": [
        {
          "id": 2667025,
          "postDate": "2024-02-24T20:11:56.020Z",
          "content": "<p>Hi! <a href=\"https://www.kaggle.com/usmansafdar09\" target=\"_blank\">@usmansafdar09</a> ! <br>\nIf the output of the notebooks is throwing error, I suggest you use <a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">Kaggle API</a> in your notebook to directly create the dataset once the patches are made</p>",
          "rawMarkdown": "Hi! @usmansafdar09 ! \nIf the output of the notebooks is throwing error, I suggest you use [Kaggle API](https://github.com/Kaggle/kaggle-api) in your notebook to directly create the dataset once the patches are made",
          "votes": 1
        }
      ]
    },
    {
      "id": 2515560,
      "postDate": "2023-11-07T02:50:39.450Z",
      "content": "<p>In the 256_13568 folder, the left part of the tissue seems to be missing from the extracted patches. Did you pad the images to be divisible by cropsize before creating tiles?</p>",
      "rawMarkdown": "In the 256_13568 folder, the left part of the tissue seems to be missing from the extracted patches. Did you pad the images to be divisible by cropsize before creating tiles?"
    },
    {
      "id": 2507749,
      "postDate": "2023-11-01T08:16:14.440Z",
      "content": "<p>Have you scaled the image? If you have scaled, what is the scalratio?<br>\nDid you crop it on the original image or on the thumbnail?<br>\nLooking forward to your reply!</p>",
      "rawMarkdown": "Have you scaled the image? If you have scaled, what is the scalratio?\nDid you crop it on the original image or on the thumbnail?\nLooking forward to your reply!"
    },
    {
      "id": 2503359,
      "postDate": "2023-10-29T02:52:53.400Z",
      "content": "<p>Thank you for sharing!<br>\nHow do you run the Go language during inference?</p>",
      "rawMarkdown": "Thank you for sharing!\nHow do you run the Go language during inference?"
    },
    {
      "id": 2493575,
      "postDate": "2023-10-23T13:39:47.257Z",
      "content": "<p>Did you employ the tile images as-is for training, or did you alter them to resemble a circular shape? and how many images are you considering in your training dataset, as there is a data imbalance within the classes?</p>",
      "rawMarkdown": "Did you employ the tile images as-is for training, or did you alter them to resemble a circular shape? and how many images are you considering in your training dataset, as there is a data imbalance within the classes?",
      "replies": [
        {
          "id": 2493582,
          "postDate": "2023-10-23T13:42:41.810Z",
          "content": "<p>Good question, I will release a tile based training-inference baseline submission this week. That should give a basic idea of the usage</p>",
          "rawMarkdown": "Good question, I will release a tile based training-inference baseline submission this week. That should give a basic idea of the usage",
          "votes": 1,
          "replies": [
            {
              "id": 2493589,
              "postDate": "2023-10-23T13:45:32.243Z",
              "content": "<p>Also the TMA images are captured at a 40x zoom level, but I think the tiled images might have been zoomed in slightly more. What do you think?</p>",
              "rawMarkdown": "Also the TMA images are captured at a 40x zoom level, but I think the tiled images might have been zoomed in slightly more. What do you think?",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2491951,
      "postDate": "2023-10-22T06:53:23.573Z",
      "content": "<p>How do I use these tiled images? Do you train it with the tiled image itself as input? </p>",
      "rawMarkdown": "How do I use these tiled images? Do you train it with the tiled image itself as input? "
    },
    {
      "id": 2487225,
      "postDate": "2023-10-18T13:01:36.243Z",
      "content": "<p>Thought I didn't know how to use the dataset in my submission correctly,but it really work in my model training.Good job.</p>",
      "rawMarkdown": "Thought I didn't know how to use the dataset in my submission correctly,but it really work in my model training.Good job.",
      "replies": [
        {
          "id": 2496135,
          "postDate": "2023-10-23T18:34:58.707Z",
          "content": "<p>did you find better results by using those tiled images?</p>",
          "rawMarkdown": "did you find better results by using those tiled images?",
          "replies": [
            {
              "id": 2496296,
              "postDate": "2023-10-23T23:40:22.603Z",
              "content": "<p>I tried to cropped the thumbnail images myself last week and it seems that it can work will in images which is not “tma”.But am facing time limit and ram limit when I crop test images in the submission.It is really a challenge for me.Now I am trying the use the dataset because it will take a lot of time if I cropped myself.</p>",
              "rawMarkdown": "I tried to cropped the thumbnail images myself last week and it seems that it can work will in images which is not “tma”.But am facing time limit and ram limit when I crop test images in the submission.It is really a challenge for me.Now I am trying the use the dataset because it will take a lot of time if I cropped myself."
            }
          ]
        }
      ]
    },
    {
      "id": 2485310,
      "postDate": "2023-10-17T05:35:14.623Z",
      "content": "<p>very helpful</p>",
      "rawMarkdown": "very helpful"
    },
    {
      "id": 2519499,
      "postDate": "2023-11-10T05:06:18.980Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2498350,
      "postDate": "2023-10-25T08:54:53.943Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/pjmathematician\" target=\"_blank\">@pjmathematician</a> I might be wrong on this but images containing inside train_thumbnail consists of black background as result if you take top K tiles most of them pick Top K tiles having Black-background. How did you solved that part. ?</p>",
      "rawMarkdown": "Hi @pjmathematician I might be wrong on this but images containing inside train_thumbnail consists of black background as result if you take top K tiles most of them pick Top K tiles having Black-background. How did you solved that part. ?\n",
      "isDeleted": true,
      "replies": [
        {
          "id": 2498369,
          "postDate": "2023-10-25T09:09:08.817Z",
          "content": "<p>Good question <a href=\"https://www.kaggle.com/asteyagaur\" target=\"_blank\">@asteyagaur</a> . Instead of picking <em>Top</em> k, you should pick the <em>middle</em> k. But after a closer look, in some images there are black tiles in the middle as well. My solution as of now randomly samples + takes middle tiles</p>",
          "rawMarkdown": "Good question @asteyagaur . Instead of picking *Top* k, you should pick the *middle* k. But after a closer look, in some images there are black tiles in the middle as well. My solution as of now randomly samples + takes middle tiles",
          "replies": [
            {
              "id": 2498984,
              "postDate": "2023-10-25T16:34:28.817Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2498987,
              "postDate": "2023-10-25T16:38:08.233Z",
              "content": "<p><a href=\"https://www.kaggle.com/pjmathematician\" target=\"_blank\">@pjmathematician</a> I converted black background with white one. Using this simple formula ( image_c[image_c==0] = 255 ) As black pixel denotes '0' and white pixel denotes '255'. Do you think it will work now as or it would have some complication on overall image ?</p>",
              "rawMarkdown": "@pjmathematician I converted black background with white one. Using this simple formula ( image_c[image_c==0] = 255 ) As black pixel denotes '0' and white pixel denotes '255'. Do you think it will work now as or it would have some complication on overall image ?",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2519496,
      "postDate": "2023-11-10T05:01:10.560Z",
      "content": "<p>Thanks for amazing share!</p>",
      "rawMarkdown": "Thanks for amazing share!"
    },
    {
      "id": 2480964,
      "postDate": "2023-10-13T17:04:45.333Z",
      "content": "<p>thanks a lot!</p>",
      "rawMarkdown": "thanks a lot!",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2478643,
      "author_name": "Noli Alonso",
      "author_url": "",
      "post_date": "2023-10-12T06:04:39.883000",
      "content": "<p>This is very helpful. Thank you!<br>\nBut there may be a little issue, upon browsing the training set thumbnails and full-size images, there is a lot of data cleaning that needs to be done. There are areas with colored markers, processing artifacts like scratches, holes, folds, dust, uneven staining etc., blurry areas, older overly pink slides, tumors located in other parts of the body instead of the ovary, poorly aligned annotations (some are flipped or inverted or translated horizontally), some have a single tiny tile with the tumor (the rest is normal tissue), and one slide is vertically stretched. Also, many slides have 2 or more slices of the same tumor, that are essentially the same. It would be a truly laborious task to clean it up so that these tiles can be optimal.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 2478699,
          "author_name": "pjmathematician",
          "author_url": "",
          "post_date": "2023-10-12T06:53:19.073000",
          "content": "<p>good point. Currently what i am doing is, I selected the middle K tiles from every image for training. And changing the position/K according to the CV score. Using every tile would be very inefficient as there are a total of ~500k tiles.</p>",
          "votes": 10,
          "replies": [
            {
              "id": 2554266,
              "author_name": "Pranav Belhekar",
              "author_url": "",
              "post_date": "2023-12-09T01:38:39.317000",
              "content": "<p>Ah! good idea. Thanks for sharing!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2499693,
      "author_name": "Jirka",
      "author_url": "",
      "post_date": "2023-10-26T07:23:08.463000",
      "content": "<p>I have here also 512*512 with scale 0.25<br>\n<a href=\"https://www.kaggle.com/datasets/jirkaborovec/tiles-of-cancer-2048px-scale-0-25\" target=\"_blank\">https://www.kaggle.com/datasets/jirkaborovec/tiles-of-cancer-2048px-scale-0-25</a></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2491605,
      "author_name": "Ivan Panshin",
      "author_url": "",
      "post_date": "2023-10-21T19:06:32.057000",
      "content": "<p>Thanks for the datasets! </p>\n<p>I had to convert everything myself. You know it's gonna be a fun competition when you have to buy another 4TB disk to play around with a single dataset</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2570054,
      "author_name": "vauclin etienne",
      "author_url": "",
      "post_date": "2023-12-21T20:36:48.097000",
      "content": "<p>Hello, thank you for your efforts. I'm interested in understanding whether it's necessary to implement tile processing on the test set as well.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2491559,
      "author_name": "sunil thite",
      "author_url": "",
      "post_date": "2023-10-21T17:52:28.403000",
      "content": "<p>Good work . Keep it up</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2665170,
      "author_name": "USMAN SAFDAR",
      "author_url": "",
      "post_date": "2024-02-23T13:35:46.167000",
      "content": "<p><a href=\"https://www.kaggle.com/pjmathematician\" target=\"_blank\">@pjmathematician</a> The datasets you provided have been incredibly helpful. Currently, I'm using this code to generate my patches. Could you please explain how you converted the outputs into datasets? I'm encountering an error when trying to create a dataset directly from the output. Any guidance would be appreciated!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2667025,
          "author_name": "pjmathematician",
          "author_url": "",
          "post_date": "2024-02-24T20:11:56.020000",
          "content": "<p>Hi! <a href=\"https://www.kaggle.com/usmansafdar09\" target=\"_blank\">@usmansafdar09</a> ! <br>\nIf the output of the notebooks is throwing error, I suggest you use <a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">Kaggle API</a> in your notebook to directly create the dataset once the patches are made</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2515560,
      "author_name": "Tahsin",
      "author_url": "",
      "post_date": "2023-11-07T02:50:39.450000",
      "content": "<p>In the 256_13568 folder, the left part of the tissue seems to be missing from the extracted patches. Did you pad the images to be divisible by cropsize before creating tiles?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2507749,
      "author_name": "WangXuC",
      "author_url": "",
      "post_date": "2023-11-01T08:16:14.440000",
      "content": "<p>Have you scaled the image? If you have scaled, what is the scalratio?<br>\nDid you crop it on the original image or on the thumbnail?<br>\nLooking forward to your reply!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2503359,
      "author_name": "tyaba",
      "author_url": "",
      "post_date": "2023-10-29T02:52:53.400000",
      "content": "<p>Thank you for sharing!<br>\nHow do you run the Go language during inference?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2493575,
      "author_name": "Tushar",
      "author_url": "",
      "post_date": "2023-10-23T13:39:47.257000",
      "content": "<p>Did you employ the tile images as-is for training, or did you alter them to resemble a circular shape? and how many images are you considering in your training dataset, as there is a data imbalance within the classes?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2493582,
          "author_name": "pjmathematician",
          "author_url": "",
          "post_date": "2023-10-23T13:42:41.810000",
          "content": "<p>Good question, I will release a tile based training-inference baseline submission this week. That should give a basic idea of the usage</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2493589,
              "author_name": "Tushar",
              "author_url": "",
              "post_date": "2023-10-23T13:45:32.243000",
              "content": "<p>Also the TMA images are captured at a 40x zoom level, but I think the tiled images might have been zoomed in slightly more. What do you think?</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2491951,
      "author_name": "JaewooChoi",
      "author_url": "",
      "post_date": "2023-10-22T06:53:23.573000",
      "content": "<p>How do I use these tiled images? Do you train it with the tiled image itself as input? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2487225,
      "author_name": "LuoZiqian",
      "author_url": "",
      "post_date": "2023-10-18T13:01:36.243000",
      "content": "<p>Thought I didn't know how to use the dataset in my submission correctly,but it really work in my model training.Good job.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2496135,
          "author_name": "Tushar",
          "author_url": "",
          "post_date": "2023-10-23T18:34:58.707000",
          "content": "<p>did you find better results by using those tiled images?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2496296,
              "author_name": "LuoZiqian",
              "author_url": "",
              "post_date": "2023-10-23T23:40:22.603000",
              "content": "<p>I tried to cropped the thumbnail images myself last week and it seems that it can work will in images which is not “tma”.But am facing time limit and ram limit when I crop test images in the submission.It is really a challenge for me.Now I am trying the use the dataset because it will take a lot of time if I cropped myself.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2485310,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-17T05:35:14.623000",
      "content": "<p>very helpful</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2519499,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-11-10T05:06:18.980000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2498350,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-25T08:54:53.943000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/pjmathematician\" target=\"_blank\">@pjmathematician</a> I might be wrong on this but images containing inside train_thumbnail consists of black background as result if you take top K tiles most of them pick Top K tiles having Black-background. How did you solved that part. ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2498369,
          "author_name": "pjmathematician",
          "author_url": "",
          "post_date": "2023-10-25T09:09:08.817000",
          "content": "<p>Good question <a href=\"https://www.kaggle.com/asteyagaur\" target=\"_blank\">@asteyagaur</a> . Instead of picking <em>Top</em> k, you should pick the <em>middle</em> k. But after a closer look, in some images there are black tiles in the middle as well. My solution as of now randomly samples + takes middle tiles</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2498984,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-10-25T16:34:28.817000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2498987,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-10-25T16:38:08.233000",
              "content": "<p><a href=\"https://www.kaggle.com/pjmathematician\" target=\"_blank\">@pjmathematician</a> I converted black background with white one. Using this simple formula ( image_c[image_c==0] = 255 ) As black pixel denotes '0' and white pixel denotes '255'. Do you think it will work now as or it would have some complication on overall image ?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2519496,
      "author_name": "Autumn",
      "author_url": "",
      "post_date": "2023-11-10T05:01:10.560000",
      "content": "<p>Thanks for amazing share!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2480964,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-13T17:04:45.333000",
      "content": "<p>thanks a lot!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2476032": "Hello! Based on this amazing notebook : https://www.kaggle.com/code/dheerajmpai/ucb-ocean-process-large-images-golang-50x-fast \nI have resized and tiled the images, with removal of *almost black* images, using the following notebooks:\n\nhttps://www.kaggle.com/code/pjmathematician/ucbo-tilemaker\nhttps://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p2\nhttps://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p3\nhttps://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p4\nhttps://www.kaggle.com/code/pjmathematician/ucbo-tilemaker-p5\n\nThe datasets can be found here:\nhttps://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-1\nhttps://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-2\nhttps://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-3\nhttps://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-4\nhttps://www.kaggle.com/datasets/pjmathematician/ucbo-tiles-256-5\n\nAn example notebook to load the datasets is:\nhttps://www.kaggle.com/pjmathematician/ucbo-256-tiles-loading\n\nI will create 512x512 datasets as well very soon!!\nThank you for this amazing competition, looking forward to participating in it!\n\n",
    "2478643": "This is very helpful. Thank you!\nBut there may be a little issue, upon browsing the training set thumbnails and full-size images, there is a lot of data cleaning that needs to be done. There are areas with colored markers, processing artifacts like scratches, holes, folds, dust, uneven staining etc., blurry areas, older overly pink slides, tumors located in other parts of the body instead of the ovary, poorly aligned annotations (some are flipped or inverted or translated horizontally), some have a single tiny tile with the tumor (the rest is normal tissue), and one slide is vertically stretched. Also, many slides have 2 or more slices of the same tumor, that are essentially the same. It would be a truly laborious task to clean it up so that these tiles can be optimal.",
    "2499693": "I have here also 512*512 with scale 0.25\nhttps://www.kaggle.com/datasets/jirkaborovec/tiles-of-cancer-2048px-scale-0-25",
    "2491605": "Thanks for the datasets! \n\nI had to convert everything myself. You know it's gonna be a fun competition when you have to buy another 4TB disk to play around with a single dataset",
    "2570054": "\nHello, thank you for your efforts. I'm interested in understanding whether it's necessary to implement tile processing on the test set as well.",
    "2491559": "Good work . Keep it up",
    "2665170": "@pjmathematician The datasets you provided have been incredibly helpful. Currently, I'm using this code to generate my patches. Could you please explain how you converted the outputs into datasets? I'm encountering an error when trying to create a dataset directly from the output. Any guidance would be appreciated!",
    "2515560": "In the 256_13568 folder, the left part of the tissue seems to be missing from the extracted patches. Did you pad the images to be divisible by cropsize before creating tiles?",
    "2507749": "Have you scaled the image? If you have scaled, what is the scalratio?\nDid you crop it on the original image or on the thumbnail?\nLooking forward to your reply!",
    "2503359": "Thank you for sharing!\nHow do you run the Go language during inference?",
    "2493575": "Did you employ the tile images as-is for training, or did you alter them to resemble a circular shape? and how many images are you considering in your training dataset, as there is a data imbalance within the classes?",
    "2491951": "How do I use these tiled images? Do you train it with the tiled image itself as input? ",
    "2487225": "Thought I didn't know how to use the dataset in my submission correctly,but it really work in my model training.Good job.",
    "2485310": "very helpful",
    "2519499": "",
    "2498350": "Hi @pjmathematician I might be wrong on this but images containing inside train_thumbnail consists of black background as result if you take top K tiles most of them pick Top K tiles having Black-background. How did you solved that part. ?\n",
    "2519496": "Thanks for amazing share!",
    "2480964": "thanks a lot!"
  }
}