{
  "id": 338299,
  "title": "Mayo Tile Images Dataset [size=1024, N=16]",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/338299",
  "author_name": "Y.Nakama",
  "post_date": "2022-07-20T02:29:44.473000",
  "votes": 40,
  "comment_count": 24,
  "views": 0,
  "content": "<p>Tile method is already introduced by <a href=\"https://www.kaggle.com/analokamus\" target=\"_blank\">@analokamus</a> in this <a href=\"https://www.kaggle.com/code/analokamus/a-fast-tile-generation\" target=\"_blank\">notebook</a>.</p>\n<p>This time I prepared tile images that size=1024, N=16 using kaggle notebook.</p>\n<ul>\n<li>Notebook: <a href=\"https://www.kaggle.com/code/yasufuminakama/mayo-train-images-size-1024-n-16-1\" target=\"_blank\">https://www.kaggle.com/code/yasufuminakama/mayo-train-images-size-1024-n-16-1</a></li>\n<li>Dataset: <a href=\"https://www.kaggle.com/datasets/yasufuminakama/mayo-train-images-size1024-n16\" target=\"_blank\">https://www.kaggle.com/datasets/yasufuminakama/mayo-train-images-size1024-n16</a></li>\n</ul>\n<pre><code>def tile(img, sz=128, N=16):\n    shape = img.shape\n    pad0,pad1 = (sz - shape[0]%sz)%sz, (sz - shape[1]%sz)%sz\n    img = np.pad(img,[[pad0//2,pad0-pad0//2],[pad1//2,pad1-pad1//2],[0,0]],constant_values=255)\n    img = img.reshape(img.shape[0]//sz,sz,img.shape[1]//sz,sz,3)\n    img = img.transpose(0,2,1,3,4).reshape(-1,sz,sz,3)\n    if len(img) &lt; N:\n        img = np.pad(img,[[0,N-len(img)],[0,0],[0,0],[0,0]],constant_values=255)\n    idxs = np.argsort(img.reshape(img.shape[0],-1).sum(-1))[:N]\n    img = img[idxs]\n    return img\n\n\ndef save_dataset(\n    df: pd.DataFrame, \n    N=16,\n    max_size=20000, \n    crop_size=1024, \n    image_dir='../input/mayo-clinic-strip-ai/train', \n    out_dir='train_images.zip',\n):\n    format_to_dtype = {\n       'uchar': np.uint8,\n       'char': np.int8,\n       'ushort': np.uint16,\n       'short': np.int16,\n       'uint': np.uint32,\n       'int': np.int32,\n       'float': np.float32,\n       'double': np.float64,\n       'complex': np.complex64,\n       'dpcomplex': np.complex128,\n    }\n    def vips2numpy(vi):\n        return np.ndarray(\n            buffer=vi.write_to_memory(),\n            dtype=format_to_dtype[vi.format],\n            shape=[vi.height, vi.width, vi.bands])\n    with zipfile.ZipFile(out_dir, \"w\") as out_image:\n        tk0 = tqdm(enumerate(df[\"image_id\"].values), total=len(df))\n        for i, image_id in tk0:\n            print(f\"[{i+1}/{len(df)}] image_id: {image_id}\")\n            image = pyvips.Image.thumbnail(f'{image_dir}/{image_id}.tif', max_size)\n            image = vips2numpy(image)\n            width, height, c = image.shape\n            print(f\"Input width: {width} height: {height}\")\n            images = tile(image, sz=crop_size, N=N)\n            for idx, img in enumerate(images):\n                img = cv2.cvtColor(img, cv2.COLOR_RGB2BGR)\n                img = cv2.imencode(\".jpg\", img, [cv2.IMWRITE_JPEG_QUALITY, 100])[1]\n                out_image.writestr(f\"{image_id}_{idx}.jpg\", img)\n            del img, image, images; gc.collect()\n\ni = 1 # 1~8\n\nif i == 8:\n    df = train[(i-1)*100:]\nelse:\n    df = train[(i-1)*100:i*100]\n\nsave_dataset(\n    df,\n    N=16, \n    max_size=20000,\n    crop_size=1024, \n    image_dir='../input/mayo-clinic-strip-ai/train', \n    out_dir=f'train_images_{i}.zip'\n)\n</code></pre>",
  "messages": [
    {
      "id": 1862799,
      "postDate": "2022-07-20T02:29:44.473Z",
      "content": "<p>Tile method is already introduced by <a href=\"https://www.kaggle.com/analokamus\" target=\"_blank\">@analokamus</a> in this <a href=\"https://www.kaggle.com/code/analokamus/a-fast-tile-generation\" target=\"_blank\">notebook</a>.</p>\n<p>This time I prepared tile images that size=1024, N=16 using kaggle notebook.</p>\n<ul>\n<li>Notebook: <a href=\"https://www.kaggle.com/code/yasufuminakama/mayo-train-images-size-1024-n-16-1\" target=\"_blank\">https://www.kaggle.com/code/yasufuminakama/mayo-train-images-size-1024-n-16-1</a></li>\n<li>Dataset: <a href=\"https://www.kaggle.com/datasets/yasufuminakama/mayo-train-images-size1024-n16\" target=\"_blank\">https://www.kaggle.com/datasets/yasufuminakama/mayo-train-images-size1024-n16</a></li>\n</ul>\n<pre><code>def tile(img, sz=128, N=16):\n    shape = img.shape\n    pad0,pad1 = (sz - shape[0]%sz)%sz, (sz - shape[1]%sz)%sz\n    img = np.pad(img,[[pad0//2,pad0-pad0//2],[pad1//2,pad1-pad1//2],[0,0]],constant_values=255)\n    img = img.reshape(img.shape[0]//sz,sz,img.shape[1]//sz,sz,3)\n    img = img.transpose(0,2,1,3,4).reshape(-1,sz,sz,3)\n    if len(img) &lt; N:\n        img = np.pad(img,[[0,N-len(img)],[0,0],[0,0],[0,0]],constant_values=255)\n    idxs = np.argsort(img.reshape(img.shape[0],-1).sum(-1))[:N]\n    img = img[idxs]\n    return img\n\n\ndef save_dataset(\n    df: pd.DataFrame, \n    N=16,\n    max_size=20000, \n    crop_size=1024, \n    image_dir='../input/mayo-clinic-strip-ai/train', \n    out_dir='train_images.zip',\n):\n    format_to_dtype = {\n       'uchar': np.uint8,\n       'char': np.int8,\n       'ushort': np.uint16,\n       'short': np.int16,\n       'uint': np.uint32,\n       'int': np.int32,\n       'float': np.float32,\n       'double': np.float64,\n       'complex': np.complex64,\n       'dpcomplex': np.complex128,\n    }\n    def vips2numpy(vi):\n        return np.ndarray(\n            buffer=vi.write_to_memory(),\n            dtype=format_to_dtype[vi.format],\n            shape=[vi.height, vi.width, vi.bands])\n    with zipfile.ZipFile(out_dir, \"w\") as out_image:\n        tk0 = tqdm(enumerate(df[\"image_id\"].values), total=len(df))\n        for i, image_id in tk0:\n            print(f\"[{i+1}/{len(df)}] image_id: {image_id}\")\n            image = pyvips.Image.thumbnail(f'{image_dir}/{image_id}.tif', max_size)\n            image = vips2numpy(image)\n            width, height, c = image.shape\n            print(f\"Input width: {width} height: {height}\")\n            images = tile(image, sz=crop_size, N=N)\n            for idx, img in enumerate(images):\n                img = cv2.cvtColor(img, cv2.COLOR_RGB2BGR)\n                img = cv2.imencode(\".jpg\", img, [cv2.IMWRITE_JPEG_QUALITY, 100])[1]\n                out_image.writestr(f\"{image_id}_{idx}.jpg\", img)\n            del img, image, images; gc.collect()\n\ni = 1 # 1~8\n\nif i == 8:\n    df = train[(i-1)*100:]\nelse:\n    df = train[(i-1)*100:i*100]\n\nsave_dataset(\n    df,\n    N=16, \n    max_size=20000,\n    crop_size=1024, \n    image_dir='../input/mayo-clinic-strip-ai/train', \n    out_dir=f'train_images_{i}.zip'\n)\n</code></pre>",
      "rawMarkdown": "Tile method is already introduced by @analokamus in this [notebook](https://www.kaggle.com/code/analokamus/a-fast-tile-generation).\n\nThis time I prepared tile images that size=1024, N=16 using kaggle notebook.\n\n- Notebook: https://www.kaggle.com/code/yasufuminakama/mayo-train-images-size-1024-n-16-1\n- Dataset: https://www.kaggle.com/datasets/yasufuminakama/mayo-train-images-size1024-n16\n\n```\ndef tile(img, sz=128, N=16):\n    shape = img.shape\n    pad0,pad1 = (sz - shape[0]%sz)%sz, (sz - shape[1]%sz)%sz\n    img = np.pad(img,[[pad0//2,pad0-pad0//2],[pad1//2,pad1-pad1//2],[0,0]],constant_values=255)\n    img = img.reshape(img.shape[0]//sz,sz,img.shape[1]//sz,sz,3)\n    img = img.transpose(0,2,1,3,4).reshape(-1,sz,sz,3)\n    if len(img) < N:\n        img = np.pad(img,[[0,N-len(img)],[0,0],[0,0],[0,0]],constant_values=255)\n    idxs = np.argsort(img.reshape(img.shape[0],-1).sum(-1))[:N]\n    img = img[idxs]\n    return img\n\n\ndef save_dataset(\n    df: pd.DataFrame, \n    N=16,\n    max_size=20000, \n    crop_size=1024, \n    image_dir='../input/mayo-clinic-strip-ai/train', \n    out_dir='train_images.zip',\n):\n    format_to_dtype = {\n       'uchar': np.uint8,\n       'char': np.int8,\n       'ushort': np.uint16,\n       'short': np.int16,\n       'uint': np.uint32,\n       'int': np.int32,\n       'float': np.float32,\n       'double': np.float64,\n       'complex': np.complex64,\n       'dpcomplex': np.complex128,\n    }\n    def vips2numpy(vi):\n        return np.ndarray(\n            buffer=vi.write_to_memory(),\n            dtype=format_to_dtype[vi.format],\n            shape=[vi.height, vi.width, vi.bands])\n    with zipfile.ZipFile(out_dir, \"w\") as out_image:\n        tk0 = tqdm(enumerate(df[\"image_id\"].values), total=len(df))\n        for i, image_id in tk0:\n            print(f\"[{i+1}/{len(df)}] image_id: {image_id}\")\n            image = pyvips.Image.thumbnail(f'{image_dir}/{image_id}.tif', max_size)\n            image = vips2numpy(image)\n            width, height, c = image.shape\n            print(f\"Input width: {width} height: {height}\")\n            images = tile(image, sz=crop_size, N=N)\n            for idx, img in enumerate(images):\n                img = cv2.cvtColor(img, cv2.COLOR_RGB2BGR)\n                img = cv2.imencode(\".jpg\", img, [cv2.IMWRITE_JPEG_QUALITY, 100])[1]\n                out_image.writestr(f\"{image_id}_{idx}.jpg\", img)\n            del img, image, images; gc.collect()\n\ni = 1 # 1~8\n\nif i == 8:\n    df = train[(i-1)*100:]\nelse:\n    df = train[(i-1)*100:i*100]\n\nsave_dataset(\n    df,\n    N=16, \n    max_size=20000,\n    crop_size=1024, \n    image_dir='../input/mayo-clinic-strip-ai/train', \n    out_dir=f'train_images_{i}.zip'\n)\n```",
      "votes": 40
    },
    {
      "id": 1865539,
      "postDate": "2022-07-21T23:35:20.300Z",
      "content": "<p>hello <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>, in the notebook of <a href=\"https://www.kaggle.com/analokamus\" target=\"_blank\">@analokamus</a> it says <code>an efficient way of removing background (=making tiles) is very important</code>. I tried a similar approach but removed background either by setting a threshold or using a classifier, this dataset includes backgrounds or just clots? Thanks in advance</p>",
      "rawMarkdown": "hello @yasufuminakama, in the notebook of @analokamus it says `an efficient way of removing background (=making tiles) is very important`. I tried a similar approach but removed background either by setting a threshold or using a classifier, this dataset includes backgrounds or just clots? Thanks in advance",
      "votes": 1,
      "replies": [
        {
          "id": 1866558,
          "postDate": "2022-07-22T15:20:39.633Z",
          "content": "<p>It picks clots images as possible but in some cases probably background images are included.<br>\nFor example: {id}_1 ~ {id}_12 are clots images but {id}_13 ~ {id}_16 are background images.</p>",
          "rawMarkdown": "It picks clots images as possible but in some cases probably background images are included.\nFor example: {id}_1 ~ {id}_12 are clots images but {id}_13 ~ {id}_16 are background images."
        }
      ]
    },
    {
      "id": 1863169,
      "postDate": "2022-07-20T07:28:24Z",
      "content": "<p>awesome work !! <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>  Thanks for sharing !</p>",
      "rawMarkdown": "awesome work !! @yasufuminakama  Thanks for sharing !",
      "votes": 1
    },
    {
      "id": 1863070,
      "postDate": "2022-07-20T06:26:09.137Z",
      "content": "<p>Do you have use the entire dataset to train?</p>",
      "rawMarkdown": "Do you have use the entire dataset to train?",
      "votes": 1,
      "replies": [
        {
          "id": 1863092,
          "postDate": "2022-07-20T06:37:46.980Z",
          "content": "<p>Yes, my current LB: 0.4 model is trained using this entire train dataset.</p>",
          "rawMarkdown": "Yes, my current LB: 0.4 model is trained using this entire train dataset.",
          "votes": 3
        },
        {
          "id": 1935373,
          "postDate": "2022-09-12T04:27:35.113Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> , How did you validate?</p>",
          "rawMarkdown": "Hi @yasufuminakama , How did you validate?"
        },
        {
          "id": 1935494,
          "postDate": "2022-09-12T06:46:30.747Z",
          "content": "<p>I used StratifiedGroupkfold(target, center_id).</p>",
          "rawMarkdown": "I used StratifiedGroupkfold(target, center_id).",
          "votes": 1
        }
      ]
    },
    {
      "id": 1968232,
      "postDate": "2022-10-03T02:05:54.920Z",
      "content": "<p>The good thing about the tile function is the argsort that throw the too much white images.</p>",
      "rawMarkdown": "The good thing about the tile function is the argsort that throw the too much white images."
    },
    {
      "id": 1968230,
      "postDate": "2022-10-03T02:04:08.170Z",
      "content": "<p>I am new in code competition. Is it possible to install pyvips in the code submission since we have not the internet access?</p>",
      "rawMarkdown": "I am new in code competition. Is it possible to install pyvips in the code submission since we have not the internet access?",
      "replies": [
        {
          "id": 1969070,
          "postDate": "2022-10-03T10:26:05.157Z",
          "content": "<p>See <a href=\"https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline\" target=\"_blank\">https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline</a>.</p>",
          "rawMarkdown": "See https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline."
        }
      ]
    },
    {
      "id": 1908688,
      "postDate": "2022-08-21T22:33:52.850Z",
      "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> I was able to train with tiles but when I want to infer the notebook crash because the tile function consumes a lot of memory, please any idea ?, I saw some notebook where test images are transformed  prior to the submission, the issue here is I suppose that the test images will change afterwards.</p>\n<p>Thanks again</p>",
      "rawMarkdown": "@yasufuminakama I was able to train with tiles but when I want to infer the notebook crash because the tile function consumes a lot of memory, please any idea ?, I saw some notebook where test images are transformed  prior to the submission, the issue here is I suppose that the test images will change afterwards.\n\nThanks again"
    },
    {
      "id": 1908682,
      "postDate": "2022-08-21T22:15:35.083Z",
      "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>  when you train with tiles you also need to tile test images?</p>\n<p>Thanks a lot</p>",
      "rawMarkdown": "@yasufuminakama  when you train with tiles you also need to tile test images?\n\nThanks a lot",
      "replies": [
        {
          "id": 1908684,
          "postDate": "2022-08-21T22:17:21.733Z",
          "content": "<blockquote>\n  <p>when you train with tiles you also need to tile test images?</p>\n</blockquote>\n<p>Yes, I also tile test images.</p>",
          "rawMarkdown": "> when you train with tiles you also need to tile test images?\n\nYes, I also tile test images.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1898766,
      "postDate": "2022-08-14T19:42:35.650Z",
      "content": "<p>Hi, how do I do to train and predict when I'm using tiled images?.</p>\n<p>Thanks in advanced!</p>",
      "rawMarkdown": "Hi, how do I do to train and predict when I'm using tiled images?.\n\nThanks in advanced!",
      "replies": [
        {
          "id": 1899004,
          "postDate": "2022-08-15T02:00:58.667Z",
          "content": "<p><a href=\"https://www.kaggle.com/code/analokamus/a-sample-of-multi-instance-learning-model\" target=\"_blank\">This notebook</a> would be helpful.</p>",
          "rawMarkdown": "[This notebook](https://www.kaggle.com/code/analokamus/a-sample-of-multi-instance-learning-model) would be helpful.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1894361,
      "postDate": "2022-08-11T12:47:15.290Z",
      "content": "<p>I see that Dataset made it from 1 to 8. Thank you!</p>",
      "rawMarkdown": "I see that Dataset made it from 1 to 8. Thank you!"
    },
    {
      "id": 1881745,
      "postDate": "2022-08-02T18:27:58.993Z",
      "content": "<p>Thanks for sharing! I'll be using this dataset to train my model.</p>",
      "rawMarkdown": "Thanks for sharing! I'll be using this dataset to train my model."
    },
    {
      "id": 1868381,
      "postDate": "2022-07-24T00:09:15.857Z",
      "content": "<p>Thanks for sharing your notebook! This is really helpful. I have a question, though - why do you recommend using a crop size of 1024? It seems like a really large crop size and I'm not sure how it would benefit the training process. Thanks in advance!</p>",
      "rawMarkdown": "Thanks for sharing your notebook! This is really helpful. I have a question, though - why do you recommend using a crop size of 1024? It seems like a really large crop size and I'm not sure how it would benefit the training process. Thanks in advance!",
      "replies": [
        {
          "id": 1868395,
          "postDate": "2022-07-24T00:21:45.833Z",
          "content": "<p>I think its a tradeoff, on the one hand a larger crop means less processing which can translate in shorter training times and inference since you would have less crops. If you crop smaller means more training but you preserve more information since you do not have to downsize. I would like to see whats the best cropping size but rather than 1024 no one has shared an alternative.</p>",
          "rawMarkdown": "I think its a tradeoff, on the one hand a larger crop means less processing which can translate in shorter training times and inference since you would have less crops. If you crop smaller means more training but you preserve more information since you do not have to downsize. I would like to see whats the best cropping size but rather than 1024 no one has shared an alternative."
        }
      ]
    },
    {
      "id": 1864819,
      "postDate": "2022-07-21T10:18:54.880Z",
      "content": "<p>Hi! Really helpful dataset, thanks! Could you please suggest what val strategy have you used? </p>",
      "rawMarkdown": "Hi! Really helpful dataset, thanks! Could you please suggest what val strategy have you used? ",
      "replies": [
        {
          "id": 1866561,
          "postDate": "2022-07-22T15:23:03.867Z",
          "content": "<p>I'm using StratifiedGroupkfold(target, center_id).</p>",
          "rawMarkdown": "I'm using StratifiedGroupkfold(target, center_id)."
        },
        {
          "id": 1911444,
          "postDate": "2022-08-24T05:26:03.137Z",
          "content": "<p>Your current performance is from a single model or from an ensemble ? </p>",
          "rawMarkdown": "Your current performance is from a single model or from an ensemble ? "
        },
        {
          "id": 1918144,
          "postDate": "2022-08-29T11:01:09.850Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>! you mean stratify by label and group by <code>center_id</code> right? <br>\nwhat about the <code>patient_id</code> ? shouldn't we make sure that same patient is not in train/val sets or it doesn't matter for this one?  <br>\nsorry for basic/newbie question I have just joined and trying to figure out</p>",
          "rawMarkdown": "Thanks @yasufuminakama! you mean stratify by label and group by `center_id` right? \nwhat about the `patient_id` ? shouldn't we make sure that same patient is not in train/val sets or it doesn't matter for this one?  \nsorry for basic/newbie question I have just joined and trying to figure out",
          "votes": 1
        }
      ]
    },
    {
      "id": 1863020,
      "postDate": "2022-07-20T05:43:40.857Z",
      "content": "<p>Awesome, thanks 💪</p>",
      "rawMarkdown": "Awesome, thanks 💪"
    }
  ],
  "comments": [
    {
      "id": 1865539,
      "author_name": "moth",
      "author_url": "",
      "post_date": "2022-07-21T23:35:20.300000",
      "content": "<p>hello <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>, in the notebook of <a href=\"https://www.kaggle.com/analokamus\" target=\"_blank\">@analokamus</a> it says <code>an efficient way of removing background (=making tiles) is very important</code>. I tried a similar approach but removed background either by setting a threshold or using a classifier, this dataset includes backgrounds or just clots? Thanks in advance</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1866558,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2022-07-22T15:20:39.633000",
          "content": "<p>It picks clots images as possible but in some cases probably background images are included.<br>\nFor example: {id}_1 ~ {id}_12 are clots images but {id}_13 ~ {id}_16 are background images.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1863169,
      "author_name": "Arun Purakkatt",
      "author_url": "",
      "post_date": "2022-07-20T07:28:24",
      "content": "<p>awesome work !! <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>  Thanks for sharing !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1863070,
      "author_name": "alex",
      "author_url": "",
      "post_date": "2022-07-20T06:26:09.137000",
      "content": "<p>Do you have use the entire dataset to train?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1863092,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2022-07-20T06:37:46.980000",
          "content": "<p>Yes, my current LB: 0.4 model is trained using this entire train dataset.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1935373,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2022-09-12T04:27:35.113000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> , How did you validate?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1935494,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2022-09-12T06:46:30.747000",
          "content": "<p>I used StratifiedGroupkfold(target, center_id).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1968232,
      "author_name": "Pierre Tisseur",
      "author_url": "",
      "post_date": "2022-10-03T02:05:54.920000",
      "content": "<p>The good thing about the tile function is the argsort that throw the too much white images.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1968230,
      "author_name": "Pierre Tisseur",
      "author_url": "",
      "post_date": "2022-10-03T02:04:08.170000",
      "content": "<p>I am new in code competition. Is it possible to install pyvips in the code submission since we have not the internet access?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1969070,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2022-10-03T10:26:05.157000",
          "content": "<p>See <a href=\"https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline\" target=\"_blank\">https://www.kaggle.com/code/analokamus/how-to-use-pyvips-offline</a>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1908688,
      "author_name": "Pablo Larrosa",
      "author_url": "",
      "post_date": "2022-08-21T22:33:52.850000",
      "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> I was able to train with tiles but when I want to infer the notebook crash because the tile function consumes a lot of memory, please any idea ?, I saw some notebook where test images are transformed  prior to the submission, the issue here is I suppose that the test images will change afterwards.</p>\n<p>Thanks again</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1908682,
      "author_name": "Pablo Larrosa",
      "author_url": "",
      "post_date": "2022-08-21T22:15:35.083000",
      "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>  when you train with tiles you also need to tile test images?</p>\n<p>Thanks a lot</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1908684,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2022-08-21T22:17:21.733000",
          "content": "<blockquote>\n  <p>when you train with tiles you also need to tile test images?</p>\n</blockquote>\n<p>Yes, I also tile test images.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1898766,
      "author_name": "Pablo Larrosa",
      "author_url": "",
      "post_date": "2022-08-14T19:42:35.650000",
      "content": "<p>Hi, how do I do to train and predict when I'm using tiled images?.</p>\n<p>Thanks in advanced!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1899004,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2022-08-15T02:00:58.667000",
          "content": "<p><a href=\"https://www.kaggle.com/code/analokamus/a-sample-of-multi-instance-learning-model\" target=\"_blank\">This notebook</a> would be helpful.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1894361,
      "author_name": "min fuka",
      "author_url": "",
      "post_date": "2022-08-11T12:47:15.290000",
      "content": "<p>I see that Dataset made it from 1 to 8. Thank you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1881745,
      "author_name": "Nghi Huynh",
      "author_url": "",
      "post_date": "2022-08-02T18:27:58.993000",
      "content": "<p>Thanks for sharing! I'll be using this dataset to train my model.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1868381,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-07-24T00:09:15.857000",
      "content": "<p>Thanks for sharing your notebook! This is really helpful. I have a question, though - why do you recommend using a crop size of 1024? It seems like a really large crop size and I'm not sure how it would benefit the training process. Thanks in advance!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1868395,
          "author_name": "moth",
          "author_url": "",
          "post_date": "2022-07-24T00:21:45.833000",
          "content": "<p>I think its a tradeoff, on the one hand a larger crop means less processing which can translate in shorter training times and inference since you would have less crops. If you crop smaller means more training but you preserve more information since you do not have to downsize. I would like to see whats the best cropping size but rather than 1024 no one has shared an alternative.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1864819,
      "author_name": "Vladislav Ostankovich",
      "author_url": "",
      "post_date": "2022-07-21T10:18:54.880000",
      "content": "<p>Hi! Really helpful dataset, thanks! Could you please suggest what val strategy have you used? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1866561,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2022-07-22T15:23:03.867000",
          "content": "<p>I'm using StratifiedGroupkfold(target, center_id).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1911444,
          "author_name": "Balaji Selvaraj",
          "author_url": "",
          "post_date": "2022-08-24T05:26:03.137000",
          "content": "<p>Your current performance is from a single model or from an ensemble ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1918144,
          "author_name": "Ioannis M",
          "author_url": "",
          "post_date": "2022-08-29T11:01:09.850000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>! you mean stratify by label and group by <code>center_id</code> right? <br>\nwhat about the <code>patient_id</code> ? shouldn't we make sure that same patient is not in train/val sets or it doesn't matter for this one?  <br>\nsorry for basic/newbie question I have just joined and trying to figure out</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1863020,
      "author_name": "Ronaldo S.A. Batista",
      "author_url": "",
      "post_date": "2022-07-20T05:43:40.857000",
      "content": "<p>Awesome, thanks 💪</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1862799": "Tile method is already introduced by @analokamus in this [notebook](https://www.kaggle.com/code/analokamus/a-fast-tile-generation).\n\nThis time I prepared tile images that size=1024, N=16 using kaggle notebook.\n\n- Notebook: https://www.kaggle.com/code/yasufuminakama/mayo-train-images-size-1024-n-16-1\n- Dataset: https://www.kaggle.com/datasets/yasufuminakama/mayo-train-images-size1024-n16\n\n```\ndef tile(img, sz=128, N=16):\n    shape = img.shape\n    pad0,pad1 = (sz - shape[0]%sz)%sz, (sz - shape[1]%sz)%sz\n    img = np.pad(img,[[pad0//2,pad0-pad0//2],[pad1//2,pad1-pad1//2],[0,0]],constant_values=255)\n    img = img.reshape(img.shape[0]//sz,sz,img.shape[1]//sz,sz,3)\n    img = img.transpose(0,2,1,3,4).reshape(-1,sz,sz,3)\n    if len(img) < N:\n        img = np.pad(img,[[0,N-len(img)],[0,0],[0,0],[0,0]],constant_values=255)\n    idxs = np.argsort(img.reshape(img.shape[0],-1).sum(-1))[:N]\n    img = img[idxs]\n    return img\n\n\ndef save_dataset(\n    df: pd.DataFrame, \n    N=16,\n    max_size=20000, \n    crop_size=1024, \n    image_dir='../input/mayo-clinic-strip-ai/train', \n    out_dir='train_images.zip',\n):\n    format_to_dtype = {\n       'uchar': np.uint8,\n       'char': np.int8,\n       'ushort': np.uint16,\n       'short': np.int16,\n       'uint': np.uint32,\n       'int': np.int32,\n       'float': np.float32,\n       'double': np.float64,\n       'complex': np.complex64,\n       'dpcomplex': np.complex128,\n    }\n    def vips2numpy(vi):\n        return np.ndarray(\n            buffer=vi.write_to_memory(),\n            dtype=format_to_dtype[vi.format],\n            shape=[vi.height, vi.width, vi.bands])\n    with zipfile.ZipFile(out_dir, \"w\") as out_image:\n        tk0 = tqdm(enumerate(df[\"image_id\"].values), total=len(df))\n        for i, image_id in tk0:\n            print(f\"[{i+1}/{len(df)}] image_id: {image_id}\")\n            image = pyvips.Image.thumbnail(f'{image_dir}/{image_id}.tif', max_size)\n            image = vips2numpy(image)\n            width, height, c = image.shape\n            print(f\"Input width: {width} height: {height}\")\n            images = tile(image, sz=crop_size, N=N)\n            for idx, img in enumerate(images):\n                img = cv2.cvtColor(img, cv2.COLOR_RGB2BGR)\n                img = cv2.imencode(\".jpg\", img, [cv2.IMWRITE_JPEG_QUALITY, 100])[1]\n                out_image.writestr(f\"{image_id}_{idx}.jpg\", img)\n            del img, image, images; gc.collect()\n\ni = 1 # 1~8\n\nif i == 8:\n    df = train[(i-1)*100:]\nelse:\n    df = train[(i-1)*100:i*100]\n\nsave_dataset(\n    df,\n    N=16, \n    max_size=20000,\n    crop_size=1024, \n    image_dir='../input/mayo-clinic-strip-ai/train', \n    out_dir=f'train_images_{i}.zip'\n)\n```",
    "1865539": "hello @yasufuminakama, in the notebook of @analokamus it says `an efficient way of removing background (=making tiles) is very important`. I tried a similar approach but removed background either by setting a threshold or using a classifier, this dataset includes backgrounds or just clots? Thanks in advance",
    "1863169": "awesome work !! @yasufuminakama  Thanks for sharing !",
    "1863070": "Do you have use the entire dataset to train?",
    "1968232": "The good thing about the tile function is the argsort that throw the too much white images.",
    "1968230": "I am new in code competition. Is it possible to install pyvips in the code submission since we have not the internet access?",
    "1908688": "@yasufuminakama I was able to train with tiles but when I want to infer the notebook crash because the tile function consumes a lot of memory, please any idea ?, I saw some notebook where test images are transformed  prior to the submission, the issue here is I suppose that the test images will change afterwards.\n\nThanks again",
    "1908682": "@yasufuminakama  when you train with tiles you also need to tile test images?\n\nThanks a lot",
    "1898766": "Hi, how do I do to train and predict when I'm using tiled images?.\n\nThanks in advanced!",
    "1894361": "I see that Dataset made it from 1 to 8. Thank you!",
    "1881745": "Thanks for sharing! I'll be using this dataset to train my model.",
    "1868381": "Thanks for sharing your notebook! This is really helpful. I have a question, though - why do you recommend using a crop size of 1024? It seems like a really large crop size and I'm not sure how it would benefit the training process. Thanks in advance!",
    "1864819": "Hi! Really helpful dataset, thanks! Could you please suggest what val strategy have you used? ",
    "1863020": "Awesome, thanks 💪"
  }
}