{
  "id": 463328,
  "title": "Tile Generation Timing",
  "url": "/competitions/UBC-OCEAN/discussion/463328",
  "author_name": "chemdatafarmer",
  "post_date": "2023-12-24T16:35:14.288000",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi Everyone,</p>\n<p>I'm having some difficulty with getting the image generation part of my submission script to run quickly enough. I trained an efficient_net classifier (keras-cv with jax backend) on individual tiles and I'm using pyvips to cut the tiles into a regular grid ( &lt;= 32 tiles per image) for inference, then predicting by using a simple sum of the logits. My NN whips through these at about 5-6 seconds per image (less on the smaller batches), but my total processing time varies widely. Naturally, TMAs don't take that long (~3-4 seconds total processing time), but larger images take ~30-100 seconds, depending on the size.</p>\n<p>I'm just trying to get a baseline in, but my last submission timed out. I was hopeful that the majority of the test set was TMAs, but it seems there aren't enough to compensate for the time. With an average 30 seconds of processing time X 2000 images, you end up with about 17 hours of total processing time give or take a bit.</p>\n<p>What sort of timings are others here seeing and are you willing to share some tips and tricks on how you were able to achieve it? I was hoping to use this competition to learn how to build a model that generalizes well and can perform outlier detection, but I've spent quite a lot of time fighting with submissions and trying to get the most out of pyvips…Any advice would be greatly appreciated.</p>",
  "messages": [
    {
      "id": 2574327,
      "postDate": "2023-12-25T20:21:12.793Z",
      "content": "<p>something like:</p>\n<pre><code> concurrent.futures \n\n (): \n    do tiling stuff...\n     tiles\n\n concurrent.futures.ThreadPoolExecutor(max_workers=)  executor:\n    results = ((executor.(process_tile) )\n</code></pre>\n<p>this is a dummy example, you need to adapt for your method and process!</p>",
      "rawMarkdown": "something like:\n\n\n```python\nimport concurrent.futures \n\ndef process_tile(): \n    do tiling stuff...\n    return tiles\n    \nwith concurrent.futures.ThreadPoolExecutor(max_workers=4) as executor:\n    results = list((executor.map(process_tile) )\n```\n\nthis is a dummy example, you need to adapt for your method and process!",
      "votes": 3,
      "replies": [
        {
          "id": 2574998,
          "postDate": "2023-12-26T12:52:35.337Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/aliabbasi\" target=\"_blank\">@aliabbasi</a> </p>",
          "rawMarkdown": "Thanks @aliabbasi "
        }
      ]
    },
    {
      "id": 2573127,
      "postDate": "2023-12-24T18:27:43.127Z",
      "content": "<p>almost had same issue, <br>\nnow I'm tiling images with pyvips on 4 core of CPUs, each core cropping a tile, and it runs at lease 25% faster. <br>\nyou can use python naive concurrent module </p>",
      "rawMarkdown": "almost had same issue, \nnow I'm tiling images with pyvips on 4 core of CPUs, each core cropping a tile, and it runs at lease 25% faster. \nyou can use python naive concurrent module ",
      "votes": 3,
      "replies": [
        {
          "id": 2573145,
          "postDate": "2023-12-24T18:43:26.797Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/aliabbasi\" target=\"_blank\">@aliabbasi</a> , thanks for the tip! Your pyvips offline notebook gave me a great starting point with this. However, I don't see anything in that notebook that shows how explicitly use a core to crop a tile. I'm also not sure what you mean by \"you can use python naive concurrent module\". Do you have an example of such a module so I can go learn how to use it?</p>",
          "rawMarkdown": "Hey @aliabbasi , thanks for the tip! Your pyvips offline notebook gave me a great starting point with this. However, I don't see anything in that notebook that shows how explicitly use a core to crop a tile. I'm also not sure what you mean by \"you can use python naive concurrent module\". Do you have an example of such a module so I can go learn how to use it?",
          "replies": [
            {
              "id": 2574852,
              "postDate": "2023-12-26T09:37:47.297Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/aliabbasi\" target=\"_blank\">@aliabbasi</a> While converting the whole images into tiles I did not get any issue related to the time . Are you converting images from train_images folder or from thumbnail folder?</p>",
              "rawMarkdown": "Hi @aliabbasi While converting the whole images into tiles I did not get any issue related to the time . Are you converting images from train_images folder or from thumbnail folder?",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2573025,
      "postDate": "2023-12-24T16:35:14.287Z",
      "content": "<p>Hi Everyone,</p>\n<p>I'm having some difficulty with getting the image generation part of my submission script to run quickly enough. I trained an efficient_net classifier (keras-cv with jax backend) on individual tiles and I'm using pyvips to cut the tiles into a regular grid ( &lt;= 32 tiles per image) for inference, then predicting by using a simple sum of the logits. My NN whips through these at about 5-6 seconds per image (less on the smaller batches), but my total processing time varies widely. Naturally, TMAs don't take that long (~3-4 seconds total processing time), but larger images take ~30-100 seconds, depending on the size.</p>\n<p>I'm just trying to get a baseline in, but my last submission timed out. I was hopeful that the majority of the test set was TMAs, but it seems there aren't enough to compensate for the time. With an average 30 seconds of processing time X 2000 images, you end up with about 17 hours of total processing time give or take a bit.</p>\n<p>What sort of timings are others here seeing and are you willing to share some tips and tricks on how you were able to achieve it? I was hoping to use this competition to learn how to build a model that generalizes well and can perform outlier detection, but I've spent quite a lot of time fighting with submissions and trying to get the most out of pyvips…Any advice would be greatly appreciated.</p>",
      "rawMarkdown": "Hi Everyone,\n\nI'm having some difficulty with getting the image generation part of my submission script to run quickly enough. I trained an efficient_net classifier (keras-cv with jax backend) on individual tiles and I'm using pyvips to cut the tiles into a regular grid ( <= 32 tiles per image) for inference, then predicting by using a simple sum of the logits. My NN whips through these at about 5-6 seconds per image (less on the smaller batches), but my total processing time varies widely. Naturally, TMAs don't take that long (~3-4 seconds total processing time), but larger images take ~30-100 seconds, depending on the size.\n\nI'm just trying to get a baseline in, but my last submission timed out. I was hopeful that the majority of the test set was TMAs, but it seems there aren't enough to compensate for the time. With an average 30 seconds of processing time X 2000 images, you end up with about 17 hours of total processing time give or take a bit.\n\nWhat sort of timings are others here seeing and are you willing to share some tips and tricks on how you were able to achieve it? I was hoping to use this competition to learn how to build a model that generalizes well and can perform outlier detection, but I've spent quite a lot of time fighting with submissions and trying to get the most out of pyvips...Any advice would be greatly appreciated.",
      "votes": 4
    }
  ],
  "comments": [
    {
      "id": 2574327,
      "author_name": "Ali",
      "author_url": "",
      "post_date": "2023-12-25T20:21:12.793000",
      "content": "<p>something like:</p>\n<pre><code> concurrent.futures \n\n (): \n    do tiling stuff...\n     tiles\n\n concurrent.futures.ThreadPoolExecutor(max_workers=)  executor:\n    results = ((executor.(process_tile) )\n</code></pre>\n<p>this is a dummy example, you need to adapt for your method and process!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2574998,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2023-12-26T12:52:35.337000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/aliabbasi\" target=\"_blank\">@aliabbasi</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2573127,
      "author_name": "Ali",
      "author_url": "",
      "post_date": "2023-12-24T18:27:43.127000",
      "content": "<p>almost had same issue, <br>\nnow I'm tiling images with pyvips on 4 core of CPUs, each core cropping a tile, and it runs at lease 25% faster. <br>\nyou can use python naive concurrent module </p>",
      "votes": 3,
      "replies": [
        {
          "id": 2573145,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2023-12-24T18:43:26.797000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/aliabbasi\" target=\"_blank\">@aliabbasi</a> , thanks for the tip! Your pyvips offline notebook gave me a great starting point with this. However, I don't see anything in that notebook that shows how explicitly use a core to crop a tile. I'm also not sure what you mean by \"you can use python naive concurrent module\". Do you have an example of such a module so I can go learn how to use it?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2574852,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-12-26T09:37:47.297000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/aliabbasi\" target=\"_blank\">@aliabbasi</a> While converting the whole images into tiles I did not get any issue related to the time . Are you converting images from train_images folder or from thumbnail folder?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2574327": "something like:\n\n\n```python\nimport concurrent.futures \n\ndef process_tile(): \n    do tiling stuff...\n    return tiles\n    \nwith concurrent.futures.ThreadPoolExecutor(max_workers=4) as executor:\n    results = list((executor.map(process_tile) )\n```\n\nthis is a dummy example, you need to adapt for your method and process!",
    "2573127": "almost had same issue, \nnow I'm tiling images with pyvips on 4 core of CPUs, each core cropping a tile, and it runs at lease 25% faster. \nyou can use python naive concurrent module ",
    "2573025": "Hi Everyone,\n\nI'm having some difficulty with getting the image generation part of my submission script to run quickly enough. I trained an efficient_net classifier (keras-cv with jax backend) on individual tiles and I'm using pyvips to cut the tiles into a regular grid ( <= 32 tiles per image) for inference, then predicting by using a simple sum of the logits. My NN whips through these at about 5-6 seconds per image (less on the smaller batches), but my total processing time varies widely. Naturally, TMAs don't take that long (~3-4 seconds total processing time), but larger images take ~30-100 seconds, depending on the size.\n\nI'm just trying to get a baseline in, but my last submission timed out. I was hopeful that the majority of the test set was TMAs, but it seems there aren't enough to compensate for the time. With an average 30 seconds of processing time X 2000 images, you end up with about 17 hours of total processing time give or take a bit.\n\nWhat sort of timings are others here seeing and are you willing to share some tips and tricks on how you were able to achieve it? I was hoping to use this competition to learn how to build a model that generalizes well and can perform outlier detection, but I've spent quite a lot of time fighting with submissions and trying to get the most out of pyvips...Any advice would be greatly appreciated."
  }
}