{
  "id": 611778,
  "title": "predict() signature in the submission notebook limits parallel processing - why not just mounting the test data?",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/611778",
  "author_name": "Stefan Denner",
  "post_date": "2025-10-14T13:06:59.801000",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi, <br>\nThroughout the challenge we were struggling a lot with the 12h time limit. Apart from the problems now towards the end with the high server load, the main issue was the strict predict() implementation in the submission notebook.<br>\nThis interface enforces sequential processing of the entire series end-to-end.</p>\n<p>We wanted to split the test set across two GPUs (each GPU processes a disjoint subset of series), but the required predict() signature and its invocation pattern prevented this.</p>\n<p>We eventually used both GPUs by splitting patches across devices inside a single process, but this was awkward, added a lot of boilerplate, and still blocked parallel data loading / preprocessing / model inference at the series level.</p>\n<p>Practically, this design reduced our achievable throughput and contributed to timeouts for otherwise working solutions.</p>\n<p>I may be missing the rationale here, but I don’t understand why the interface mandates a single predict() that streams series one-by-one. What’s the advantage of enforcing that signature instead of giving participants freedom to orchestrate test-time execution (e.g., multi-process, multi-GPU, series sharding), as long as the outputs and timing/resource constraints are respected?</p>\n<p>Best, <br>\nStefan</p>",
  "messages": [
    {
      "id": 3301897,
      "postDate": "2025-10-14T13:06:59.800Z",
      "content": "<p>Hi, <br>\nThroughout the challenge we were struggling a lot with the 12h time limit. Apart from the problems now towards the end with the high server load, the main issue was the strict predict() implementation in the submission notebook.<br>\nThis interface enforces sequential processing of the entire series end-to-end.</p>\n<p>We wanted to split the test set across two GPUs (each GPU processes a disjoint subset of series), but the required predict() signature and its invocation pattern prevented this.</p>\n<p>We eventually used both GPUs by splitting patches across devices inside a single process, but this was awkward, added a lot of boilerplate, and still blocked parallel data loading / preprocessing / model inference at the series level.</p>\n<p>Practically, this design reduced our achievable throughput and contributed to timeouts for otherwise working solutions.</p>\n<p>I may be missing the rationale here, but I don’t understand why the interface mandates a single predict() that streams series one-by-one. What’s the advantage of enforcing that signature instead of giving participants freedom to orchestrate test-time execution (e.g., multi-process, multi-GPU, series sharding), as long as the outputs and timing/resource constraints are respected?</p>\n<p>Best, <br>\nStefan</p>",
      "rawMarkdown": "Hi, \nThroughout the challenge we were struggling a lot with the 12h time limit. Apart from the problems now towards the end with the high server load, the main issue was the strict predict() implementation in the submission notebook.\nThis interface enforces sequential processing of the entire series end-to-end.\n\nWe wanted to split the test set across two GPUs (each GPU processes a disjoint subset of series), but the required predict() signature and its invocation pattern prevented this.\n\nWe eventually used both GPUs by splitting patches across devices inside a single process, but this was awkward, added a lot of boilerplate, and still blocked parallel data loading / preprocessing / model inference at the series level.\n\nPractically, this design reduced our achievable throughput and contributed to timeouts for otherwise working solutions.\n\nI may be missing the rationale here, but I don’t understand why the interface mandates a single predict() that streams series one-by-one. What’s the advantage of enforcing that signature instead of giving participants freedom to orchestrate test-time execution (e.g., multi-process, multi-GPU, series sharding), as long as the outputs and timing/resource constraints are respected?\n\nBest, \nStefan",
      "votes": 6
    },
    {
      "id": 3302793,
      "postDate": "2025-10-16T14:18:14.733Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/st3v3d\" target=\"_blank\">@st3v3d</a>,</p>\n<p>For most competitions, we provide the entire test set at once so that you are free to orchestrate the test-time execution as you describe. For this competition, there were strong correlations among groups of series that, if fully exploited, would have made the task much easier. The desire however was for each series to be treated independently, motivated both by the application and also the need for a larger effective sample size.</p>\n<p>I agree though that with a dual GPU setup enabling more parallelism would have been helpful. It's possible we could have streamed series in batches of four, say, and still achieved our goals. I appreciate the feedback, and we'll consider it in the design of future competitions.</p>\n<p>And congrats on the gold!</p>",
      "rawMarkdown": "Hi @st3v3d,\n\nFor most competitions, we provide the entire test set at once so that you are free to orchestrate the test-time execution as you describe. For this competition, there were strong correlations among groups of series that, if fully exploited, would have made the task much easier. The desire however was for each series to be treated independently, motivated both by the application and also the need for a larger effective sample size.\n\nI agree though that with a dual GPU setup enabling more parallelism would have been helpful. It's possible we could have streamed series in batches of four, say, and still achieved our goals. I appreciate the feedback, and we'll consider it in the design of future competitions.\n\nAnd congrats on the gold!",
      "votes": 2,
      "replies": [
        {
          "id": 3302857,
          "postDate": "2025-10-16T15:47:35.213Z",
          "content": "<p>Thanks so much for the detailed explanation. Makes perfect sense! And indeed, batching would have helped us already quite a bit :) </p>",
          "rawMarkdown": "Thanks so much for the detailed explanation. Makes perfect sense! And indeed, batching would have helped us already quite a bit :) "
        }
      ]
    },
    {
      "id": 3301903,
      "postDate": "2025-10-14T13:14:58.460Z",
      "content": "<p>I second this. It's quite the bottleneck on how submissions are processed, especially when there's multiple models that require different processing for inputs. </p>",
      "rawMarkdown": "I second this. It's quite the bottleneck on how submissions are processed, especially when there's multiple models that require different processing for inputs. ",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 3302793,
      "author_name": "Ryan Holbrook",
      "author_url": "",
      "post_date": "2025-10-16T14:18:14.733000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/st3v3d\" target=\"_blank\">@st3v3d</a>,</p>\n<p>For most competitions, we provide the entire test set at once so that you are free to orchestrate the test-time execution as you describe. For this competition, there were strong correlations among groups of series that, if fully exploited, would have made the task much easier. The desire however was for each series to be treated independently, motivated both by the application and also the need for a larger effective sample size.</p>\n<p>I agree though that with a dual GPU setup enabling more parallelism would have been helpful. It's possible we could have streamed series in batches of four, say, and still achieved our goals. I appreciate the feedback, and we'll consider it in the design of future competitions.</p>\n<p>And congrats on the gold!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3302857,
          "author_name": "Stefan Denner",
          "author_url": "",
          "post_date": "2025-10-16T15:47:35.213000",
          "content": "<p>Thanks so much for the detailed explanation. Makes perfect sense! And indeed, batching would have helped us already quite a bit :) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3301903,
      "author_name": "Satwik",
      "author_url": "",
      "post_date": "2025-10-14T13:14:58.460000",
      "content": "<p>I second this. It's quite the bottleneck on how submissions are processed, especially when there's multiple models that require different processing for inputs. </p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3301897": "Hi, \nThroughout the challenge we were struggling a lot with the 12h time limit. Apart from the problems now towards the end with the high server load, the main issue was the strict predict() implementation in the submission notebook.\nThis interface enforces sequential processing of the entire series end-to-end.\n\nWe wanted to split the test set across two GPUs (each GPU processes a disjoint subset of series), but the required predict() signature and its invocation pattern prevented this.\n\nWe eventually used both GPUs by splitting patches across devices inside a single process, but this was awkward, added a lot of boilerplate, and still blocked parallel data loading / preprocessing / model inference at the series level.\n\nPractically, this design reduced our achievable throughput and contributed to timeouts for otherwise working solutions.\n\nI may be missing the rationale here, but I don’t understand why the interface mandates a single predict() that streams series one-by-one. What’s the advantage of enforcing that signature instead of giving participants freedom to orchestrate test-time execution (e.g., multi-process, multi-GPU, series sharding), as long as the outputs and timing/resource constraints are respected?\n\nBest, \nStefan",
    "3302793": "Hi @st3v3d,\n\nFor most competitions, we provide the entire test set at once so that you are free to orchestrate the test-time execution as you describe. For this competition, there were strong correlations among groups of series that, if fully exploited, would have made the task much easier. The desire however was for each series to be treated independently, motivated both by the application and also the need for a larger effective sample size.\n\nI agree though that with a dual GPU setup enabling more parallelism would have been helpful. It's possible we could have streamed series in batches of four, say, and still achieved our goals. I appreciate the feedback, and we'll consider it in the design of future competitions.\n\nAnd congrats on the gold!",
    "3301903": "I second this. It's quite the bottleneck on how submissions are processed, especially when there's multiple models that require different processing for inputs. "
  }
}