{
  "topic": {
    "id": 737696,
    "title": "RSNA Raptor Weights",
    "authorName": "Dread Development",
    "commentCount": 14,
    "votes": 18,
    "postDate": "2026-08-26T23:40:02.832000"
  },
  "comments": [
    {
      "id": 3520983,
      "authorName": "Jack",
      "votes": 1,
      "postDate": "2026-09-04T14:51:27.547000",
      "content": "<p>A couple questions..</p>\n<ol>\n<li>Is your 0.936-0.938 model trained on the same labels as the 0.924?</li>\n<li>Is your 0.936-0.938 model trained on the same datasets as the 0.924?</li>\n<li>Is your 0.936-0.938 model using the same inference pipeline as the 0.924 or did you make improvements to it?</li>\n<li>Does your SWA checkpoint from the 0.924 model set score better or worse than best epoch model? I'd answer this myself but I'm out of submissions from testing a bunch of other things lol</li>\n</ol>\n<p>I've been testing a bunch of things with your previous SWA checkpoint and have it scoring well and quite efficiently. Your work has been insightful and a good starting point for this competition. Although I'd happily mess around with whatever new weights you release, don't feel pressured to release more weights, especially if they're your best or close to your best. This is a competition after all 😅</p>"
    },
    {
      "id": 3521127,
      "authorName": "Dread Development",
      "votes": 4,
      "postDate": "2026-09-04T23:23:01.233000",
      "content": "<ol>\n<li><p>Labels - yes, every checkpoint published  is trained on the same label set with the only difference being the corpus geometry &amp; window count.</p></li>\n<li><p>Dataset - Same comp data, nothing external with the only change in preprocessing. The 0.924 set is 64 slices,140mm crop, span 0.06-0.94, 336px. Later ones widen the span, go to 80-96 slices or move to 384 px.</p></li>\n<li><p>Pipeline - Same architecture and same windowing code</p></li>\n<li><p>SWA - slightly worse. On our 58 gold labeled studies the final checkpoint reads 0.9167 and the SWA one 0.9150. Across six SWA/non-SWA pairs it averages out to nothing, and on the board an SWA arm scored 0.937 where its non-SWA twin scored 0.938.</p></li>\n</ol>\n<p>I have not been able to make nearly as much progress in the last two weeks due to the OpenAI  security comp consuming all of my attention, however as that has concluded I am back to focusing on RSNA so I expect to gain headway this weekend.\nI do … want very much so to release more raptor weights &amp; I do have quite a few to release even as of right now, I just dont want a repeat of my last drop.\nI had made a few datasets public as well as the weights &amp; I watched the entire board shift.. once I can asses where I stand I will release them, just being careful in doing so.</p>\n<p>I am also trying to get back on track for Biohub, Solar filament, UMUD &amp; another small lang model thats wrapping up so I am still diving time between several comps though these are no where near the headache the OpenAI  Security comp was so expect the release sooner (maybe sunday)😀</p>"
    },
    {
      "id": 3521141,
      "authorName": "Jack",
      "votes": 1,
      "postDate": "2026-09-05T00:00:32.870000",
      "content": "<p>Thanks for the response. I've got the solo swa checkpoint from <a href=\"https://www.kaggle.com/datasets/dreaddevelopment/raptor-knee-widedense\" target=\"_blank\">https://www.kaggle.com/datasets/dreaddevelopment/raptor-knee-widedense</a> scoring 0.930 - you're telling me this was worse than 0.924? Or am I confused about your datasets and notebooks?</p>"
    },
    {
      "id": 3521142,
      "authorName": "Dread Development",
      "votes": 0,
      "postDate": "2026-09-05T00:44:57.570000",
      "content": "<p>You're not confused about the datasets, 0.924 isnt a property of the checkpoint but rather how many windows it was scored with. That number was the corpus at 42 windows per study.\nThe identical weights at 62 windows scored 0.927 on our board so window count alone was worth +0.003 &amp; full coverage is higher still. A solo SWA checkpoint reading 0.930 is consistent with that - You're just scoring more of each study than my early run did.\nI have another checkpoint public already that pushes those exact weights to 0.932 If im not mistaken.</p>\n<p>I am scoring a few subs right now and expect them back soon, I've made some tweaks &amp; am not entirely sure but very optimistic about them.</p>\n<ul>\n<li>I have 7 in flight scoring right now</li>\n</ul>"
    },
    {
      "id": 3521143,
      "authorName": "Dread Development",
      "votes": 1,
      "postDate": "2026-09-05T00:46:21.333000",
      "content": "<p>This is the solo SWA that scored 0.932 - <a href=\"https://www.kaggle.com/datasets/dreaddevelopment/raptor-knee-finespacing\" target=\"_blank\">https://www.kaggle.com/datasets/dreaddevelopment/raptor-knee-finespacing</a></p>"
    },
    {
      "id": 3517670,
      "authorName": "Cody_Null",
      "votes": 3,
      "postDate": "2026-08-27T20:10:23.553000",
      "content": "<p>Is this saying you have a 5 fold raptor sub with .936-.938 or is that a larger ensemble? Thanks for sharing in any case!</p>"
    },
    {
      "id": 3517697,
      "authorName": "Dread Development",
      "votes": 4,
      "postDate": "2026-08-27T22:21:11.370000",
      "content": "<p>Great question Cody! So neither &amp; that may not of been what you were expecting.\nits a single model, not a 5 + - fold, one CoAtNet 384 run over the full labelled set with a small fully labelled holdout to pick the checkpoint.\nA couple of the subs average the best three epochs of that same run, but there is no fold ensembeling in it.</p>\n<p>I will release updated weights as soon as possible, I am just really crunched for time at the moment as I have another competition wrapping up (OpenAI Security) &amp; I would rather not finish sub 20 on the public board so I have diverted all of my attention there until end of that comp.\nI hope this answered your question!!!</p>"
    },
    {
      "id": 3517769,
      "authorName": "Cody_Null",
      "votes": 1,
      "postDate": "2026-08-28T05:57:31.773000",
      "content": "<p>Very cool! Thanks for sharing, interested in seeing your next update!</p>"
    },
    {
      "id": 3518136,
      "authorName": "nguyen214",
      "votes": 1,
      "postDate": "2026-08-29T16:09:12.673000",
      "content": "<p>May I ask about the inference time of your pipeline? I think your leaderboard score is quite strong, but I didn’t see your name on the efficiency leaderboard, so I was curious about the runtime.</p>"
    },
    {
      "id": 3518170,
      "authorName": "Dread Development",
      "votes": 0,
      "postDate": "2026-08-29T18:18:40.450000",
      "content": "<p>Sure, its around 9 hours for the full test set.\nAlmost all of that is the window count rather than models. We score 94 windows per study, so our CoAtNet ends up costing about as much as the entire 5 fold dino half we blend it with.\nwe did run a trimmed version a few days back, 24 windows on one T4 scored 0.02 and 42 windows got 0.930, so the cheap end is only about 0.01 behind.\nJust havent selected one of those for the effiecienc side.</p>"
    },
    {
      "id": 3521688,
      "authorName": "Dread Development",
      "votes": 1,
      "postDate": "2026-09-06T19:05:16.017000",
      "content": "<p>Wanted to update you on the timing - I have dropped scoring time to under 2 hours now &amp; am continuing to work both the timing &amp; accuracy in parallel.</p>\n<p>Update 9/7/2026 at 11:51 EST - Time down from 2 hours to 29 minutes</p>"
    },
    {
      "id": 3524672,
      "authorName": "Ziad Ahmed",
      "votes": 0,
      "postDate": "2026-09-15T15:24:11.287000",
      "content": "<p>Thanks for publishing these, and for the geometry table on the dataset pages - it's what makes them usable without burning submissions to guess. We run maxspan-v5 and native384dense-v10 as a two-arm blend over a shared decode.</p>\n<p>A question about the one page that omits the geometry. Nine of the ten Raptor datasets state slices / span / crop / stored px; raptor-knee-finespacing says only \"the Fine Spacing preprocessing of the corpus\". Given your note above that the later corpora \"go to 80-96 slices\", I'd rather not feed v9 a volume it never saw:</p>\n<ol>\n<li><p>What slice count and span does the Fine Spacing corpus use?</p></li>\n<li><p>Was the 0.932 scored at your 94 windows per study, or at a lower count?</p></li>\n</ol>\n<p>The reason for (2) is your own fullspan result: widening the span at a fixed 44 slices cost 0.009, with the conclusion that widening only pays if slices are added to keep the sampling dense. If Fine Spacing is a 96-slice corpus, then scoring it on a 64-slot volume at 62 windows is that same density mismatch in the same direction - so the 0.932 wouldn't be available at 62 windows, and the checkpoint swap and the window count would be two different levers rather than one.</p>\n<p>Also, if you happen to remember: \"42 windows got 0.930\" for the trimmed run - was that on the Fine Spacing corpus or the 64-slice one?</p>\n<p>Happy to report back what we measure either way.</p>"
    },
    {
      "id": 3524687,
      "authorName": "Dread Development",
      "votes": 0,
      "postDate": "2026-09-15T16:05:36.510000",
      "content": "<p>Good catch. That page went up without the geometry line and I've just fixed it, so it's on the dataset now along with a note about window counts.\nFine Spacing is 80 slices per study across 2 to 98 percent of each series, a 140 mm crop, stored at 336 pixels. The slots are 22 sagittal fluid, 18 sagittal, 15 coronal fluid, 10 coronal, 15 axial. Full window coverage is 78.\nThe 0.932 was scored at 78, so no, it isn't available at 62, and your reasoning about why is right. Same effect as the full span result: 44 slices over 2 to 98 percent scored 0.917 against 0.926 for 44 over 6 to 94, so spreading the same slices thinner cost 0.009. Putting 64 slices over the wide span brought it to 0.928, and 80 slices to 0.932. Density is the lever rather than span, and 1.20 percent of the stack per slice is the finest I've measured and also the best. Scoring an 80 slice volume at 62 windows leaves about a fifth of it unread, and at that point the checkpoint swap and the coverage change stop being separable, which is exactly what you said.\nOn the 42 window 0.930, I think two numbers from this thread have got stuck together. My 42 window results top out at 0.926. The 0.924 is the 64 slice corpus at 42 windows, and the identical weights read 0.927 at 62. The 0.930 is Jack's, further up this thread, from running the widedense SWA checkpoint with more coverage than my original run used.\nPlease do report back what you get, I'd be interested either way.</p>"
    },
    {
      "id": 3525865,
      "authorName": "Ziad Ahmed",
      "votes": 0,
      "postDate": "2026-09-19T08:11:02.677000",
      "content": "<p>Reporting back as promised. Your density curve doesn't transfer to our pipeline, and I think the reason is instructive.\nWe tested it at fixed everything-else on one fold: 8 slices/slot → 10 slices/slot, which on our 0.12–0.88 band moves us from 1.58% to 1.27% of the stack per slice into the range you measured as best. It came back −0.0045 (plateau −0.0046) against a 0.003 noise threshold, so a real, small negative. We also tried tightening the crop 140 → 110 mm at the denser sampling: +0.0005, inside noise, unresolved.\nThe difference is probably structural rather than a contradiction. You sample 80 slices across whole series and score every window position, so extra slices buy genuinely new coverage. We sample 6 per-series slots into a masked attention head that pools 48 tokens, so extra slices mostly dilute the softmax rather than adding field of view the head already sees the whole span, just coarsely. Density is a coverage lever for you and a pooling-composition change for us.\nTwo other one-variable results from the same week, in case they're useful: swapping RadImageNet for ImageNet ResNet-50 at identical architecture and correct per-corpus normalisation cost us −0.063, and replacing our attention head with masked mean pooling cost −0.014. So for this pipeline the medical pretraining and the learned pooling are both doing real work.\nThanks again for publishing the geometry it saved us from spending a submission finding out that 0.932 wasn't available at 62 windows.</p>"
    }
  ],
  "index": {
    "id": "737696",
    "title": "RSNA Raptor Weights",
    "authorName": "",
    "commentCount": "14",
    "votes": "18",
    "postDate": "2026-08-26 23:40:02.832000"
  },
  "competition": "rsna-knee-abnormality-detection"
}