{"cells":[{"cell_type":"markdown","id":"4f998332","metadata":{},"source":"# RSNA Knee 0.937: 22 submissions, 6 lessons\n\nThis project recorded 22 submissions from August 25 through September 5, 2026 (12 calendar dates, inclusive). A public **0.936** baseline was reproduced on August 30; the best subsequent public score was **0.937**. This notebook records which additions failed and examines an expected +0.004 gain that was not observed on the public leaderboard. It does not establish the cause of that forecast error.\n\nThis notebook contains the complete 22-submission score trail, small reproducible calculations, and an explicit separation between inherited public models and my own experiments. It does **not** claim medical validity or a private-leaderboard improvement.\n\n**Correction — September 7, 2026 (Version 2):** the reproduced base is submission **9**, not 12. The M42/43/44 ensemble scored **0.899**, or **+0.009 versus the recorded M44 score of 0.890**; the earlier +0.021 used a different K-series model. The V6 weights and Gold58 replay estimates originated in **prvsiyan's public Notebook**. Dates now convert the API's UTC timestamps to JST. Repeated use of Gold58 is a limitation, not a demonstrated explanation for the forecast error.\n\n**AI assistance:** implementation, execution support, analysis and writing in this project were AI-assisted. The recorded results should not be interpreted as proof of unaided implementation or independent discovery of upstream methods.\n"},{"cell_type":"markdown","id":"f7a75d3f","metadata":{},"source":"## tl;dr\n\n- Successive project-trained configurations reached **0.862 and 0.910**; this is not a single-factor comparison. Adding the 0.910 ensemble to a public 0.936 base produced a displayed delta of **0.000** at weights 0.10 and 0.15.\n- Adding a 0.874 model produced **0.933** at weight 0.20 and **0.935** at weight 0.10. These tested mixtures were harmful despite the expectation of complementary predictions.\n- Renta K.'s public weak-label meniscus route reported **0.937**. Our later Bee V6 submission also scored **0.937**; this ledger has no separate unchanged 0.937-parent submission isolating the specialist's effect.\n- Upstream V6 weights and Gold58 estimates motivated our **0.941** forecast. The observed score was **0.937**, a forecast error of **−0.004**. A Gold58 delta is not a calibrated public-LB prediction.\n- The practical stopping rule: require a production-equivalent, case-level incremental test before spending leaderboard submissions on another blend coefficient.\n"},{"cell_type":"markdown","id":"7916cfbf","metadata":{},"source":"## Context & Methods\n\nThe competition metric is the macro-average ROC AUC over twelve knee-MRI findings. The table below is my complete public-leaderboard history from **2026-08-25 through 2026-09-05 JST**, retrieved from the Kaggle submissions API on 2026-09-07.\n\n### Key assumptions\n\n1. The API snapshot reports scores to three decimals. Equality of displayed scores does not prove equality of unrounded performance; exact sub-0.001 differences cannot be recovered here.\n2. The 58 expert-labelled studies are useful for diagnosis, but they were repeatedly inspected and therefore are not an independent selection set.\n3. A leaderboard pair is treated as a useful comparison only when the model graph and one intended factor are traceable. The full sequence is observational, not a 22-arm randomized experiment.\n"},{"cell_type":"code","execution_count":1,"id":"2c22b511","metadata":{"execution":{"iopub.execute_input":"2026-09-06T19:27:25.511331Z","iopub.status.busy":"2026-09-06T19:27:25.511184Z","iopub.status.idle":"2026-09-06T19:27:26.265269Z","shell.execute_reply":"2026-09-06T19:27:26.264685Z"}},"outputs":[],"source":"import json\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom IPython.display import display\n\nrecords = json.loads(r'''[{\"experiment\": 1, \"submission_ref\": 55763493, \"date\": \"2026-08-25\", \"score\": 0.5, \"description\": \"Plumbing probe v1 - constant 0.5, no model. Measures decode coverage, transfer-syntax mix on the real test set, and wall-clock per study against the 9h budget. Expected LB ~0.5 by design.\"}, {\"experiment\": 2, \"submission_ref\": 55791822, \"date\": \"2026-08-26\", \"score\": 0.862, \"description\": \"v12 best checkpoint (epoch 5, gold58 macro 0.8465). First trained-model submission. Epoch 5 is a +4.4 sigma spike over the plateau 0.8203 +/- 0.0059, so this tests whether best-epoch selection on n=58 transfers.\"}, {\"experiment\": 3, \"submission_ref\": 55791852, \"date\": \"2026-08-26\", \"score\": 0.855, \"description\": \"v12 last checkpoint (epoch 9, gold58 macro 0.8250). Paired with the best-epoch submission to decide 82149809: if best does not beat last on the hidden test, the +4.4 sigma epoch-5 peak was a selection artefact of n=58.\"}, {\"experiment\": 4, \"submission_ref\": 55866896, \"date\": \"2026-08-29\", \"score\": 0.849, \"description\": \"K42 (local-trained, 150mm crop なし, v2 teacher) rsna-knee-infer\"}, {\"experiment\": 5, \"submission_ref\": 55866897, \"date\": \"2026-08-29\", \"score\": 0.856, \"description\": \"K42 (local-trained, 150mm crop なし, v2 teacher) rsna-knee-infer-avg\"}, {\"experiment\": 6, \"submission_ref\": 55870185, \"date\": \"2026-08-30\", \"score\": 0.874, \"description\": \"M42: v3 teacher + 150mm crop + slot-aware selection, best(ep5)\"}, {\"experiment\": 7, \"submission_ref\": 55879031, \"date\": \"2026-08-30\", \"score\": 0.864, \"description\": \"E1: K43 ep5 (same config as K42, seed 43) - LB seed band measurement\"}, {\"experiment\": 8, \"submission_ref\": 55879034, \"date\": \"2026-08-30\", \"score\": 0.878, \"description\": \"E1: K44 ep5 (same config as K42, seed 44) - LB seed band measurement\"}, {\"experiment\": 9, \"submission_ref\": 55879474, \"date\": \"2026-08-30\", \"score\": 0.936, \"description\": \"Control: unmodified fork of public dinosaur-v4 (LB 0.936) - reproduction check\"}, {\"experiment\": 10, \"submission_ref\": 55879814, \"date\": \"2026-08-30\", \"score\": 0.933, \"description\": \"Blend v1: public dinosaur-v4 base (0.936) + our M42 (0.874, decorrelated) rank-average w=0.20\"}, {\"experiment\": 11, \"submission_ref\": 55885991, \"date\": \"2026-08-30\", \"score\": 0.935, \"description\": \"Blend v2: same as v1 but w=0.10 (dose-response: base 0.936, w=0.20 -> 0.933)\"}, {\"experiment\": 12, \"submission_ref\": 55903055, \"date\": \"2026-08-31\", \"score\": 0.89, \"description\": \"M44: bundle (v3 teacher + 150mm crop + slot-aware) on best seed 44. M42=0.874, K44=0.878\"}, {\"experiment\": 13, \"submission_ref\": 55904225, \"date\": \"2026-08-31\", \"score\": 0.897, \"description\": \"Seed ensemble: M42+M44 probability average (members 0.874 / 0.890 solo)\"}, {\"experiment\": 14, \"submission_ref\": 55913163, \"date\": \"2026-08-31\", \"score\": 0.899, \"description\": \"3-seed ensemble M42+43+44 (2-seed was 0.897)\"}, {\"experiment\": 15, \"submission_ref\": 55922160, \"date\": \"2026-09-01\", \"score\": 0.936, \"description\": \"Blend retry: public base 0.936 + 3-seed ensemble 0.899 (was 0.874 solo), rank avg w=0.10. Prior dose-response at gap 0.062: 0.935. Gap now 0.037.\"}, {\"experiment\": 16, \"submission_ref\": 55928396, \"date\": \"2026-09-01\", \"score\": 0.936, \"description\": \"Label-selective blend: LM/PFOA/LatOA only at w=0.20 (uniform w=0.10 was neutral 0.936)\"}, {\"experiment\": 17, \"submission_ref\": 55940691, \"date\": \"2026-09-01\", \"score\": 0.892, \"description\": \"T2 metrology slot: labelwise readout+z single seed42 (dial -0.011 vs gold58 +0.007 vs M42 - settling which local metric tracks LB)\"}, {\"experiment\": 18, \"submission_ref\": 56023037, \"date\": \"2026-09-05\", \"score\": 0.904, \"description\": \"T2D44 efficiency screening single model (sample extrapolation 96 min; final runtime determined by scoring run)\"}, {\"experiment\": 19, \"submission_ref\": 56023245, \"date\": \"2026-09-05\", \"score\": 0.91, \"description\": \"T2 3-seed ensemble: seed42/43/44 ep5 probability mean\"}, {\"experiment\": 20, \"submission_ref\": 56031539, \"date\": \"2026-09-05\", \"score\": 0.936, \"description\": \"T2 ens rank blend w=0.10; gated\"}, {\"experiment\": 21, \"submission_ref\": 56031651, \"date\": \"2026-09-05\", \"score\": 0.936, \"description\": \"T2 ens rank blend w=0.15; gated\"}, {\"experiment\": 22, \"submission_ref\": 56034990, \"date\": \"2026-09-05\", \"score\": 0.937, \"description\": \"Bee V6 fixed outer weights; expected 0.941 from exact V38 nested Gold58 estimate; kernel v1 output sha 2cac0e29\"}]''')\nhistory = pd.DataFrame(records)\nby_ref = history.set_index(\"submission_ref\")\ndef score(ref):\n    return float(by_ref.loc[ref, \"score\"])\n\nbase_ref = 55879474\nbase_score = score(base_ref)\nbase_order = int(by_ref.loc[base_ref, \"experiment\"])\nbee_ref = 56034990\nm44_ref = 55903055\nm_ensemble_ref = 55913163\nassert base_order == 9\nassert history[\"submission_ref\"].is_unique\nassert np.isclose(score(m_ensemble_ref) - score(m44_ref), 0.009)\n\nchecks = {\n    \"submission rows\": len(history),\n    \"unique experiment ids\": int(history[\"experiment\"].nunique()),\n    \"missing scores\": int(history[\"score\"].isna().sum()),\n    \"first score\": float(history.iloc[0][\"score\"]),\n    \"best score\": float(history[\"score\"].max()),\n}\nassert checks == {\n    \"submission rows\": 22,\n    \"unique experiment ids\": 22,\n    \"missing scores\": 0,\n    \"first score\": 0.5,\n    \"best score\": 0.937,\n}\npd.Series(checks, name=\"value\").to_frame()\n"},{"cell_type":"markdown","id":"08a3e3d6","metadata":{},"source":"## Data\n\nEach row below is one completed Kaggle submission from the saved September 7 snapshot. API timestamps were converted from UTC to JST before extracting dates. Descriptions are preserved as historical submission notes, including claims or expectations that may subsequently have proved incorrect; they are not independently verified conclusions. No hidden-test labels, private leaderboard values, or patient data are included.\n"},{"cell_type":"code","execution_count":2,"id":"5284ad36","metadata":{"execution":{"iopub.execute_input":"2026-09-06T19:27:26.267024Z","iopub.status.busy":"2026-09-06T19:27:26.266804Z","iopub.status.idle":"2026-09-06T19:27:26.315362Z","shell.execute_reply":"2026-09-06T19:27:26.314963Z"}},"outputs":[],"source":"display(history[[\"experiment\", \"submission_ref\", \"date\", \"score\", \"description\"]].style.format({\"score\": \"{:.3f}\"}))\n"},{"cell_type":"markdown","id":"773f5de5","metadata":{},"source":"## Results\n\n### 1. The apparent jump is mostly a change of starting point\n\nThe 0.500 submission was a plumbing probe. Experiments 2–8 are project-trained model and inference variants. Experiment **9** reproduces an external public 0.936 system; experiments 10–11 test M42 blends. Later submissions alternate between standalone models, ensembles and additions to the stronger base. The historical order alone does not isolate a training improvement.\n"},{"cell_type":"code","execution_count":3,"id":"ae759dc4","metadata":{"execution":{"iopub.execute_input":"2026-09-06T19:27:26.316699Z","iopub.status.busy":"2026-09-06T19:27:26.316529Z","iopub.status.idle":"2026-09-06T19:27:26.438925Z","shell.execute_reply":"2026-09-06T19:27:26.438361Z"}},"outputs":[],"source":"fig, ax = plt.subplots(figsize=(11, 4.8))\nax.plot(history[\"experiment\"], history[\"score\"], marker=\"o\", lw=2, color=\"#315A8A\")\nax.axhline(base_score, color=\"#B24C3E\", ls=\"--\", lw=1.5, label=f\"reproduced public base: {base_score:.3f}\")\nax.scatter([base_order, int(by_ref.loc[bee_ref, \"experiment\"])], [base_score, score(bee_ref)], s=90, color=[\"#B24C3E\", \"#2B7A4B\"], zorder=3)\nax.annotate(\"public base reproduced\", (base_order, base_score), xytext=(4.5, 0.919), arrowprops={\"arrowstyle\": \"->\"})\nax.annotate(\"best: 0.937\", (22, 0.937), xytext=(18.0, 0.916), arrowprops={\"arrowstyle\": \"->\"})\nax.set(title=\"Public LB history: self-trained models improved, but did not lift the strong base\",\n       xlabel=\"submission order\", ylabel=\"public macro AUC\")\nax.set_xticks(range(1, 23))\nax.set_ylim(0.82, 0.945)\nax.grid(axis=\"y\", alpha=0.25)\nax.legend(loc=\"lower right\")\nplt.tight_layout()\nplt.show()\n"},{"cell_type":"markdown","id":"10aa4e64","metadata":{},"source":"The plumbing probe is omitted from the vertical range so the differences among useful models remain visible. The exact values remain in the table.\n\n### 2. Six decision units are more informative than 22 isolated scores\n"},{"cell_type":"code","execution_count":4,"id":"7903b8ba","metadata":{"execution":{"iopub.execute_input":"2026-09-06T19:27:26.440552Z","iopub.status.busy":"2026-09-06T19:27:26.440334Z","iopub.status.idle":"2026-09-06T19:27:26.446393Z","shell.execute_reply":"2026-09-06T19:27:26.44597Z"}},"outputs":[],"source":"decisions = pd.DataFrame([\n    [\"Pipeline\", \"constant 0.5\", score(55763493), \"Completed scoring; no model-quality claim\"],\n    [\"First trained model\", \"best vs last checkpoint\", score(55791822), f\"Displayed best-minus-last difference: {score(55791822)-score(55791852):+.3f}\"],\n    [\"Seed sensitivity\", \"K43 vs K44\", score(55879034), f\"Two recorded K variants: {score(55879031):.3f} and {score(55879034):.3f}; K42 inference variants are not pooled here\"],\n    [\"Ensembling\", \"M42/43/44 mean vs recorded M44\", score(m_ensemble_ref), f\"{score(m_ensemble_ref)-score(m44_ref):+.3f} vs M44; no standalone M43 score in this ledger, so no best-member claim\"],\n    [\"Second project model\", \"T2 three-seed ensemble\", score(56023245), \"Higher standalone score; tested blends matched displayed 0.936\"],\n    [\"Public-derived candidate\", \"Bee V6\", score(bee_ref), \"Matched upstream reported 0.937 parent; forecast 0.941 was not reached\"],\n], columns=[\"decision\", \"comparison\", \"reference_public_score\", \"observed lesson\"])\ndisplay(decisions.style.format({\"reference_public_score\": \"{:.3f}\"}))\n"},{"cell_type":"markdown","id":"16e543f3","metadata":{},"source":"### 3. A decorrelated weak model can still be harmful\n\nThe same reproduced base scored 0.936. Mixing a 0.874 component by rank average gave 0.933 at weight 0.20 and 0.935 at 0.10. After improving the component to 0.899 and then 0.910, a 0.10–0.15 blend became neutral at displayed precision, but never positive.\n"},{"cell_type":"code","execution_count":5,"id":"ac142a46","metadata":{"execution":{"iopub.execute_input":"2026-09-06T19:27:26.44763Z","iopub.status.busy":"2026-09-06T19:27:26.447508Z","iopub.status.idle":"2026-09-06T19:27:26.531108Z","shell.execute_reply":"2026-09-06T19:27:26.530718Z"}},"outputs":[],"source":"blend_tests = pd.DataFrame([\n    [\"M42 single seed\", score(55870185), 0.20, score(55879814)],\n    [\"M42 single seed\", score(55870185), 0.10, score(55885991)],\n    [\"M42/43/44 ensemble\", score(m_ensemble_ref), 0.10, score(55922160)],\n    [\"T2 3-seed ensemble\", score(56023245), 0.10, score(56031539)],\n    [\"T2 3-seed ensemble\", score(56023245), 0.15, score(56031651)],\n], columns=[\"added component\", \"component_score\", \"blend_weight\", \"blend_score\"])\nblend_tests[\"delta_vs_0.936\"] = blend_tests[\"blend_score\"] - base_score\ndisplay(blend_tests.style.format({\n    \"component_score\": \"{:.3f}\", \"blend_weight\": \"{:.2f}\",\n    \"blend_score\": \"{:.3f}\", \"delta_vs_0.936\": \"{:+.3f}\",\n}))\n\nfig, ax = plt.subplots(figsize=(8.8, 4.6))\ncolors = {\"M42 single seed\": \"#B24C3E\", \"M42/43/44 ensemble\": \"#C68B2C\", \"T2 3-seed ensemble\": \"#315A8A\"}\nfor name, group in blend_tests.groupby(\"added component\", sort=False):\n    ax.scatter(group[\"blend_weight\"], group[\"delta_vs_0.936\"] * 1000,\n               s=90, label=f\"{name} (solo {group.iloc[0]['component_score']:.3f})\", color=colors[name])\nax.axhline(0, color=\"black\", lw=1)\nax.set(title=\"Incremental value to the 0.936 base: stronger components became neutral, not positive\",\n       xlabel=\"rank-blend weight\", ylabel=\"displayed delta vs base (×0.001 AUC)\")\nax.set_xticks([0.10, 0.15, 0.20])\nax.grid(alpha=0.25)\nax.legend(frameon=False)\nplt.tight_layout()\nplt.show()\n"},{"cell_type":"markdown","id":"8797fa7c","metadata":{},"source":"### 4. An upstream Gold58 estimate did not predict our public score\n\n**Gold58** denotes the 58 competition studies with expert labels, used here as an exploratory evaluation set. It has been repeatedly inspected for model selection and is not a fresh holdout.\n\nThe five outer weights and exact V38 Gold58 replay came from [prvsiyan's Bee Notebook](https://www.kaggle.com/code/prvsiyan/the-bee-s-knees-final-rsna-push). Its source reported **0.941932 → 0.947966** on Gold58 and an estimated nested gain of about **+0.004**. These are upstream reports, not an independently reproduced exact-V38 replay in this project. Our team used that estimate to forecast **0.941** from the source-reported **0.937** parent, replaced a private dependency with a matching public checkpoint, removed the sample-only V7 graft, and submitted V6. Our completed submission scored **0.937**.\n\nThe local delta was not a calibrated forecast of the public leaderboard. Gold58 reuse, differences between evaluation and deployed components, sampling variation and distribution differences are possible explanations; this single aggregate result does not identify their contributions. Score rounding also limits the comparison. The ledger has no separate submission of the unchanged 0.937 parent, so the table does not claim a controlled zero effect of V6.\n"},{"cell_type":"code","execution_count":6,"id":"1bf3693e","metadata":{"execution":{"iopub.execute_input":"2026-09-06T19:27:26.532872Z","iopub.status.busy":"2026-09-06T19:27:26.532643Z","iopub.status.idle":"2026-09-06T19:27:26.539088Z","shell.execute_reply":"2026-09-06T19:27:26.538719Z"}},"outputs":[],"source":"prediction_check = pd.DataFrame([\n    [\"our Bee V6 submission (upstream weights)\", 0.941, score(bee_ref)],\n], columns=[\"candidate\", \"expected_public_score\", \"observed_public_score\"])\nprediction_check[\"error\"] = prediction_check[\"observed_public_score\"] - prediction_check[\"expected_public_score\"]\ndisplay(prediction_check.style.format({\n    \"expected_public_score\": \"{:.3f}\", \"observed_public_score\": \"{:.3f}\", \"error\": \"{:+.3f}\"\n}))\nassert np.isclose(prediction_check.loc[0, \"error\"], -0.004)\n"},{"cell_type":"markdown","id":"433ee24b","metadata":{},"source":"## What is inherited and what is original\n\n| Part | Provenance | Claim made here |\n|---|---|---|\n| 0.936 multi-arm base | Public notebooks by Pilkwang Kim, prvsiyan, Roman Tamrazov and collaborators | Reproduced the displayed public score as a control; not my model |\n| 0.937 meniscus specialist | Public weak-label DINOv2 residual by Renta K. | Used as the stronger public parent; not my trained specialist |\n| K42/M42/T2 models | Project-specific training and submissions with AI assistance | Standalone scores and documented blend comparisons; not a claim of unaided implementation |\n| Five outer-weight changes and Gold58 estimates | prvsiyan's Bee V6 public source | Our dependency replacement, V7 removal and submission audit; no claim to have originated these weights or estimates |\n| This notebook | AI-assisted project analysis and writing | Score trail, corrected comparisons, limitations and stopping rule |\n\nThe distinction matters: a high score inherited from public code is a starting point. The contribution being offered is evidence about **incremental value**, including failed additions.\n\n## Takeaways\n\n1. **Reproduce the strong parent first.** Experiment 9 supplied the displayed 0.936 control for the later blend comparisons.\n2. **Measure a component alone and on top of the same parent.** A move from 0.874 to 0.910 alone did not imply positive ensemble value.\n3. **Do not spend precision you do not have.** Reusing 58 studies can diagnose obvious failures; it cannot reliably choose among +0.004 candidates after repeated inspection.\n4. **Publish negative dose-response.** The 0.20 → 0.10 blend pair prevented further blind tuning of the same weak component.\n5. **Require case-level production parity.** The next useful specialist needs predictions from the exact deployed graph on an independent cohort, not a substitute proxy.\n6. **Stop when the experiment no longer identifies a cause.** More blend coefficients would consume submissions without separating teacher quality, image coverage and validation reuse.\n\n## Limitations\n\n- All outcomes are public-leaderboard scores; private-leaderboard behavior is unknown.\n- Scores are rounded to three decimals and cannot resolve micro-gains.\n- Most rows are sequential research decisions, not controlled single-factor experiments.\n- The 58 expert studies were repeatedly reused, so their estimates are exploratory.\n- The upstream 0.937 parent score and exact V38 replay are source-reported. Our 22-row ledger does not independently establish a specialist-only or V6-only gain over that parent.\n- Error explanations are hypotheses; no causal attribution is established by a single aggregate LB result.\n- This is a model-selection audit, not evidence of clinical diagnostic accuracy.\n\n## Public sources and attribution\n\n- [Pilkwang Kim — RSNA Knee baseline v1](https://www.kaggle.com/code/pilkwang/rsna-knee-baseline-v1)\n- [prvsiyan — Head and shoulders, knees and toes](https://www.kaggle.com/code/prvsiyan/head-and-shoulders-knees-and-toes)\n- [Roman Tamrazov — RSNA Knee DINOsaur V4](https://www.kaggle.com/code/romantamrazov/rsna-knee-dinosaur-v4)\n- [Renta K. — 0.937 weak-label DINOv2 meniscus residual](https://www.kaggle.com/code/renta0426/rsna-knee-0-937-weak-label-dinov2-meniscus-resid)\n- [prvsiyan — The bee's knees: final RSNA push (V6 weights and Gold58 estimates)](https://www.kaggle.com/code/prvsiyan/the-bee-s-knees-final-rsna-push)\n- [RSNA Knee Abnormality Detection](https://www.kaggle.com/competitions/rsna-knee-abnormality-detection)\n\nThe calculations can be rerun on CPU without network access. The table preserves the original submission descriptions; the Version 2 correction above supersedes any conflicting interpretation.\n"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.13.9"}},"nbformat":4,"nbformat_minor":5}