The headline isn't "which model wins" — it's what the wrapper does. Testing the same model, same case, with and without an agent scaffold:
Same anonymized case, graded against the same guideline set. "Wrapper" = how the model was used.
| # | Model | Wrapper | Best score /100 | Key note |
|---|---|---|---|---|
| 1 | DeepSeek V4 Pro | WorkBuddy agent | 50 | Only model to catch the staging trap (pT1aN0M0 IA) |
| 2 | Kimi K2.6 | WorkBuddy agent | 46 | Stability flag: refused on first attempt in 1/3 runs |
| 3 | GLM 5.2 (Zhipu) | WorkBuddy agent | 44 | Largest floor lift (+39) — see wrapper effect |
| 4 | Doubao | web | 41 | Strong web baseline without an agent |
| 5 | Kimi K3 | WorkBuddy agent | 40 | Re-scored 27 → 40 after a published correction |
| 6 | DeepSeek V4 Flash | agent / web | 34 | Web 28 → agent 34 (+6) |
| 7 | DeepSeek (web, fast tier) | web | 28 | — |
| 8 | ChatGLM / GLM (web) | web | 5 | Baseline floor — wrap it in an agent before trusting |
| Model | Web /100 | Agent /100 | Wrapper Δ | Read |
|---|---|---|---|---|
| GLM 5.2 | 5 | 44 | +39 | Scaffolding turns a non-answer into a usable one |
| Kimi K2.6 | 37 | 46 | +9 | Already coherent; agent adds polish + a cited guideline |
| DeepSeek V4 Flash | 28 | 34 | +6 | Small lift — web baseline already moderate |
Only models tested in both wrappers are listed (series rule). Full web-only runs for Kimi K3 and DeepSeek V4 Flash are pending and will extend this grid.
A clinician's quick routing. All scores are on the same case; your mileage varies with the prompt and the guideline version.
Subscribe and I'll send the one-page PDF cheat card (which free model for which clinical task, updated as the leaderboard grows). Free, roughly monthly, unsubscribe anytime.
Powered by Buttondown. Your email is never sold and never changes my independent conclusions.
One anonymized post-op breast-cancer case; identical prompt; a 16-item / 100-point rubric; reference standard = 2026 CBCS / CSCO 2024 / NCCN 2025. Every run is signed and reproducible. The case image, prompt, and score sheet are published in the Benchmark Data appendix so anyone can re-score any model. See the full Model Reviews series (V1 → V4) for transcripts.
This leaderboard is meant to grow beyond my own runs. You can add any free model you've tested — no code, no GitHub account needed.
Everything you need is already published:
Clicking opens your email app pre-filled to jstz1983@163.com with your answers. Please attach 1–2 screenshots of the model's answer so I can re-score and verify before publishing. Nothing is posted automatically — I review every submission.