⚠️ Affiliate Disclosure / 利益披露: This page may contain affiliate or promotional links. They never increase your price and never change my independent conclusions. Details

Free-Model Clinical LLM Leaderboard

One anonymized post-op breast-cancer case · same prompt · same 16-item / 100-point rubric · 2026 CBCS / CSCO 2024 / NCCN 2025.
Not a marketing leaderboard — a surgeon's reproducible one.
Last updated 2026-07-31 · updated monthly
By Tan Haosheng, MD/PhD · 副主任医师 (Associate Chief Physician) · 甲状腺乳腺外科 (Thyroid & Breast Surgery), 泰州市人民医院 (Taizhou People's Hospital). Scores are signed and reproducible; see the Benchmark Data appendix.

The one finding that matters

The headline isn't "which model wins" — it's what the wrapper does. Testing the same model, same case, with and without an agent scaffold:

Agent wrapper raises the floor, not the ceiling. GLM 5.2 went from 5 → 44 (+39), DeepSeek V4 Flash 28 → 34 (+6), Kimi K2.6 37 → 46 (+9). The higher a model's web baseline, the smaller the wrapper lift. A model that flounders alone gains enormously from scaffolding; a model already coherent gains polish.
Bar chart of agent-wrapper lift in points: GLM 5.2 +39, Kimi K2.6 +9, DeepSeek V4 Flash +6. The agent raises the floor, not the ceiling — higher web baseline means smaller wrapper gain. Grouped bar chart: the same anonymized breast-cancer case scored 0-100 with a 16-item rubric vs 2026 CBCS. GLM 5.2 web 5 vs agent 44; DeepSeek V4 Flash web 28 vs agent 34; Kimi K2.6 web 37 vs agent 46.

Ranked leaderboard — best score per model

Same anonymized case, graded against the same guideline set. "Wrapper" = how the model was used.

#ModelWrapperBest score /100Key note
1DeepSeek V4 ProWorkBuddy agent50Only model to catch the staging trap (pT1aN0M0 IA)
2Kimi K2.6WorkBuddy agent46Stability flag: refused on first attempt in 1/3 runs
3GLM 5.2 (Zhipu)WorkBuddy agent44Largest floor lift (+39) — see wrapper effect
4Doubaoweb41Strong web baseline without an agent
5Kimi K3WorkBuddy agent40Re-scored 27 → 40 after a published correction
6DeepSeek V4 Flashagent / web34Web 28 → agent 34 (+6)
7DeepSeek (web, fast tier)web28
8ChatGLM / GLM (web)web5Baseline floor — wrap it in an agent before trusting

Wrapper effect — same model, web vs agent

ModelWeb /100Agent /100Wrapper ΔRead
GLM 5.2544+39Scaffolding turns a non-answer into a usable one
Kimi K2.63746+9Already coherent; agent adds polish + a cited guideline
DeepSeek V4 Flash2834+6Small lift — web baseline already moderate

Only models tested in both wrappers are listed (series rule). Full web-only runs for Kimi K3 and DeepSeek V4 Flash are pending and will extend this grid.

Free-model cheat card — which for which job

A clinician's quick routing. All scores are on the same case; your mileage varies with the prompt and the guideline version.

50
DeepSeek V4 Pro · agent
Safest single answer on a tricky staging case — the only model that caught the trap.
+39
GLM 5.2 · wrap in agent
If you only have a plain web chat with a weak model, scaffold it — the floor lift is largest.
46
Kimi K2.6 · agent
Consistent + decent, but watch the refusal flag — keep a human in the loop.
14 / 5
Grok 4.5 fast / ChatGLM web
Avoid for unassisted clinical decisions — copied the erroneous "Stage 0" line.

📬 Get the printable free-model cheat card

Subscribe and I'll send the one-page PDF cheat card (which free model for which clinical task, updated as the leaderboard grows). Free, roughly monthly, unsubscribe anytime.

Powered by Buttondown. Your email is never sold and never changes my independent conclusions.

Methodology in one line

One anonymized post-op breast-cancer case; identical prompt; a 16-item / 100-point rubric; reference standard = 2026 CBCS / CSCO 2024 / NCCN 2025. Every run is signed and reproducible. The case image, prompt, and score sheet are published in the Benchmark Data appendix so anyone can re-score any model. See the full Model Reviews series (V1 → V4) for transcripts.

Submit your own run (community)

This leaderboard is meant to grow beyond my own runs. You can add any free model you've tested — no code, no GitHub account needed.

① Get the public test kit (free, reusable)

Everything you need is already published:

② Fill in your result

Clicking opens your email app pre-filled to jstz1983@163.com with your answers. Please attach 1–2 screenshots of the model's answer so I can re-score and verify before publishing. Nothing is posted automatically — I review every submission.

Confidentiality: submit only results produced from the public framework kit. Do not include any patient data, unpublished manuscripts, or private prompts. By submitting you agree the result may be published on this leaderboard with credit (or anonymously, on request).
— Tan Haosheng, MD/PhD · 副主任医师 · 甲状腺乳腺外科 (Thyroid & Breast Surgery), 泰州市人民医院 (Taizhou People's Hospital). This is research-methodology documentation, not medical advice.