🔁 Benchmark Repeats · Repeated Test #6 / 2026-08-02
Agnes-2.0-flash vs DeepSeek V4 Flash: 20-Run Head-to-Head on One Case
🎯 Core Philosophy
AI is an assistant to the physician, not a replacement. We use AI to help clinicians cross-check against guidelines, surface points that may be missed, and point to specific guideline sections and page numbers. Final treatment decisions rest with the physician, based on evidence-based medicine.
By Tan Haosheng · Associate Chief Physician, MD/PhD · Thyroid & Breast Surgery, Taizhou People's Hospital · Published 2026-08-02 · Updated 2026-08-02
⚠️ Medical disclaimer: All clinical case analyses, model evaluations, and treatment discussions on this site are for educational and research purposes only. They are based on a single anonymized case and do not constitute individual medical advice, diagnosis, or treatment. For any real patient, always consult your own qualified physician and follow current guidelines (NCCN / CSCO / CBCS). AI outputs are not a substitute for professional clinical judgment. Last reviewed: 2026-08-02.
TL;DR (30-second verdict)
Same standard text prompt, same 16-item/100 rubric — Agnes-2.0-flash and DeepSeek V4 Flash each run 20 times (both via API, no web search)
DeepSeek mean 92.8 leads Agnes mean 85.5 by 7.3 (median 94 vs 86)
Stability: both ranges ≥ 8 → both unstable (Agnes range 22, DeepSeek range 17)
Gap is in endocrine-therapy details (3.2a–d) and long-term management (4.1/4.2); core staging, surgery, chemo, targeted and confidence are near-perfect for both
Speed: Agnes median ~6.5s vs DeepSeek ~150s — Agnes ~20× faster and more concise
1. Why put these two models head-to-head
One is our own Agnes-2.0-flash (free, clinical-decision API agent); the other is currently the cheapest frontier API, DeepSeek V4 Flash. They are close in positioning — both cheap, fast, directly usable for clinical Q&A — which makes them a fair same-tier match.
More importantly, this group deliberately isolates the web-search variable from earlier column pieces (e.g. the DeepSeek web 20-run test): both models run via API with web search off, so what is compared is each model's own parametric clinical knowledge, not retrieval strength. That is also why this group's scores cannot be subtracted from the earlier DeepSeek web test — input mode and retrieval conditions differ.
2. Method (same case, same rubric as the whole series)
Input: standard text prompt (sent verbatim, not a screenshot) — the series-wide standard prompt, pure-text reasoning, no visual/OCR noise
Models: Agnes-2.0-flash (apihub API) vs DeepSeek V4 Flash (api.deepseek.com); web search disabled for both, parametric knowledge only
Runs: 20 each, fresh session per run, no context carry-over; framework auto-sends prompt, extracts full response, raw data archived
Scoring: sole Rubric = clinical-benchmark/framework 16-item/100; each response scored item-by-item (0/1/2 scale, converted to percentage)
Stats: mean / median / SD / range; stability by framework rule (range ≥ 8 = unstable) + academic rule (range > 1.96×SD also unstable)
3. Total scores & distribution
🟣 Agnes-2.0-flash (API)
85.5
Median 86 · SD 5.35 · range 22 (76–98) 20 runs · no web search unstable
🟢 DeepSeek V4 Flash (API)
92.8
Median 94 · SD 5.83 · range 17 (83–100) 20 runs · no web search leads +7.3
Statistic
Agnes-2.0-flash
DeepSeek V4 Flash
Mean
85.5 / 100
92.8 / 100
Median
86
94
SD
5.35
5.83
Range (max−min)
22 (98 / 76)
17 (100 / 83)
CV
6.3%
6.3%
Response time (median / range)
~6.5s / 0.8–35s
~150s / 108–192s
⚠️ Stability: both unstable Framework rule (range ≥ 8): Agnes 22, DeepSeek 17, both over. Academic rule (1.96×SD): thresholds 10.5 / 11.4, both ranges still exceed. I.e. both models had clear crashes across 20 runs; a single output does not represent their level.
Each dimension's percentage = sum of 20-run item means in that group ÷ max. Both model shapes overlaid to show strengths and the source of the gap at a glance.
🟣 Agnes-2.0-flash🟢 DeepSeek V4 Flash% = sum of 20-run item means ÷ max
4.2 Strengths both share (20/20 near-perfect ✅)
1.1 / 1.3 staging & subtype: challenge TisN0M0, correctly call Luminal A — both 20/20
2.1 / 2.2 surgery & RT: no ALND, no RT — both 20/20
3.3 anti-HER2: HER2-0 → no targeted indication — both 20/20
4.3 / 4.4 follow-up & contralateral: both 20/20
5 confidence: both give uncertainty notes, 20/20
Takeaway: On "guideline-clear, decision-tree-clean" core clinical decisions, both models are near-perfect and highly stable — that is not the selection divide.
4.3 Source of the gap: endocrine details + long-term management
Dimension
Agnes
DeepSeek
Gap
Endocrine therapy (3.2a–d)
76.2%
85.7%
−9.5
Long-term management (4.1/4.2)
74.2%
86.7%
−12.5
Staging (1.2)
92.5%
100%
−7.5
DeepSeek's 7.3-point lead comes almost entirely from endocrine-therapy details and long-term management:
3.2a endocrine necessity + lumpectomy/mastectomy difference: Agnes only 2/20 full marks (6.6/12), DeepSeek 7/20 (8.1/12) — both often miss the "invasive cancer strong-recommend vs DCIS post-mastectomy = chemoprevention" distinction
4.2 bone density / lifestyle / BRCA: Agnes only 4/20 full marks (1.9/4), its weakest item; DeepSeek 11/20 (3.1/4) — Agnes far more often omits lifestyle and BRCA cues
4.1 AE monitoring: Agnes only 3/20 full marks (3.4/6), DeepSeek 10/20 (4.5/6)
This reflects: Agnes's parametric knowledge recalls guideline-edge details (dosing, monitoring frequency, lifestyle, genetic counseling) more weakly than DeepSeek, but is not behind on主干 decisions. Note neither model had web search here, so this is a difference in each model's own knowledge surface, not retrieval.
5. Judge fairness (expert ruling)
✅ Judge ruling: fair. This group was scored by an LLM-as-Judge based on Agnes; an early draft flagged "same-source as Agnes, possible bias toward Agnes." After the site author (associate chief physician, MD/PhD, 20+ yrs clinical) spot-checked run by run, the judge is certified fair:
The result itself rebuts bias: had the judge systematically favored the same-source model, Agnes would be inflated; but the judged model DeepSeek (92.8) actually scored above Agnes (85.5) — no evidence of pro-Agnes inflation.
Expert spot-check consistent: DeepSeek full-mark runs (8/12/16/19) and Agnes high run (run1=98) were verified item-by-item; scores match clinical judgment, no compression or inflation found.
Unified scoring protocol: both models share one 0/1/2 Rubric and one conversion formula, with anti-anchoring (each run scored independently, not referenced to the other model).
For transparency we still disclose: the judge is an Agnes-based LLM, certified fair by the author; readers wanting a fully third-party judge can review the full raw responses and judge themselves.
6. Verdict & usage advice
Overall: DeepSeek V4 Flash 92.8 leads Agnes-2.0-flash 85.5 (+7.3). The lead is almost all in endocrine details and long-term management; on core clinical decisions both are near-perfect.
Stability: ❌ both ranges ≥ 8, both unstable — must be stated on publish/quote; a single output is not a basis for结论.
Speed/UX: Agnes ~20× faster (~6.5s vs ~150s) and more concise; DeepSeek slower but more complete on guideline细节.
Usage advice: both fit as clinical screening / second-opinion assistants. For guideline-edge details — endocrine dosing (3.2c/d), follow-up refinement (4.1), BRCA/lifestyle (4.2) — a clinician must verify; never trust any single output alone. If workflow priority is "fast + concise", Agnes is smoother; if "complete guideline recall", DeepSeek is more thorough.
7. Data & reproducibility
40 raw responses (text + screenshots): public page (each run independently accessible, full text copyable, not a local path)
Run config: Agnes-2.0-flash (apihub API) and DeepSeek V4 Flash (api.deepseek.com), 20 runs each, web search off
Frequently asked questions
DeepSeek scores 7.3 higher — is it simply better overall?
No. The gap comes almost entirely from endocrine-therapy details (3.2a–d) and long-term management (4.1/4.2). DeepSeek wins on parametric guideline recall there, but both are near-perfect on core staging, surgery, chemo, targeted and confidence. Agnes also responds ~20× faster and more concisely.
Is the judge (scorer) fair?
Scoring used an LLM-as-Judge based on Agnes; an early draft flagged possible same-source bias toward Agnes. But the judged model DeepSeek scored higher than Agnes itself — had there been systematic favoritism, Agnes would be inflated. The site author (associate chief physician, MD/PhD) spot-checked full-score and borderline runs and confirmed scores match clinical judgment; the judge is certified fair.
Is this the same setup as the earlier DeepSeek 20-run test?
No. The earlier test used the web UI fast mode (auto web search each run). Here both models ran via API with web search off, isolating each model's own parametric knowledge. The two scores must not be subtracted.
Both unstable — still usable?
By framework (range ≥ 8) and academic (range > 1.96×SD) rules both are unstable. But point losses concentrate in guideline-detail items; core decision items are highly consistent. Use as screening/second-opinion assistants; have a clinician verify endocrine dosing, follow-up and BRCA details; never trust a single output alone.
Why is Agnes so fast yet lower-scored?
Speed comes from a more compact model size and inference path; the lower score is mainly weaker recall of guideline-edge details (dosing, monitoring frequency, lifestyle, genetic counseling) versus DeepSeek, while主干 decisions are not behind. Also Agnes had no web search here to補 recall of guideline text.