🇨🇳 阅读中文版 · This article is also available in Chinese.

🔁 Benchmark Repeats · Repeat Test #1 / 2026-08-01

DeepSeek Web 20-Run Complete Repeat Test: Same Case, Same Rubric — Is the Answer Stable?

By Tan Haosheng · Associate Chief Physician / MD, PhD · Thyroid & Breast Surgery, Taizhou People's Hospital · Published 2026-08-01
⚠️ Medical disclaimer: All clinical case analyses, model evaluations, and treatment discussions on this site are for educational and research purposes only. They are based on a single anonymized case and do not constitute individual medical advice, diagnosis, or treatment. For any real patient, always consult your own qualified physician and follow current guidelines (NCCN / CSCO / CBCS). AI outputs are not a substitute for professional clinical judgment. Last reviewed: 2026-08-01.
TL;DR (30-second takeaway)

1. Why 20 runs: from "which model is best" to "is the answer stable?"

The Model Reviews series (V1–V5) answers the selection question: which free model performs better clinically. It runs each model only 1–3 times — a small-sample test group — which tells us the "average level" but not another, more critical question:

Same model, same standardized text prompt, run 20 times — will it crash?

For clinical use, one good run does not prove reliability. If a model tells you "no radiotherapy needed" in 15 runs but "strongly recommend radiotherapy" in 5, its high mean is useless for independent decisions. That is the purpose of this column (Benchmark Repeats): use repeated runs to expose occasional and systematic problems.

2. Method (same case, same rubric as Model Reviews)

3. 20 total scores and distribution

RunScore distributionTotalTier
Run 1█████████████░░░░░░░65RED
Run 2█████████████░░░░░░░69AMBER
Run 3█████████████░░░░░░░68AMBER
Run 4██████████████░░░░░░72GREEN
Run 5███████████████░░░░░75GREEN
Run 6████████████░░░░░░░░63RED
Run 7█████████████░░░░░░░66RED
Run 8█████████████░░░░░░░66RED
Run 9██████████████░░░░░░70AMBER
Run 10███████████████░░░░░75GREEN
Run 11██████████████░░░░░░72GREEN
Run 12█████████████░░░░░░░68AMBER
Run 13████████████░░░░░░░░64RED
Run 14██████████████░░░░░░71AMBER
Run 15██████████████░░░░░░72GREEN
Run 16██████████████░░░░░░73GREEN
Run 17██████████████░░░░░░72GREEN
Run 18██████████████░░░░░░72GREEN
Run 19██████████████░░░░░░72GREEN
Run 20██████████████░░░░░░72GREEN
StatisticValue
Mean69.8 / 100
Median71.5
Std Dev (SD)3.54
Variance12.56
Range (max−min)12 (max 75 / min 63)
Coeff. of Variation (CV)5.1%
DeepSeek web run 15 result (score 72, top tier)
Fig: DeepSeek web run 15, score 72 (top tier), showing "read 8 web pages". That run misread the invasive focus as 0.6cm → misjudged pT1b, an instance of text-value parsing fluctuation.
⚠️ Stability verdict: UNSTABLE Range 12 ≥ framework threshold 8. Also unstable by the academic criterion (range > 1.96×SD = 6.94). Across 20 runs: high tier (72–75) 10 times, mid tier (67–71) 5 times, low tier (63–66) 5 times — 1 in 4 runs fell into the low tier.

4. Item-by-item scores: what's stable, what's a blind spot

ItemMaxMeanFull-mark runsZero runs
1.1 Agree with TisN0M065.919/200/20
1.2 Staging correction107.02/200/20
1.3 Molecular subtype44.020/200/20
2.1 Surgery rationale + ALND87.919/200/20
2.2 Adjuvant radiotherapy43.818/200/20
3.1 Chemotherapy + multigene test107.90/200/20
3.2a Endocrine necessity + mastectomy/BCS difference ★126.00/200/20
3.2b AI vs TAM + age cutoff ★106.00/200/20
3.2c Drug dose & duration62.50/200/20
3.2d OFS / extension / CDK4/661.50/200/20
3.3 Anti-HER244.020/200/20
4.1 Endocrine adverse monitoring62.50/200/20
4.2 Bone density & lifestyle41.91/200/20
4.3 Follow-up plan43.511/200/20
4.4 Contralateral breast + BRCA43.49/200/20
5 Self-assessment + uncertainty22.020/200/20

Capability radar (7 dimensions aggregated from the 16-item Rubric, 20-run mean)

The 16 scoring items are collapsed into 7 capability dimensions; each dimension's percentage = sum of that group's 20-run item means ÷ max. Single-model profile — see strengths and blind spots at a glance.

诊断分期DS 85局部治疗DS 98化疗决策DS 79内分泌治疗DS 47靶向治疗DS 100长期管理DS 63把握度DS 100
DeepSeek web (fast)pct = 20-run item-mean sum ÷ max
⭐ Overall rating (16-item mean sum ÷ 100)
DeepSeek web (fast)69.8/100★★★½☆3.5/5

Stability verdict across 20 runs: UNSTABLE (range 12 ≥ 8). Endocrine details (3.2a–d) scored 0/20 full marks, mean only 47% — not safe for independent complex endocrine decisions without clinician review; standard decisions (subtype, anti-HER2, self-assessment) were stable full marks all 20 runs.

4.1 Strengths (highly consistent full marks across 20 runs ✅)

4.2 Systematic blind spots (0/20 full marks — reproducible, not occasional)

ItemMean/MaxLostProblem
3.2a Endocrine necessity + mastectomy/BCS difference ★6.0/12−6All 20 runs failed to distinguish "strong recommendation for invasive cancer vs. chemoprevention after DCIS mastectomy (consider/optional)"
3.2b AI vs TAM + age cutoff ★6.0/10−4All recommended AI, but none mentioned the 60-year age cutoff (≥60 AI≈TAM)
3.2c Drug dose & duration2.5/6−3.5Only wrote "5 years" without specific drug + dose (letrozole 2.5mg/d, etc.)
3.2d OFS/extension/CDK4/61.5/6−4.5Most omitted that OFS is not applicable (postmenopausal), CDK4/6 not indicated
4.1 Adverse monitoring2.5/6−3.5Only vaguely "monitor bone density/lipids", lacking specific adverse events and monitoring frequency
4.2 Bone density & lifestyle1.9/4−2.1Lacked calcium + Vit D + weight-bearing exercise lifestyle advice
1.2 Staging correction7.0/10−3.00.2cm already exceeds microinvasion definition yet most judged pT1mi (OCR reading issue); Run 15 misread 0.6cm → pT1b

4.3 Fluctuating items (occasional loss — source of instability)

5. Clinical consistency observations

6. Conclusion & usage recommendations

Overall: mean 69.8/100 (good, not excellent). Strengths concentrate on "guideline-clear, decision-tree-clear" items (staging recognition, surgery, chemo, targeted, self-assessment); the core weakness is endocrine-treatment details (3.2a/b/c/d combined lost ~18 points).

Stability: ❌ range 12, judged UNSTABLE by the framework rule — must be stated when reporting or publishing.

Usage recommendation: DeepSeek web can serve as a first-pass assistant for clinical decisions; but endocrine-detail and age-stratified decisions must be manually reviewed; high-risk decisions like radiotherapy have occasional crashes and must not be trusted from a single output.

7. Data & reproduction

Frequently asked questions

Why does the same model score 12 points apart across 20 runs?

LLM sampling is stochastic, and DeepSeek web auto-searches the web each run (reading 7–15 pages), and the different evidence retrieved can shift staging and radiotherapy decisions. Across 20 runs the range reached 12 points (63–75) across 3 tiers, showing a single run is not representative.

Can I directly compare this column with Model Reviews (V1–V5)?

No. V1–V5 are small-sample groups (1–3 runs/model, mixed prompt styles); this column is a 20-run repeated standardized test. Adding or subtracting scores across columns yields wrong conclusions; read them separately.

What stable blind spots does DeepSeek show across 20 runs?

All 4 endocrine-detail items scored 0/20 full marks: 3.2a never distinguished "strong recommendation for invasive cancer vs. chemoprevention after DCIS mastectomy"; 3.2b always recommended AI but never mentioned the 60-year age cutoff; 3.2c/3.2d lacked specific drug doses and OFS/CDK4/6 assessment. Reproducible blind spots.

Why does the staging conclusion differ each time?

The largest invasive focus is 0.2cm (2mm), just over the microinvasion definition (≤1mm). Across 20 runs most judged pT1mi, 2 correctly pT1a, 1 misread as pT1b — this is a difference in how the model parses the value "0.2cm" in the text prompt, not a visual/OCR limitation. Run 15's screenshot (Fig 3-1) is exactly this misread-0.6cm output, confirming text-parsing fluctuation.

Why is the input a text prompt rather than a screenshot?

The whole column uses a standardized text prompt (not an anonymized case screenshot); the model reads directly from text. This tests pure text reasoning and clinical decision-making, and makes "reading fluctuation" attribution clearer — fluctuations come from text-value parsing, not OCR error.

🔁 Back to Benchmark Repeats · 🔬 Model Reviews (V1–V5) · 📊 Scoring framework