🇨🇳 阅读中文版 · This article is also available in Chinese.
🔁 Benchmark Repeats · Repeat Test #1 / 2026-08-01
DeepSeek Web 20-Run Complete Repeat Test: Same Case, Same Rubric — Is the Answer Stable?
By Tan Haosheng · Associate Chief Physician / MD, PhD · Thyroid & Breast Surgery, Taizhou People's Hospital · Published 2026-08-01
⚠️ Medical disclaimer: All clinical case analyses, model evaluations, and treatment discussions on this site are for educational and research purposes only. They are based on a single anonymized case and do not constitute individual medical advice, diagnosis, or treatment. For any real patient, always consult your own qualified physician and follow current guidelines (NCCN / CSCO / CBCS). AI outputs are not a substitute for professional clinical judgment. Last reviewed: 2026-08-01.
TL;DR (30-second takeaway)
DeepSeek web (fast mode) was fed the same standardized text prompt 20 times, scored item-by-item with the 16-item/100 rubric
Result: mean 69.8 (median 71.5), SD 3.54, range 12 points (63–75)
Per the framework stability rule (range ≥ 8 = unstable): judged UNSTABLE — about 1 in 4 runs fell into the low tier
Stable blind spot: all 4 endocrine-detail items (3.2a/b/c/d) scored 0/20 full marks, losing ~18 points combined — a reproducible model blind spot, not偶然
1. Why 20 runs: from "which model is best" to "is the answer stable?"
The Model Reviews series (V1–V5) answers the selection question: which free model performs better clinically. It runs each model only 1–3 times — a small-sample test group — which tells us the "average level" but not another, more critical question:
Same model, same standardized text prompt, run 20 times — will it crash?
For clinical use, one good run does not prove reliability. If a model tells you "no radiotherapy needed" in 15 runs but "strongly recommend radiotherapy" in 5, its high mean is useless for independent decisions. That is the purpose of this column (Benchmark Repeats): use repeated runs to expose occasional and systematic problems.
2. Method (same case, same rubric as Model Reviews)
Input: standardized text prompt (sent verbatim, not a screenshot) — the whole column uses the standard prompt, pure-text reasoning comparison, excluding visual/OCR error
Case: 60-year-old woman, right mastectomy + SLNB; invasive carcinoma with low-grade DCIS and ductal papillary carcinoma components, max diameter 0.6cm, with 3 clusters of stromal invasion (approximately 0.05cm, 0.06cm, 0.2cm); ER 3+(90%), PR 3+(90%), HER2 0, Ki-67 5%; sentinel 0/2 + 1 negative; discharge mislabeled TisN0M0
Model: DeepSeek web fast mode; auto web search every run ("read 7–15 web pages"), no API, no expert mode
Run: 20 runs each with a new browser session (each time manually verifying login then pasting the prompt), no context carryover; the framework auto-sends instructions and extracts the full response, with raw data archived for traceability
Scoring: the only Rubric = clinical-benchmark/framework 16-item/100; each response scored item-by-item by a clinician
Statistics: mean / median / SD / variance / range / coefficient of variation; stability judged by the framework rule (range ≥ 8 = unstable), and also by the academic criterion (range > 1.96×SD also = unstable)
3. 20 total scores and distribution
Run
Score distribution
Total
Tier
Run 1
█████████████░░░░░░░
65
RED
Run 2
█████████████░░░░░░░
69
AMBER
Run 3
█████████████░░░░░░░
68
AMBER
Run 4
██████████████░░░░░░
72
GREEN
Run 5
███████████████░░░░░
75
GREEN
Run 6
████████████░░░░░░░░
63
RED
Run 7
█████████████░░░░░░░
66
RED
Run 8
█████████████░░░░░░░
66
RED
Run 9
██████████████░░░░░░
70
AMBER
Run 10
███████████████░░░░░
75
GREEN
Run 11
██████████████░░░░░░
72
GREEN
Run 12
█████████████░░░░░░░
68
AMBER
Run 13
████████████░░░░░░░░
64
RED
Run 14
██████████████░░░░░░
71
AMBER
Run 15
██████████████░░░░░░
72
GREEN
Run 16
██████████████░░░░░░
73
GREEN
Run 17
██████████████░░░░░░
72
GREEN
Run 18
██████████████░░░░░░
72
GREEN
Run 19
██████████████░░░░░░
72
GREEN
Run 20
██████████████░░░░░░
72
GREEN
Statistic
Value
Mean
69.8 / 100
Median
71.5
Std Dev (SD)
3.54
Variance
12.56
Range (max−min)
12 (max 75 / min 63)
Coeff. of Variation (CV)
5.1%
Fig: DeepSeek web run 15, score 72 (top tier), showing "read 8 web pages". That run misread the invasive focus as 0.6cm → misjudged pT1b, an instance of text-value parsing fluctuation.
⚠️ Stability verdict: UNSTABLE Range 12 ≥ framework threshold 8. Also unstable by the academic criterion (range > 1.96×SD = 6.94). Across 20 runs: high tier (72–75) 10 times, mid tier (67–71) 5 times, low tier (63–66) 5 times — 1 in 4 runs fell into the low tier.
4. Item-by-item scores: what's stable, what's a blind spot
Capability radar (7 dimensions aggregated from the 16-item Rubric, 20-run mean)
The 16 scoring items are collapsed into 7 capability dimensions; each dimension's percentage = sum of that group's 20-run item means ÷ max. Single-model profile — see strengths and blind spots at a glance.
DeepSeek web (fast)pct = 20-run item-mean sum ÷ max
⭐ Overall rating (16-item mean sum ÷ 100)
DeepSeek web (fast)69.8/100★★★½☆3.5/5
Stability verdict across 20 runs: UNSTABLE (range 12 ≥ 8). Endocrine details (3.2a–d) scored 0/20 full marks, mean only 47% — not safe for independent complex endocrine decisions without clinician review; standard decisions (subtype, anti-HER2, self-assessment) were stable full marks all 20 runs.
4.1 Strengths (highly consistent full marks across 20 runs ✅)
All 20 runs failed to distinguish "strong recommendation for invasive cancer vs. chemoprevention after DCIS mastectomy (consider/optional)"
3.2b AI vs TAM + age cutoff ★
6.0/10
−4
All recommended AI, but none mentioned the 60-year age cutoff (≥60 AI≈TAM)
3.2c Drug dose & duration
2.5/6
−3.5
Only wrote "5 years" without specific drug + dose (letrozole 2.5mg/d, etc.)
3.2d OFS/extension/CDK4/6
1.5/6
−4.5
Most omitted that OFS is not applicable (postmenopausal), CDK4/6 not indicated
4.1 Adverse monitoring
2.5/6
−3.5
Only vaguely "monitor bone density/lipids", lacking specific adverse events and monitoring frequency
4.2 Bone density & lifestyle
1.9/4
−2.1
Lacked calcium + Vit D + weight-bearing exercise lifestyle advice
1.2 Staging correction
7.0/10
−3.0
0.2cm already exceeds microinvasion definition yet most judged pT1mi (OCR reading issue); Run 15 misread 0.6cm → pT1b
4.3 Fluctuating items (occasional loss — source of instability)
2.2 Radiotherapy (3.8/4): Run 1 judged "needs radiotherapy (strongly recommend)" ❌, Run 6 "strongly consider" — opposite to the correct answer (no indication). This is the most dangerous occasional-error type
4.3 Follow-up (3.5/4): some runs recommended unnecessary routine imaging/tumor-marker screening
4.4 BRCA (3.4/4): 4 runs "recommend testing" too aggressively (standard: no family history → not routine)
5. Clinical consistency observations
Text input causes numeric-parsing fluctuation: Run 15 misread "invasive focus 0.2cm" as 0.6cm → misjudged pT1b (see screenshot in Section 3); most runs correctly judged pT1mi; Runs 5/10 correctly judged pT1a — this is the model's difference in parsing the numeric value in the text prompt, not a visual/OCR limitation
Web search is an extra variable: all 20 runs auto-searched the web (7–15 pages), with mixed citation numbers (-1-, -2-), making it impossible to confirm they cited the exact 2026 CBCS page/section — the framework's requirement to "cite guideline section/page" is hard to verify in web fast mode
Core items stably lose points (strong): 3.2a/3.2b are deliberate framework discriminators (post-mastectomy DCIS endocrine = chemoprevention; 60-year AI/TAM age cutoff); DeepSeek missed both in all 20 runs — a reproducible model blind spot, suggesting systematic insufficiency in "distinguishing recommendation strength across scenarios" and "catching age cutoffs"
6. Conclusion & usage recommendations
Overall: mean 69.8/100 (good, not excellent). Strengths concentrate on "guideline-clear, decision-tree-clear" items (staging recognition, surgery, chemo, targeted, self-assessment); the core weakness is endocrine-treatment details (3.2a/b/c/d combined lost ~18 points).
Stability: ❌ range 12, judged UNSTABLE by the framework rule — must be stated when reporting or publishing.
Usage recommendation: DeepSeek web can serve as a first-pass assistant for clinical decisions; but endocrine-detail and age-stratified decisions must be manually reviewed; high-risk decisions like radiotherapy have occasional crashes and must not be trusted from a single output.
7. Data & reproduction
All 20 raw responses: public page (each run independently accessible, full text copyable — not a local path)
Scoring method: clinician scored each response item-by-item; per-item scores traceable
Frequently asked questions
Why does the same model score 12 points apart across 20 runs?
LLM sampling is stochastic, and DeepSeek web auto-searches the web each run (reading 7–15 pages), and the different evidence retrieved can shift staging and radiotherapy decisions. Across 20 runs the range reached 12 points (63–75) across 3 tiers, showing a single run is not representative.
Can I directly compare this column with Model Reviews (V1–V5)?
No. V1–V5 are small-sample groups (1–3 runs/model, mixed prompt styles); this column is a 20-run repeated standardized test. Adding or subtracting scores across columns yields wrong conclusions; read them separately.
What stable blind spots does DeepSeek show across 20 runs?
All 4 endocrine-detail items scored 0/20 full marks: 3.2a never distinguished "strong recommendation for invasive cancer vs. chemoprevention after DCIS mastectomy"; 3.2b always recommended AI but never mentioned the 60-year age cutoff; 3.2c/3.2d lacked specific drug doses and OFS/CDK4/6 assessment. Reproducible blind spots.
Why does the staging conclusion differ each time?
The largest invasive focus is 0.2cm (2mm), just over the microinvasion definition (≤1mm). Across 20 runs most judged pT1mi, 2 correctly pT1a, 1 misread as pT1b — this is a difference in how the model parses the value "0.2cm" in the text prompt, not a visual/OCR limitation. Run 15's screenshot (Fig 3-1) is exactly this misread-0.6cm output, confirming text-parsing fluctuation.
Why is the input a text prompt rather than a screenshot?
The whole column uses a standardized text prompt (not an anonymized case screenshot); the model reads directly from text. This tests pure text reasoning and clinical decision-making, and makes "reading fluctuation" attribution clearer — fluctuations come from text-value parsing, not OCR error.