⚠️ Affiliate Disclosure / 利益披露: This page may contain affiliate or promotional links. If you sign up or subscribe through such a link, I may earn a commission — but it never increases your price and never changes my independent conclusions. See the Affiliate Disclosure page.
4 Free AI Models on One Breast Cancer Case — A Surgeon's Extended Test
By Tan Haosheng · 副主任医师 / 医学博士 (Associate Chief Physician, MD/PhD) · 甲状腺乳腺外科 (Thyroid & Breast Surgery), 泰州市人民医院 (Taizhou People's Hospital) · Published 2026-07-30
⚠️ Medical disclaimer / 医疗免责声明: All clinical case analyses, model evaluations, and treatment discussions on this site are for educational and research purposes only. They are based on a single anonymized case and do not constitute individual medical advice, diagnosis, or treatment. For any real patient, always consult your own qualified physician and follow current guidelines (NCCN / CSCO / CBCS). AI outputs are not a substitute for professional clinical judgment. Last reviewed / 最后审阅: 2026-07-31.
⚠️ Correction note (2026-07-31). Two clarifications that affect how this article should be read, both surfaced while writing Part 3:
Part 2 and Part 3 use the same case, not different cases. Earlier versions of both articles said otherwise. Every model across the series was given the identical anonymized pathology report. What differs between the two articles is the scoring sheet — 16 items here, a stricter 10 items in Part 3 — so scores are comparable within an article but only directional across them.
The three small foci (0.05 / 0.06 / 0.2 cm) are interstitial (stromal) invasion, not lymphovascular invasion. The report explicitly states "未见明确脉管瘤栓及神经侵犯" — no LVI, no perineural invasion. The largest invasive focus is therefore 2 mm and the correct stage is pT1aN0M0, Stage ⅠA (Luminal A-like), not pT1b. The report's own "TisN0M0 / Stage 0" summary line is a genuine error in the source document — catching it is exactly what the staging item scores.
The 16-item scores in this article are unchanged: this rubric never graded the LVI-vs-interstitial distinction. The score that did change is Kimi K3's in Part 3, re-scored from 27 to 40. Full details in the Part 3 correction note.
Why this follow-up exists: My first test pitted DeepSeek web (fast mode) against ChatGLM on a real, anonymized post-op breast cancer case and scored them with a fixed 16-item / 100-point rubric against the 2026 CBCS guideline. Two models were not enough. Readers — and I — wanted to know where Doubao (豆包) and the new DeepSeek V4 Pro inside the WorkBuddy agent land. So I ran the identical case, identical prompt, identical rubric on all four and put them side by side. Spoiler: only one model caught the single most important error hidden in the record — and it wasn't the one you'd guess.
What changed since Part 1
Same case, same 16-part prompt, same scoring key (correct + cited guideline = full; correct but uncited = half; wrong = 0). What's new:
豆包 Doubao (mobile web) — doubao.com. Three runs, all in English.
DeepSeek V4 Pro (via WorkBuddy agent) — the DeepSeek V4 Pro model wrapped in the WorkBuddy free agent, asked the same Chinese prompt. Three runs (the first returned a refusal, then two substantive answers, then a third).
DeepSeek web (fast mode) and ChatGLM web scores are carried over from Part 1 and re-expressed on the same scale so the four-way table is apples-to-apples.
The case (the trap is the whole point)
A 60-year-old woman, post-op day 1 after right breast cancer surgery. Pathology: three foci of stromal invasion, largest focus 0.2 cm (2 mm), ER 90%+, PR 90%+, HER2 0, Ki-67 5%. Sentinel nodes negative (0/2 + 1/1). Discharge diagnosis: TisN0M0 (stage 0, DCIS).
The trap: "3 foci of invasion, largest 0.2 cm" should set off an alarm. Microinvasion is defined as a single focus ≤1 mm. A 2 mm focus already exceeds that → at minimum pT1a (Stage IA), not stage 0. Any model that rubber-stamps "TisN0M0" is missing the single most clinically important discrepancy in the record. The guideline reference is CBCS 2026, TNM staging pp. 33–34, DCIS definition p. 50.
Runs: structured ×3
Language: English (!)
Caught staging trap: NO
Guideline citations: 0 / 16
🔵 DeepSeek (web fast)
28/100
Runs: 28 · 28 · 28 (median 28)
Language: Chinese
Caught staging trap: NO
Guideline citations: 0 / 16
🟢 ChatGLM 5.2 (web)
5/100
Runs: 10 · 2 · 5 (median 5)
Language: English (!)
Caught staging trap: NO
Guideline citations: 0 / 16
The 16-item score sheet (all four, same rubric)
Item (max)
DS V4 Pro
Doubao
DS web
ChatGLM
1.1 Challenge TisN0M0 (6)
3
0
0
0
1.2 Staging correction (10)
5
0
0
0
1.3 Molecular subtype (4)
2
2
2
1
2.1 Surgery + ALND (8)
4
4
4
2
2.2 Adjuvant radiation (4)
2
2
2
2
3.1 Chemo + genomic (10)
5
5
5
2
3.2a Endocrine necessity (12)
6
6
0
0
3.2b AI vs TAM + age (10)
5
5
0
0
3.2c Dose & duration (6)
3
3
3
1
3.2d OFS / CDK4-6 (6)
3
3
3
1
3.3 Anti-HER2 (4)
2
2
2
1
4.1 Adverse monitoring (6)
3
3
2
1
4.2 Bone density (4)
2
2
2
1
4.3 Follow-up plan (4)
2
2
2
1
4.4 Contralateral + BRCA (4)
2
2
2
1
5 Self-assessment (2)
1
0
1
0
TOTAL / 100
50
41
28
5
Scoring rule applied uniformly: a correct answer that did not cite a specific CBCS chapter/page was capped at half credit (this is why even the winner stays at 50 — none of the four ever cited the guideline).
Overall ranking — all four, same case
All four models were scored on the same Part 2 case with the same 16-item / 100-point rubric, so this ranking is apples-to-apples (no cross-case fudge):
Rank
Model (wrapper)
Score
Caught the trap?
One-line verdict
#1
DeepSeek V4 Pro (WorkBuddy agent)
50 / 100
YES
The only model to flag the staging error. But 1 of 3 runs wrongly advised radiation — not safe to trust alone.
#2
豆包 Doubao (mobile web)
41 / 100
NO
Best structure among the non-trap-catchers; replied in English on a fully Chinese case.
#3
DeepSeek web (fast)
28 / 100
NO
Rock-stable (28·28·28) but missed the trap and the AI-vs-TAM age nuance.
#4
ChatGLM 5.2 (web)
5 / 100
NO
Collapsed on 2 of 3 runs, English-only, missed nearly everything.
Bottom line for V2: on this case, DeepSeek V4 Pro (via the WorkBuddy agent) is clearly the strongest, but its 1/3-run radiation error means it is not certified safe. Doubao is the surprise runner-up. ChatGLM 5.2 is not usable for this work.
The headline: only DeepSeek V4 Pro caught the trap
DeepSeek V4 Pro (WorkBuddy) was the only model to flag the staging error. Its answer opened by stating the discharge diagnosis "TisN0M0 (stage 0)" conflicts with the pathology (invasive carcinoma with a 2 mm focus) and corrected it to pT1N0(sn)M0 / IA stage — then built the whole plan on the corrected stage. That single move is worth ~16 rubric points and is the difference between "useful second opinion" and "confidently wrong."
By contrast, Doubao, DeepSeek web, and ChatGLM all accepted "TisN0M0" verbatim — including Doubao, which is otherwise the surprise of this test (see below). For a clinician, a model that silently endorses a wrong stage 0 label is the most dangerous failure mode: it looks authoritative and you may not re-check it.
DeepSeek V4 Pro — the winner, with one crack
Strengths: caught staging; correctly waived chemo (even mentioned Oncotype DX RS <20/25); correctly waived radiation for a mastectomy + N0; gave AI 5–10 years with drug names and doses; flagged bone-density/Vit-D monitoring; explicit self-assessment advising the clinician to correct the chart staging.
The crack — reproducibility: across three runs, one run wrongly recommended postoperative radiotherapy, erroneously describing the surgery as "breast-conserving" (保乳) when the record clearly states a simple mastectomy (右乳单纯切除). Post-mastectomy with negative margins and N0 has no radiation indication. So V4 Pro is not perfectly stable — 2 of 3 runs were excellent, 1 made a serious error. I scored the model on its best substantive answer but flag this as a real reliability concern: the same prompt can return a dangerous contradiction.
Doubao — the pleasant surprise, with two flaws
Stronger than expected: Doubao's three answers were well-structured, correctly handled the mastectomy → no-radiation logic, correctly waived chemo, and gave a reasonable AI-5-year plan with bone monitoring. On pure clinical structure it beat DeepSeek web and ChatGLM comfortably.
Flaw 1 — missed the trap: all three runs accepted "pTisN0M0 (stage 0)" without question, so it fell into the same staging error as the weaker models.
Flaw 2 — English on a Chinese prompt: exactly like ChatGLM in Part 1, Doubao answered a fully Chinese case in English. For a Chinese clinician that's a usability penalty: lost nuance in dosing/terminology and higher re-reading cost. (Not scored in the rubric, but noted as a dealbreaker for Chinese clinical use.)
Where all four still fail together
0 / 16 guideline citations — every model. All four paraphrase "mainstream / authoritative / latest guideline" without a single chapter or page. You cannot verify they read CBCS 2026. This is the single most important limitation and it hasn't moved since Part 1.
The AI-vs-TAM age cutoff (3.2b) is fragile. Only V4 Pro and Doubao gave a usable AI-first answer; none explicitly cited the CBCS note that AI's advantage over tamoxifen is mainly <60 years. At exactly 60, the boundary matters.
None is safe to trust alone on staging or nuanced recommendation grading. The scores rank them; they do not certify any of them.
Head-to-head: which for which task
Task
Pick
Why
Chinese clinical reasoning on this case
DeepSeek V4 Pro (WorkBuddy)
Only one that caught the staging trap; Chinese
Structured draft, English-friendly
Doubao
Clean structure, but replies in English
Quick Q&A / drug facts
DeepSeek web fast
Instant, Chinese, reproducible
Reading English literature
ChatGLM / Doubao
English output is their asset here
Staging re-interpretation
None, alone
Only V4 Pro passed; even it was unstable
Citing a specific guideline chapter
None
All scored 0/16
What I actually do with these tools
First person, no hedging: I use free LLMs as drafting and second-opinion aids, never as decision authority. After this extended test, the practical rule is sharper: when the case hides a staging or grading trap, only the V4-Pro-class model earns a glance — and even then I re-derive the stage myself. The 0/16 citation rate means I never accept a "guideline says so" claim without opening the PDF. Free models are genuinely good at the boring 80% (discharge drafts, patient-explanation scripts, English polish) and dangerously confident at the 20% that matters most.
Limitations
One case, text-only input, free tiers only, three runs per model (V4 Pro: one refusal + two strong + one flawed).
Doubao was tested on mobile web; DeepSeek V4 Pro inside the WorkBuddy agent — the agent wrapper may contribute to the better staging catch, which I cannot yet separate from the base model.
Scores are one surgeon's interpretation of a fixed rubric; the framework is published so you can disagree and re-score.
📎 Full data: the open testing framework (prompt + 16-item rubric), all four models' complete answers, and per-item score sheets live in the Clinical LLM Benchmark data appendix (bilingual). Read the long-form, then verify it yourself.
⚠️ Medical disclaimer. This article is an LLM tool-comparison evaluation. It is not medical advice, and model outputs must never be used for real patient decisions. The case is anonymized and used only to demonstrate a testing method.
Frequently asked questions
Is DeepSeek V4 Pro definitely better than DeepSeek web, or is it the WorkBuddy agent?
Good question and I can't fully separate them yet. The V4 Pro model was run inside the WorkBuddy agent, so the agent's system prompt / context may contribute to the better staging catch. I'll test V4 Pro raw (no agent) in a later part of this series to isolate the effect.
Why did Doubao answer in English?
Same behavior ChatGLM showed in Part 1 — the model defaulted to English despite a fully Chinese prompt and medical record. For Chinese clinical use that's a real usability penalty, though it didn't tank Doubao's clinical score (the rubric scores accuracy, not language).
One V4 Pro run recommended radiation — isn't that a big error?
Yes, it's serious: post-mastectomy with negative margins and N0 has no radiation indication, and the run wrongly called the surgery "breast-conserving." I scored V4 Pro on its best substantive answer but flag the inconsistency prominently — the same prompt returned both an excellent plan and a dangerous one. That instability is exactly why no model here is certified for solo use.
Can I reproduce this myself?
Yes. The full prompt and 16-item rubric are published at /clinical-benchmark/framework.html. Paste them into any model, score against the 2026 CBCS guideline, and you'll get your own numbers.
This is the second of a planned 5-part 原创实测 / Hands-on Review series — real, signed, reproducible tests of AI tools I actually use in clinical and research work. Next up: isolate whether the V4 Pro gain comes from the model or the WorkBuddy agent.