⚠️ Affiliate Disclosure / 利益披露: This page may contain affiliate or promotional links. If you sign up or subscribe through such a link, I may earn a commission — but it never increases your price and never changes my independent conclusions. See the Affiliate Disclosure page.

4 Free AI Models on One Breast Cancer Case — A Surgeon's Extended Test

By Tan Haosheng · 副主任医师 / 医学博士 (Associate Chief Physician, MD/PhD) · 甲状腺乳腺外科 (Thyroid & Breast Surgery), 泰州市人民医院 (Taizhou People's Hospital) · Published 2026-07-30
⚠️ Medical disclaimer / 医疗免责声明: All clinical case analyses, model evaluations, and treatment discussions on this site are for educational and research purposes only. They are based on a single anonymized case and do not constitute individual medical advice, diagnosis, or treatment. For any real patient, always consult your own qualified physician and follow current guidelines (NCCN / CSCO / CBCS). AI outputs are not a substitute for professional clinical judgment. Last reviewed / 最后审阅: 2026-07-31.
🧪 原创实测 #2 / 5  ·  Hands-on Review 2 of 5 — Extended 4-model showdown
⚠️ Correction note (2026-07-31). Two clarifications that affect how this article should be read, both surfaced while writing Part 3:

The 16-item scores in this article are unchanged: this rubric never graded the LVI-vs-interstitial distinction. The score that did change is Kimi K3's in Part 3, re-scored from 27 to 40. Full details in the Part 3 correction note.

Why this follow-up exists: My first test pitted DeepSeek web (fast mode) against ChatGLM on a real, anonymized post-op breast cancer case and scored them with a fixed 16-item / 100-point rubric against the 2026 CBCS guideline. Two models were not enough. Readers — and I — wanted to know where Doubao (豆包) and the new DeepSeek V4 Pro inside the WorkBuddy agent land. So I ran the identical case, identical prompt, identical rubric on all four and put them side by side. Spoiler: only one model caught the single most important error hidden in the record — and it wasn't the one you'd guess.

What changed since Part 1

Same case, same 16-part prompt, same scoring key (correct + cited guideline = full; correct but uncited = half; wrong = 0). What's new:

The case (the trap is the whole point)

A 60-year-old woman, post-op day 1 after right breast cancer surgery. Pathology: three foci of stromal invasion, largest focus 0.2 cm (2 mm), ER 90%+, PR 90%+, HER2 0, Ki-67 5%. Sentinel nodes negative (0/2 + 1/1). Discharge diagnosis: TisN0M0 (stage 0, DCIS).

The trap: "3 foci of invasion, largest 0.2 cm" should set off an alarm. Microinvasion is defined as a single focus ≤1 mm. A 2 mm focus already exceeds that → at minimum pT1a (Stage IA), not stage 0. Any model that rubber-stamps "TisN0M0" is missing the single most clinically important discrepancy in the record. The guideline reference is CBCS 2026, TNM staging pp. 33–34, DCIS definition p. 50.

Results at a glance

🟣 DeepSeek V4 Pro (WorkBuddy)
50/100
Runs: strong / strong / 1 flawed
Language: Chinese
Caught staging trap: YES
Guideline citations: 0 / 16
🟠 豆包 Doubao (mobile)
41/100
Runs: structured ×3
Language: English (!)
Caught staging trap: NO
Guideline citations: 0 / 16
🔵 DeepSeek (web fast)
28/100
Runs: 28 · 28 · 28 (median 28)
Language: Chinese
Caught staging trap: NO
Guideline citations: 0 / 16
🟢 ChatGLM 5.2 (web)
5/100
Runs: 10 · 2 · 5 (median 5)
Language: English (!)
Caught staging trap: NO
Guideline citations: 0 / 16

The 16-item score sheet (all four, same rubric)

Item (max)DS V4 ProDoubaoDS webChatGLM
1.1 Challenge TisN0M0 (6)3000
1.2 Staging correction (10)5000
1.3 Molecular subtype (4)2221
2.1 Surgery + ALND (8)4442
2.2 Adjuvant radiation (4)2222
3.1 Chemo + genomic (10)5552
3.2a Endocrine necessity (12)6600
3.2b AI vs TAM + age (10)5500
3.2c Dose & duration (6)3331
3.2d OFS / CDK4-6 (6)3331
3.3 Anti-HER2 (4)2221
4.1 Adverse monitoring (6)3321
4.2 Bone density (4)2221
4.3 Follow-up plan (4)2221
4.4 Contralateral + BRCA (4)2221
5 Self-assessment (2)1010
TOTAL / 1005041285

Scoring rule applied uniformly: a correct answer that did not cite a specific CBCS chapter/page was capped at half credit (this is why even the winner stays at 50 — none of the four ever cited the guideline).

Overall ranking — all four, same case

All four models were scored on the same Part 2 case with the same 16-item / 100-point rubric, so this ranking is apples-to-apples (no cross-case fudge):

RankModel (wrapper)ScoreCaught the trap?One-line verdict
#1DeepSeek V4 Pro (WorkBuddy agent)50 / 100YESThe only model to flag the staging error. But 1 of 3 runs wrongly advised radiation — not safe to trust alone.
#2豆包 Doubao (mobile web)41 / 100NOBest structure among the non-trap-catchers; replied in English on a fully Chinese case.
#3DeepSeek web (fast)28 / 100NORock-stable (28·28·28) but missed the trap and the AI-vs-TAM age nuance.
#4ChatGLM 5.2 (web)5 / 100NOCollapsed on 2 of 3 runs, English-only, missed nearly everything.
Bottom line for V2: on this case, DeepSeek V4 Pro (via the WorkBuddy agent) is clearly the strongest, but its 1/3-run radiation error means it is not certified safe. Doubao is the surprise runner-up. ChatGLM 5.2 is not usable for this work.

The headline: only DeepSeek V4 Pro caught the trap

DeepSeek V4 Pro (WorkBuddy) was the only model to flag the staging error. Its answer opened by stating the discharge diagnosis "TisN0M0 (stage 0)" conflicts with the pathology (invasive carcinoma with a 2 mm focus) and corrected it to pT1N0(sn)M0 / IA stage — then built the whole plan on the corrected stage. That single move is worth ~16 rubric points and is the difference between "useful second opinion" and "confidently wrong."

By contrast, Doubao, DeepSeek web, and ChatGLM all accepted "TisN0M0" verbatim — including Doubao, which is otherwise the surprise of this test (see below). For a clinician, a model that silently endorses a wrong stage 0 label is the most dangerous failure mode: it looks authoritative and you may not re-check it.

DeepSeek V4 Pro — the winner, with one crack

Strengths: caught staging; correctly waived chemo (even mentioned Oncotype DX RS <20/25); correctly waived radiation for a mastectomy + N0; gave AI 5–10 years with drug names and doses; flagged bone-density/Vit-D monitoring; explicit self-assessment advising the clinician to correct the chart staging.
The crack — reproducibility: across three runs, one run wrongly recommended postoperative radiotherapy, erroneously describing the surgery as "breast-conserving" (保乳) when the record clearly states a simple mastectomy (右乳单纯切除). Post-mastectomy with negative margins and N0 has no radiation indication. So V4 Pro is not perfectly stable — 2 of 3 runs were excellent, 1 made a serious error. I scored the model on its best substantive answer but flag this as a real reliability concern: the same prompt can return a dangerous contradiction.

Doubao — the pleasant surprise, with two flaws

Stronger than expected: Doubao's three answers were well-structured, correctly handled the mastectomy → no-radiation logic, correctly waived chemo, and gave a reasonable AI-5-year plan with bone monitoring. On pure clinical structure it beat DeepSeek web and ChatGLM comfortably.
Flaw 1 — missed the trap: all three runs accepted "pTisN0M0 (stage 0)" without question, so it fell into the same staging error as the weaker models.
Flaw 2 — English on a Chinese prompt: exactly like ChatGLM in Part 1, Doubao answered a fully Chinese case in English. For a Chinese clinician that's a usability penalty: lost nuance in dosing/terminology and higher re-reading cost. (Not scored in the rubric, but noted as a dealbreaker for Chinese clinical use.)

Where all four still fail together

Head-to-head: which for which task

TaskPickWhy
Chinese clinical reasoning on this caseDeepSeek V4 Pro (WorkBuddy)Only one that caught the staging trap; Chinese
Structured draft, English-friendlyDoubaoClean structure, but replies in English
Quick Q&A / drug factsDeepSeek web fastInstant, Chinese, reproducible
Reading English literatureChatGLM / DoubaoEnglish output is their asset here
Staging re-interpretationNone, aloneOnly V4 Pro passed; even it was unstable
Citing a specific guideline chapterNoneAll scored 0/16

What I actually do with these tools

First person, no hedging: I use free LLMs as drafting and second-opinion aids, never as decision authority. After this extended test, the practical rule is sharper: when the case hides a staging or grading trap, only the V4-Pro-class model earns a glance — and even then I re-derive the stage myself. The 0/16 citation rate means I never accept a "guideline says so" claim without opening the PDF. Free models are genuinely good at the boring 80% (discharge drafts, patient-explanation scripts, English polish) and dangerously confident at the 20% that matters most.

Limitations

📎 Full data: the open testing framework (prompt + 16-item rubric), all four models' complete answers, and per-item score sheets live in the Clinical LLM Benchmark data appendix (bilingual). Read the long-form, then verify it yourself.
一句话结论(中文): 同一个病例、同一套 16 项评分、四个免费模型——DeepSeek V4 Pro(WorkBuddy 智能体)50 分排第一,是唯一识破"把浸润癌误写成 Tis 0 期"这一陷阱的模型;豆包 41 分是惊喜(结构清晰、正确判断全切后不放疗,但照单全收错误分期 + 用英文回答);DeepSeek 网页快速模式 28 分ChatGLM 5 分延续首测。四者 引用指南章节均为 0/16,且 V4 Pro 三次里有一次误把全切当保乳推了放疗——没有哪个能单独采信。免费模型适合写、适合想,不适合替你做分期和决策。
⚠️ Medical disclaimer. This article is an LLM tool-comparison evaluation. It is not medical advice, and model outputs must never be used for real patient decisions. The case is anonymized and used only to demonstrate a testing method.

Frequently asked questions

Is DeepSeek V4 Pro definitely better than DeepSeek web, or is it the WorkBuddy agent?

Good question and I can't fully separate them yet. The V4 Pro model was run inside the WorkBuddy agent, so the agent's system prompt / context may contribute to the better staging catch. I'll test V4 Pro raw (no agent) in a later part of this series to isolate the effect.

Why did Doubao answer in English?

Same behavior ChatGLM showed in Part 1 — the model defaulted to English despite a fully Chinese prompt and medical record. For Chinese clinical use that's a real usability penalty, though it didn't tank Doubao's clinical score (the rubric scores accuracy, not language).

One V4 Pro run recommended radiation — isn't that a big error?

Yes, it's serious: post-mastectomy with negative margins and N0 has no radiation indication, and the run wrongly called the surgery "breast-conserving." I scored V4 Pro on its best substantive answer but flag the inconsistency prominently — the same prompt returned both an excellent plan and a dangerous one. That instability is exactly why no model here is certified for solo use.

Can I reproduce this myself?

Yes. The full prompt and 16-item rubric are published at /clinical-benchmark/framework.html. Paste them into any model, score against the 2026 CBCS guideline, and you'll get your own numbers.

This is the second of a planned 5-part 原创实测 / Hands-on Review series — real, signed, reproducible tests of AI tools I actually use in clinical and research work. Next up: isolate whether the V4 Pro gain comes from the model or the WorkBuddy agent.