🔁 Benchmark Repeats · #3 / 2026-08-02

GLM vs DeepSeek R1: Same Criteria, Same Depth — Who Wins on Stability?

⚠️ Medical disclaimer: All clinical case analyses are for educational and research purposes only. Based on a single anonymized case — not individual medical advice.

Key Findings

With deep thinking + web search enabled, DeepSeek R1 outperforms GLM-5.2 in both accuracy and stability:
Mean 86.5 vs 80.8 (+5.7), range 14 vs 35 (−60%), SD 4.4 vs 9.9 (−55%).
GLM-5.2 Mean
80.8
Range: 35 | SD: 9.9
DeepSeek R1 Mean
86.5
Range: 14 | SD: 4.4
R1 Advantage
+5.7
Mean +5.7 · Range −21
💡 Key Insight
GLM's outputs fluctuate wildly (63→98, spanning 4 tiers). DeepSeek R1, under the same settings, shows much tighter clustering (79→92, only 2 tiers). For clinical decision support, consistency beats occasional brilliance.

Methodology

  • Case: 60F, right mastectomy + SLNB; invasive carcinoma invasive carcinoma with low-grade DCIS and ductal papillary carcinoma components, max diameter 0.6cm, with 3 clusters of stromal invasion (approximately 0.05cm, 0.06cm, 0.2cm); ER/PR 3+, HER2-, Ki-67 5%; discharge dx TisN0M0 (incorrect)
  • Input: Canonical text prompt (same 5-question format), identical for both models
  • Settings: Both models: deep thinking ON + web search ON (same as Part 2 R1 setting)
  • Scoring: 16-item/100 rubric (clinical-benchmark/framework)
  • Runs: GLM 16 valid (1 INVALID), DeepSeek R1 20 valid

Item-by-Item Comparison (16 Rubric Items)

Item Max GLM Avg R1 Avg Diff
1.1 TisN0M0 agreement65.85.5-0.3
1.2 Staging correction108.48.5
1.3 Molecular subtype44.04.0
2.1 Surgical adequacy + ALND86.96.5-0.4
2.2 Adjuvant radiotherapy43.43.2-0.2
3.1 Chemo indication + 21-gene108.19.5+1.4
3.2a Endocrine necessity129.610.8+1.2
3.2b AI vs TAM + age cut-off107.510.0+2.5
3.2c Drug dosage & duration64.85.5+0.7
3.2d OFS/CDK4/663.24.0+0.8
3.3 Anti-HER243.83.5-0.3
4.1 ADE monitoring63.43.0-0.4
4.2 Bone density & lifestyle42.31.7-0.6
4.3 Follow-up plan43.64.0+0.4
4.4 Contralateral breast + BRCA43.13.5+0.4
5. Self-assessment + uncertainty21.81.3-0.5
Total10080.886.5+5.7

Conclusion

Under identical deep-thinking + web-search settings, DeepSeek R1 delivers both higher accuracy and greater stability than GLM-5.2.
Mean +5.7, range reduced by 60% (35→14), SD reduced by 55%.

Implications for Clinical Use

  • Stability > occasional high scores: A model with range 35 is unsuitable for independent decision support
  • Deep thinking adds value: Both models improved from quick mode, but R1's consistency is superior
  • Model selection: DeepSeek R1 is the more reliable choice for clinical decision support under these settings

Limitations

  • Sample size limited (GLM 16 runs, R1 20 runs) — needs more repetition
  • Single case tested; conclusions may not generalize
  • Scoring based on keyword coverage; may differ from human expert scoring

FAQ

Why does GLM have such a wide range?

GLM's deep thinking mode can produce either high-quality answers (98) or answers missing key items (63), depending on reasoning path consistency. R1's reasoning is more stable.

Is this comparison fair?

Yes. Same case, same prompt, same settings (deep thinking + web search ON), same rubric. Only the model differs.

Is a range of 14 considered stable?

Per our framework (range ≥ 8 = unstable), R1's 14 is still "unstable" but much better than GLM's 35. Ideal target: range < 8.

← Part 2: GLM 20 Runs · Part 1: DeepSeek Quick vs R1