๐Ÿ‡จ๐Ÿ‡ณ ้˜…่ฏปไธญๆ–‡็‰ˆ ยท ไธญๆ–‡็‰ˆๅŒๆญฅไธŠ็บฟใ€‚

๐Ÿ” Benchmark Repeats ยท Repeat #2 / 2026-08-01

GLM Web 20-Run Repeat: One Case, One Rubric โ€” Is It Stable?

By Tan Haosheng ยท Associate Chief Physician, MD/PhD ยท Thyroid & Breast Surgery, Taizhou People's Hospital ยท Published 2026-08-01
โš ๏ธ Medical disclaimer: All clinical case analyses and model evaluations are for educational and research purposes only, based on a single anonymized case, and do not constitute individual medical advice. Follow current guidelines (NCCN / CSCO / CBCS). AI outputs are not a substitute for professional judgment. Last reviewed: 2026-08-01.
โš ๏ธ DRAFT framework: This is the structural scaffold for Benchmark Repeats article #2. GLM 20-run data will be filled in after the repeat test completes and is scored.
TL;DR

1. Why 20 runs: from "which model is best" to "is the answer stable"

The Model Reviews series (V1โ€“V5) answers a selection question. It runs each model only 1โ€“3 times โ€” a small-sample test group. That tells us the "average level" but not the more critical question:

Same model, same standard prompt, run 20 times โ€” does it ever break down?

For clinical use, one good run does not prove reliability. If a model wrongly recommends radiotherapy in several of 20 runs, its high average is irrelevant for independent decisions. That is why this column (Benchmark Repeats) exists: repetition exposes occasional and systematic failures. Article #1 (DeepSeek web) showed a range of 12 points โ†’ unstable. This article tests GLM with the same case, rubric, and input protocol.

2. Method (same case, same rubric as Model Reviews & Repeat #1)

3. 20-run scores & distribution

RunDistributionScoreTier
Run 1โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 2โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 3โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 4โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 5โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 6โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 7โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 8โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 9โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 10โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 11โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 12โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 13โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 14โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 15โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 16โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 17โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 18โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 19โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
Run 20โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ€”โ€”
StatisticValue
Meanโ€”
Meanโ€”
Meanโ€”
Meanโ€”
Meanโ€”
Meanโ€”

4. Item-by-item: what is stable, what is the blind spot

ItemMaxMeanFull marksZero
1.1 ๆ˜ฏๅฆ่ฎคๅŒ TisN0M06โ€”โ€”โ€”
1.2 ๅˆ†ๆœŸไฟฎๆญฃ10โ€”โ€”โ€”
1.3 ๅˆ†ๅญๅˆ†ๅž‹4โ€”โ€”โ€”
2.1 ๆ‰‹ๆœฏๅˆ็†ๆ€ง + ALND8โ€”โ€”โ€”
2.2 ่พ…ๅŠฉๆ”พ็–—4โ€”โ€”โ€”
3.1 ๅŒ–็–—ๆŒ‡ๅพ + ๅคšๅŸบๅ› ๆฃ€ๆต‹10โ€”โ€”โ€”
3.2a ๅ†…ๅˆ†ๆณŒๅฟ…่ฆๆ€ง + ๅ…จๅˆ‡/ไฟไนณๅทฎๅผ‚ โ˜…12โ€”โ€”โ€”
3.2b AI vs TAM + ๅนด้พ„ๅˆ‡็‚น โ˜…10โ€”โ€”โ€”
3.2c ่ฏ็‰ฉๅ‰‚้‡็–—็จ‹6โ€”โ€”โ€”
3.2d OFS / ๅปถ้•ฟ / CDK4/66โ€”โ€”โ€”
3.3 ๆŠ— HER24โ€”โ€”โ€”
4.1 ๅ†…ๅˆ†ๆณŒไธ่‰ฏๅๅบ”็›‘ๆต‹6โ€”โ€”โ€”
4.2 ้ชจๅฏ†ๅบฆไธŽ็”Ÿๆดปๆ–นๅผ4โ€”โ€”โ€”
4.3 ้š่ฎฟ่ฎกๅˆ’4โ€”โ€”โ€”
4.4 ๅฏนไพงไนณ่…บ + BRCA4โ€”โ€”โ€”
5 ่‡ชๆˆ‘่ฏ„ไผฐ + ไธ็กฎๅฎšๆ€ง4โ€”โ€”โ€”

5. Clinical consistency observations

6. Conclusion & usage advice

Filled after data is ready.

7. Data & reproduction

๐Ÿ” Back to Benchmark Repeats ยท ๐Ÿ”ฌ Model Reviews (V1โ€“V5) ยท ๐Ÿ“Š Scoring framework