Home › AI Agents for Medical Research
I've spent 20+ years in thyroid and breast surgery and 10+ years publishing in oncology. In July 2026 I started asking a sharp question: can a general-purpose AI agent, with one prompt, do work I'd actually trust to inform a manuscript I'm about to submit? These case studies are my honest answer — including the parts where the agent pushed back on my own draft.
Read this first — what is and isn't published here
Every case on this page comes from unpublished work in progress. So this section reports the agent's behaviour only: how it approached the task, whether it guessed or asked, whether it pushed back, and how long it took.
- Withheld entirely: the research question, the hypotheses, any conclusion, any figure, any number, any gene, any disease context, any file or sample identifier.
- Published: the protocol, the rubric, the wall-clock timings, the agent's behavioural pass/fail per dimension, and a small number of fully-redacted illustrative screenshots that show only that output was produced — not what it said.
- Why: peer review comes first. These pages exist to answer "can an agent do this kind of work", not to preprint anything.
Why this section exists
The "AI for X" discourse is mostly testimonials. I wanted measurements on my own work, with the prompts saved and the runs recorded. Two things forced this section into existence:
- I could not remember how one of my own samples had been prepared. A real analysis can't start until you know what's in front of you — and that's exactly where AI agents tend to silently guess. I tested whether an agent would ask, work it out, or hallucinate.
- I had a draft from a different line of work lying around. I deliberately attached it to a dataset it had nothing to do with, and watched whether the agent noticed the mismatch or generated a confident cross-validation anyway.
What's here (start anywhere)
Methodology — How I test AI agents on research tasks
The protocol I use for every case: prompt, data, evaluation criteria, what counts as a pass, what counts as a hallucination, and the four failure modes I watch for.
Case 1 · 18 min 05 s, one shotReading the data, checking a draft
Two unlabelled samples handed over, plus a draft of my own. Could the agent tell the samples apart from the data itself — and did it hold the draft to the data, or just agree with me?
Case 2 · trap test · 4 minThe mismatch trap
The "wrong paper" test. I attached a draft from a completely different line of work on top of the same dataset. Does the agent notice that the data cannot speak to that draft — or does it manufacture a validation?
Case 3 · forward look · one promptFrom raw data to a research plan
Same data, no draft to defend. I asked the agent to propose what could realistically be written and where it could go. A different capability from Cases 1–2: originating a plan rather than auditing one.
Headline findings (so you can skim)
Behavioural results only. What the agent concluded scientifically is deliberately not reported.
| Test point | What the agent did | Verdict |
|---|---|---|
| Work out how two unlabelled samples differed, using the data alone | Derived the distinction from the data itself and stated it with the supporting evidence, instead of asking me or picking one at random. I knew the ground truth; it matched. | Correct |
| Hold my own draft to the data rather than agreeing with me | Went through the draft claim by claim and separated the ones the data could carry from the ones it could not. It contradicted parts of my own text to my face rather than validating everything. | Pushed back — as it should |
| Handle a draft the data has no business validating | Refused to treat the exercise as a validation. Named the specific structural reasons the dataset could not answer that draft's question before producing anything else. | Correct |
| Avoid the classic statistical traps | Distinguished per-cell from aggregated analysis rather than reporting one as the other, and accounted for technical depth before calling anything enriched. | Passed |
| Wall-clock time, one shot each | Case 1: 18 min 05 s. Case 2 (follow-up prompt, same session): ~4 min. ~22 minutes total — no re-prompting, no retries, no hand-holding. | Fast |
| Originate a research plan, not just audit one | Case 3: produced several concrete, differentiated directions with realistic ambition levels — and led with the ceiling imposed by the dataset's size rather than burying it. | Yes, with honest scoping |
| Where it was weakest | It is only as good as the framing it is given: with a vague prompt it will happily produce a well-organised answer to the wrong question. It also cannot know what it was not shown, and does not always say so unprompted. | Needs a domain expert in the loop |
What this section is not
- Not a benchmark. I'm not ranking agents. WorkBuddy Hy3 is the only one I tested under this protocol so far. See /model-reviews for cross-model LLM benchmarks.
- Not a recommendation. The findings are mine and reproducible. Your data, your manuscript, your statistical power — your call.
- Not medical advice. The cases are about research methodology, not clinical decisions. Anything a surgeon would act on needs a human pathologist, a human radiologist, and a tumor board.
- Not a preprint. No hypothesis, result, figure or data point from any unpublished study appears anywhere in this section, in any form.
About the author (so you can weigh this)
I am Tan Haosheng (谭好升), Associate Chief Physician in the Department of Thyroid & Breast Surgery at Taizhou People's Hospital (泰州市人民医院), MD/PhD trained at Peking Union Medical College / Tsinghua / National Cancer Center. I have 20+ years of clinical practice and 10+ SCI-indexed publications in surgical oncology.
That puts me in a useful position for this kind of test: I can read an agent's output and tell whether it is making a defensible statistical claim or hand-waving through one. I am not a computational biologist by training, so I am also exactly the kind of person who would benefit from an agent doing the data plumbing for me.
Full credentials and disclosures: /about.
Get the next case study by email
One case study every few weeks — the task, the protocol, the rubric score, and an honest verdict on the agent. No filler.
Powered by Buttondown. Subscribe to get new AI-agent research evaluations by email.