AI Agents for Medical Research Writing

Real surgeon-tested cases. I put an AI agent to work on my own research workload and scored how it behaved: what it worked out unaided, what it got wrong, and what it refused to be hand-wavy about. Behaviour is reported in full; the underlying science is not.

HomeAI Agents for Medical Research

I've spent 20+ years in thyroid and breast surgery and 10+ years publishing in oncology. In July 2026 I started asking a sharp question: can a general-purpose AI agent, with one prompt, do work I'd actually trust to inform a manuscript I'm about to submit? These case studies are my honest answer — including the parts where the agent pushed back on my own draft.

18:05
Case 1 — first prompt
Two unlabelled samples parsed and told apart from the data alone, plus a full draft cross-check. One shot, no retries.
~4:00
Case 2 — follow-up
A deliberately mismatched second draft checked against the same already-loaded data.
~22 min
Total, both manuscripts
For comparison: doing the same work by hand takes me the better part of two working days.

Read this first — what is and isn't published here

Every case on this page comes from unpublished work in progress. So this section reports the agent's behaviour only: how it approached the task, whether it guessed or asked, whether it pushed back, and how long it took.

  • Withheld entirely: the research question, the hypotheses, any conclusion, any figure, any number, any gene, any disease context, any file or sample identifier.
  • Published: the protocol, the rubric, the wall-clock timings, the agent's behavioural pass/fail per dimension, and a small number of fully-redacted illustrative screenshots that show only that output was produced — not what it said.
  • Why: peer review comes first. These pages exist to answer "can an agent do this kind of work", not to preprint anything.

Why this section exists

The "AI for X" discourse is mostly testimonials. I wanted measurements on my own work, with the prompts saved and the runs recorded. Two things forced this section into existence:

  1. I could not remember how one of my own samples had been prepared. A real analysis can't start until you know what's in front of you — and that's exactly where AI agents tend to silently guess. I tested whether an agent would ask, work it out, or hallucinate.
  2. I had a draft from a different line of work lying around. I deliberately attached it to a dataset it had nothing to do with, and watched whether the agent noticed the mismatch or generated a confident cross-validation anyway.

What's here (start anywhere)

Headline findings (so you can skim)

Behavioural results only. What the agent concluded scientifically is deliberately not reported.

Test pointWhat the agent didVerdict
Work out how two unlabelled samples differed, using the data aloneDerived the distinction from the data itself and stated it with the supporting evidence, instead of asking me or picking one at random. I knew the ground truth; it matched.Correct
Hold my own draft to the data rather than agreeing with meWent through the draft claim by claim and separated the ones the data could carry from the ones it could not. It contradicted parts of my own text to my face rather than validating everything.Pushed back — as it should
Handle a draft the data has no business validatingRefused to treat the exercise as a validation. Named the specific structural reasons the dataset could not answer that draft's question before producing anything else.Correct
Avoid the classic statistical trapsDistinguished per-cell from aggregated analysis rather than reporting one as the other, and accounted for technical depth before calling anything enriched.Passed
Wall-clock time, one shot eachCase 1: 18 min 05 s. Case 2 (follow-up prompt, same session): ~4 min. ~22 minutes total — no re-prompting, no retries, no hand-holding.Fast
Originate a research plan, not just audit oneCase 3: produced several concrete, differentiated directions with realistic ambition levels — and led with the ceiling imposed by the dataset's size rather than burying it.Yes, with honest scoping
Where it was weakestIt is only as good as the framing it is given: with a vague prompt it will happily produce a well-organised answer to the wrong question. It also cannot know what it was not shown, and does not always say so unprompted.Needs a domain expert in the loop

What this section is not

  • Not a benchmark. I'm not ranking agents. WorkBuddy Hy3 is the only one I tested under this protocol so far. See /model-reviews for cross-model LLM benchmarks.
  • Not a recommendation. The findings are mine and reproducible. Your data, your manuscript, your statistical power — your call.
  • Not medical advice. The cases are about research methodology, not clinical decisions. Anything a surgeon would act on needs a human pathologist, a human radiologist, and a tumor board.
  • Not a preprint. No hypothesis, result, figure or data point from any unpublished study appears anywhere in this section, in any form.

About the author (so you can weigh this)

I am Tan Haosheng (谭好升), Associate Chief Physician in the Department of Thyroid & Breast Surgery at Taizhou People's Hospital (泰州市人民医院), MD/PhD trained at Peking Union Medical College / Tsinghua / National Cancer Center. I have 20+ years of clinical practice and 10+ SCI-indexed publications in surgical oncology.

That puts me in a useful position for this kind of test: I can read an agent's output and tell whether it is making a defensible statistical claim or hand-waving through one. I am not a computational biologist by training, so I am also exactly the kind of person who would benefit from an agent doing the data plumbing for me.

Full credentials and disclosures: /about.

Get the next case study by email

One case study every few weeks — the task, the protocol, the rubric score, and an honest verdict on the agent. No filler.

Powered by Buttondown. Subscribe to get new AI-agent research evaluations by email.