HomeAgents for ResearchCase 1

Case 1 — Reading the data, checking a draft

By Tan Haosheng, MD, PhD · Last reviewed 2026-07-31

Task type · data comprehension~18 min wall-clockWorkBuddy Hy3One shot

I had two unlabelled samples in front of me and a draft of my own I wanted pressure-tested. One prompt, no scaffolding. Below is what the agent did — behaviour and timing only. The science stays out of frame.

⚠️ About this page

The underlying samples, the draft, and every conclusion drawn from them are unpublished. This page reports only: what kind of task was set, what the agent did in observable terms, and how long it took. No screenshot, number, gene, claim or figure from the underlying work is reproduced here in any form. The blurred image below is decorative — a softened snapshot of the agent's output window, with all text and values removed.

The task, paraphrased

Here are two unlabelled files from one of my research samples. I cannot remember which one had the cell-type enrichment step. Can you tell from the data? Separately, here is a draft I wrote — does the data actually support the claims in it?

One natural-language prompt. No JSON scaffold. No tool whitelist. No re-prompting — whatever the agent produced on the first pass is what is scored below.

What the agent did (behaviour only)

Observable sequence

  • Opened and parsed both files on its own.
  • Derived the distinction between the two preparations from the data itself, with a quantified breakdown as evidence — instead of asking me or guessing.
  • Read the draft I attached, claim by claim, and separated the ones the data could carry from the ones it could not.
  • Spotted an aggregation-vs-unit-level confound inside the draft, explained the mechanism, and corrected for technical depth before saying anything was "enriched".
  • Returned sentence-level edits to the over-claimed parts of the draft.
  • Did not flatten the single-patient limitation to please me.

No result, gene, percentage or hypothesis from this run is disclosed on this page.

Agent output window · illustrative · text & values removed A blurred snapshot of the agent's output window, shown only to indicate that output was produced
Why this image is blurred. Every label, axis, value and piece of text has been removed at the pixel level — this image shows only that the agent produced a structured visual output. Anything recoverable from the original is gone.

Time vs. my normal workflow

For comparison, doing the same data-comprehension + draft check in my normal R/Python workflow is a multi-session job. The agent did it in one shot.

Time saved · 3 tasks Bar chart comparing manual workflow time vs. one WorkBuddy Hy3 agent run for three research tasks
Manual workflow vs. one WorkBuddy agent run, per task. Manual times are my own estimate for the same job in my normal R/Python workflow. Agent times are wall-clock from prompt to final report, one shot each.

What this case is — and isn't

  • Is: a record that an AI agent was put on a real research workload and produced a useful behavioural result on the first try.
  • Isn't: a scientific result, a preprint, a benchmark of agent X vs Y, or a recommendation to use any specific tool without your own quality bar.

Case 2 — the mismatch trap test →