When your co-author is a lab that runs itself
Materials-lab PI · Science · July 2026
In a self-driving materials lab, an agent proposes the next experiment and files the results — every protocol signed, versioned, and walkable.

Not a memory or an argument — the answer is a diff.
The morning briefing reads like a shift handover, because that is what it is: fourteen runs completed overnight, twelve within tolerance, one failed on a pump fault, and one — run 214 — diverged from the model in a way nobody predicted. The principal investigator reads it standing up. The synthesis rigs did the chemistry while the lab slept; the agent filed every result, linked each run to the hypothesis it was testing, and queued the one genuinely interesting anomaly at the top. The night worked. Now the humans decide what it means.
This is a materials-discovery lab where much of the experimentation is automated — a self-driving lab, in the field's blunt phrase. The automation was never the hard part. The hard part is that science run at machine speed produces machine quantities of record, and a result no one can trace is not a result. Reproducibility is the whole game, and reproducibility is, at bottom, a provenance problem.
So the lab's shared memory is built on attribution. Every protocol in the workspace is signed and versioned; every run the agent files points back to the exact protocol version it executed and the hypothesis it serves. When a result surprises the team — and 214 surprised them — they walk the version history like a timeline: what changed, when, in whose hands, machine or human. The answer to did we alter the anneal step before or after the good runs is not a memory or an argument. It is a diff.
The agent is a colleague in the strict sense: it proposes. The next experiment arrives as an action — sometimes assigned to the agent itself, sometimes to a bench scientist when the step needs hands — with the reasoning attached. The PI approves, redirects, or declines; the proposing never becomes doing without a decision on the record. Mid-run, per-block locks keep the agent and a grad student from editing the same protocol at the same moment, a small mechanical courtesy that has prevented more arguments than any lab meeting.

When the PI wants to interrogate the anomaly, she opens a thread and addresses the agent directly: why did run 214 diverge from the model? What comes back is not an oracle's verdict but a working — the reasoning, the queries it ran, the linked evidence, laid out step by step in the thread. The literature has started calling this provenance reliability rather than debugging. In the lab it goes by a plainer name: the difference between a co-author and a black box.
The human part of the lab has not shrunk; it has moved up. The grad students spend their hours on interpretation and design instead of transcription. The decisions section — what the humans concluded and why — is the most-read part of the workspace, and every entry in it rests on a lattice of signed, walkable records that no one had to compile.
Run 214, three weeks later, becomes the quarter's most promising lead: the divergence was real, reproducible, and traceable to a substrate batch nobody had suspected. The paper is in draft. Somewhere in its methods section is a sentence few PIs ever expected to write with a straight face: the full, versioned experimental record is available — and it is.