N24 THE REALITY LAYER
The Case for Researcher-in-the-Loop AI
The best agent does not remove judgment. It makes judgment more leveraged and inspectable.
IN THIS NOTE · JANUARY 2026
Scientific AI is often framed as a race toward autonomy. Researchers may gain more from systems that collaborate well: planning openly, surfacing uncertainty, accepting correction and preserving the thread of an investigation.
Direction is scientific work
Choosing the question, deciding what evidence matters and interpreting an unexpected result are not clerical steps around the research. They are the research. Agents can expand the search space and execute analyses, but they need a mechanism for human context to change the plan.
A checkpoint is valuable when it appears before a consequential branch, not after the final answer.
Show the work
Researchers need access to sources, code, assumptions and intermediate outputs. This does not mean flooding the interface with every token. It means preserving an auditable path from question to conclusion.
The agent should also carry context forward. Repeatedly rebuilding project state wastes time and encourages shallow, disconnected answers.
Collaboration is a benchmark
A useful system should improve when a researcher challenges it. It should ask questions that change the analysis, acknowledge unresolved contradictions and propose the next discriminating test.
Autonomy remains valuable for bounded work. The broader objective is a research relationship in which humans and agents each make the other's contribution more rigorous.
Build an interface for epistemic control
Researcher-in-the-loop should mean more than a final approve button. The interface should expose the plan before expensive work begins, show which sources and assumptions control the conclusion, separate model output from retrieved fact and let the researcher branch the investigation without losing history. Checkpoints belong where expert context can change the path, not at arbitrary intervals designed to make the product appear supervised.
Good friction is selective. Routine transformations can run automatically in a sandbox. A conflicting endpoint definition, uncertain unit conversion or proposed use of sensitive data should stop and ask a specific question. The system should carry the answer into future steps and document its effect. This makes human attention scarce but high leverage. It also produces an audit trail that explains why the final output reflects a collaboration rather than an opaque sequence of model calls.
Measure the partnership over time
A collaborative system should improve the quality and speed of the researcher's decisions, not merely reduce keystrokes. Evaluation can measure whether users identify contradictions earlier, reproduce analyses more reliably, design stronger next tests and retain an accurate mental model of uncertainty. It should also measure automation bias: do users accept polished outputs more readily, and does the interface make challenge easier or socially costly?
Longitudinal studies matter because expertise changes the interaction. New users may need explicit scaffolding; experienced researchers may want compressed controls and direct access to intermediate artifacts. The agent should learn project context without turning past decisions into unchallengeable defaults. The best partnership preserves independent judgment on both sides: the model can surface an inconvenient pattern, and the researcher can force the model to revise a convenient story.
- Place human checkpoints before consequential branches.
- Preserve provenance and project continuity.
- Evaluate how the system responds to correction.
I would revise this if fully autonomous systems consistently outperformed collaborative systems on open-ended, consequential scientific programs without losing inspectability.
Primary and institutional sources used as the grounding layer. Interpretation and synthesis are Luca's.
01