24 / 35JANUARY 2026SCIENTIFIC AI

N24 THE REALITY LAYER

The Case for Researcher-in-the-Loop AI

The best agent does not remove judgment. It makes judgment more leveraged and inspectable.

AUTHORLUCA
READ3 MIN
EVIDENCEPRIMARY-SOURCE GROUNDED
PUBLISHED
ARCHIVE NOTE

Retrospective operator note covering January 2026. Published in September 2026 using public sources and contemporaneous working themes. It was not originally published on the archive date.

IN THIS NOTE · JANUARY 2026

Scientific AI is often framed as a race toward autonomy. Researchers may gain more from systems that collaborate well: planning openly, surfacing uncertainty, accepting correction and preserving the thread of an investigation.

01

Direction is scientific work

Choosing the question, deciding what evidence matters and interpreting an unexpected result are not clerical steps around the research. They are the research. Agents can expand the search space and execute analyses, but they need a mechanism for human context to change the plan.

A checkpoint is valuable when it appears before a consequential branch, not after the final answer.

02

Show the work

Researchers need access to sources, code, assumptions and intermediate outputs. This does not mean flooding the interface with every token. It means preserving an auditable path from question to conclusion.

The agent should also carry context forward. Repeatedly rebuilding project state wastes time and encourages shallow, disconnected answers.

03

Collaboration is a benchmark

A useful system should improve when a researcher challenges it. It should ask questions that change the analysis, acknowledge unresolved contradictions and propose the next discriminating test.

Autonomy remains valuable for bounded work. The broader objective is a research relationship in which humans and agents each make the other's contribution more rigorous.

04

Build an interface for epistemic control

Researcher-in-the-loop should mean more than a final approve button. The interface should expose the plan before expensive work begins, show which sources and assumptions control the conclusion, separate model output from retrieved fact and let the researcher branch the investigation without losing history. Checkpoints belong where expert context can change the path, not at arbitrary intervals designed to make the product appear supervised.

Good friction is selective. Routine transformations can run automatically in a sandbox. A conflicting endpoint definition, uncertain unit conversion or proposed use of sensitive data should stop and ask a specific question. The system should carry the answer into future steps and document its effect. This makes human attention scarce but high leverage. It also produces an audit trail that explains why the final output reflects a collaboration rather than an opaque sequence of model calls.

05

Measure the partnership over time

A collaborative system should improve the quality and speed of the researcher's decisions, not merely reduce keystrokes. Evaluation can measure whether users identify contradictions earlier, reproduce analyses more reliably, design stronger next tests and retain an accurate mental model of uncertainty. It should also measure automation bias: do users accept polished outputs more readily, and does the interface make challenge easier or socially costly?

Longitudinal studies matter because expertise changes the interaction. New users may need explicit scaffolding; experienced researchers may want compressed controls and direct access to intermediate artifacts. The agent should learn project context without turning past decisions into unchallengeable defaults. The best partnership preserves independent judgment on both sides: the model can surface an inconvenient pattern, and the researcher can force the model to revise a convenient story.

OPERATOR LENS
  1. Place human checkpoints before consequential branches.
  2. Preserve provenance and project continuity.
  3. Evaluate how the system responds to correction.
WHAT WOULD CHANGE MY MIND

I would revise this if fully autonomous systems consistently outperformed collaborative systems on open-ended, consequential scientific programs without losing inspectability.

EVIDENCE LEDGER

Primary and institutional sources used as the grounding layer. Interpretation and synthesis are Luca's.

01
BIOS and BioAgentsBIO AI
02
OpenLabsBIO
03
PubMedNational Library of Medicine
04
AI Risk Management FrameworkU.S. National Institute of Standards and Technology