Building the AI Scientist
Building the
AI Scientist
A complete scientific system that can compute, interpret, evaluate its own answer, and hand a scientist something they can challenge.
An AI scientist does two things
- 01Computesearch, inference, optimization
- 02Interpretevidence, mechanism, alternatives
Above — the two kinds of work. One field of evidence, resolved two ways. The computation pass converges it toward a single specified, evaluable structure. The interpretation pass reorganizes the identical matter into competing explanations with one clue missing, until a discriminating test kills one and a position survives.
Two problems define what an AI scientist must do
Computation problem versus scientific interpretation problem · the two problems an AI scientist must solve
Examine each problem
Problem 01 of 02 · Computation problem
Given an amino-acid sequence, what three-dimensional structure will the protein adopt?
Navigation across an enormous terrain toward a defined destination. The input and desired output are specified, candidates can be evaluated against learned constraints, and the space of possibilities can be narrowed.
- Input
- Specified — an amino-acid sequence
- Output
- Specified — a candidate structure
- Evaluation
- Confidence, benchmarks, experimental structures
- Capability
- Computation — search, inference, optimization
- Stopping
- Relatively clear
The industry has solved more of this than the other. A system that does this well has valuable computational capability.
The same particles, reorganized. Only the organizing rule changes.
Six things an AI scientist is
A complete scientific system
The system is made out of a model, its tools, and the scientific harness that orchestrates them.
Combines computation and interpretation
It converts data into reproducible signals, then connects those signals to prior knowledge and competing explanations.
Organized around scientific work loops
Built upward from validated vertical workflows, each with its own methods, evidence standards and rubric.
Evaluates and challenges its own output
It seeks contradicting evidence, compares alternatives, and states what would change its conclusion.
Produces reviewable, decision-grade work
A scientific position with evidence, alternatives, uncertainty, provenance and recommended next steps.
Governed by human judgment and provenance
Named reviewers, named decision owners, explicit artifact states, and claim-level traceability throughout.
A system, not a collection of AI features
The nine layers
Layer 01 · Framing
What scientific question is being answered, and what would count as an answer?
Everything downstream is judged against the objective. An AI scientist needs a scientific goal specific enough to constrain method choice, evidence standards and stopping criteria.
What it looks like when the science is being done
Twenty seconds of an omics-to-target run. Watch for the parts a demo usually hides: a confounder caught before interpretation, a branch falsified and closed, a candidate's confidence revised down, contradictions kept rather than resolved away, and a human changing the rubric mid-run.
- stage
- 0/10
- memory
- 0
- papers
- —
- confidence
- 0.71
- define objective
- ingest dataset
- quality control
- differential expression
- pathway interpretation
- candidate extraction
- evidence retrieval
- tractability + safety
- rank + critique
- assemble memo
Scripted from the omics-to-target loop model, played in real time. A real run takes hours to days; this shows the sequence, compressed. What matters is that the failures are visible: this is what a reviewable run looks like, as opposed to a confident answer with no record of how it was reached.


