Skip to content

Building the AI Scientist

Building the
AI Scientist

A complete scientific system that can compute, interpret, evaluate its own answer, and hand a scientist something they can challenge.

An AI scientist does two things

  • 01Compute
  • 02Interpret

Above — the two kinds of work. One field of evidence, resolved two ways. The computation pass converges it toward a single specified, evaluable structure. The interpretation pass reorganizes the identical matter into competing explanations with one clue missing, until a discriminating test kills one and a position survives.

01Define the category

Two problems define what an AI scientist must do

An AI scientist has to solve two different kinds of problem: one is a computation problem, the other an interpretation problem. Both are hard. The industry has solved far more of the first than of the second.
fig. 01

Computation problem versus scientific interpretation problem · the two problems an AI scientist must solve

Examine each problem

Problem 01 of 02 · Computation problem

Given an amino-acid sequence, what three-dimensional structure will the protein adopt?

Navigation across an enormous terrain toward a defined destination. The input and desired output are specified, candidates can be evaluated against learned constraints, and the space of possibilities can be narrowed.

Input
Specified — an amino-acid sequence
Output
Specified — a candidate structure
Evaluation
Confidence, benchmarks, experimental structures
Capability
Computation — search, inference, optimization
Stopping
Relatively clear

The industry has solved more of this than the other. A system that does this well has valuable computational capability.

The same particles, reorganized. Only the organizing rule changes.

02The definition

Six things an AI scientist is

A working definition the market can hold onto — and test any system against.
01

A complete scientific system

The system is made out of a model, its tools, and the scientific harness that orchestrates them.

02

Combines computation and interpretation

It converts data into reproducible signals, then connects those signals to prior knowledge and competing explanations.

03

Organized around scientific work loops

Built upward from validated vertical workflows, each with its own methods, evidence standards and rubric.

04

Evaluates and challenges its own output

It seeks contradicting evidence, compares alternatives, and states what would change its conclusion.

05

Produces reviewable, decision-grade work

A scientific position with evidence, alternatives, uncertainty, provenance and recommended next steps.

06

Governed by human judgment and provenance

Named reviewers, named decision owners, explicit artifact states, and claim-level traceability throughout.

03Make the architecture understandable

A system, not a collection of AI features

9 layers and 45 capabilities. Work loops rather than flows: evaluation returns work to execution, evidence and planning as often as it advances it, and nothing reaches a work product without passing a scientific stopping decision.

The nine layers

cycling

Layer 01 · Framing

What scientific question is being answered, and what would count as an answer?

Everything downstream is judged against the objective. An AI scientist needs a scientific goal specific enough to constrain method choice, evidence standards and stopping criteria.

04A run, in real time

What it looks like when the science is being done

Twenty seconds of an omics-to-target run. Watch for the parts a demo usually hides: a confounder caught before interpretation, a branch falsified and closed, a candidate's confidence revised down, contradictions kept rather than resolved away, and a human changing the rubric mid-run.

stage
0/10
memory
0
papers
confidence
0.71
t = 0.0s1 branch open
  1. define objective
  2. ingest dataset
  3. quality control
  4. differential expression
  5. pathway interpretation
  6. candidate extraction
  7. evidence retrieval
  8. tractability + safety
  9. rank + critique
  10. assemble memo
run logexecuting

Scripted from the omics-to-target loop model, played in real time. A real run takes hours to days; this shows the sequence, compressed. What matters is that the failures are visible: this is what a reviewable run looks like, as opposed to a confident answer with no record of how it was reached.

05Essays

Read the argument in full

An essay series developing the category from first principles. Each essay takes one part of the argument above and works it through — start wherever the question you brought here lives.