Skip to content

Assessment

Is it really an AI scientist?

Systems described as AI scientists differ enormously in what they actually do. This assessment separates the capabilities that make a system fast from the ones that make it scientifically trustworthy.
Five polished steel blocks in a row, each taller than the last, the tallest carrying an indigo inset.
Five levels, and each is a real destination rather than a failure to reach the next. Most credible scientific AI systems sit at two or three; no system we are aware of — including our own — fully occupies five.

fig. 01 — the standard, applied

Score a system, and watch a gate refuse it

Thirteen dimensions, five levels. The levels are gated, not score-banded: each one lists the capabilities that carry it, so a system cannot reach a level by scoring well elsewhere if it never falsifies a hypothesis or never stops on scientific grounds. Where a system is held back matters more than the total it scores.

Before you begin

Thirteen dimensions, four minutes

One question per capability. Answers are maturity-based rather than yes-or-no, because almost nothing in a real system is simply present or absent.

What are you assessing?

Execution

Can it do the work — sustain a run, use tools, compute a result.

Scientific cognition

Can it reason — hold context, integrate evidence, build competing explanations.

Scientific judgment

Can it assess its own science and decide what should happen next.

Governance and output

Can a human trust, challenge, own and use the result.

what gets assessed

  • 01Scientific objective definition
  • 02Long-running execution
  • 03Tool and method use
  • 04Branching and parallel exploration
  • 05Scientific memory
  • 06Evidence integration
  • 07Hypothesis generation
  • 08Contradiction and falsification
  • 09Scientific evaluation
  • 10Scientific stopping criteria
  • 11Human collaboration
  • 12Provenance and reproducibility
  • 13Decision-grade output
No sign-up. Nothing is submitted.

The model is written to be defensible on its own terms. Most credible scientific AI systems land at level 2 or 3, and no system we are aware of — including our own — fully occupies level 5.