Assessment
Is it really an AI scientist?

fig. 01 — the standard, applied
Score a system, and watch a gate refuse it
Thirteen dimensions, five levels. The levels are gated, not score-banded: each one lists the capabilities that carry it, so a system cannot reach a level by scoring well elsewhere if it never falsifies a hypothesis or never stops on scientific grounds. Where a system is held back matters more than the total it scores.
Before you begin
Thirteen dimensions, four minutes
One question per capability. Answers are maturity-based rather than yes-or-no, because almost nothing in a real system is simply present or absent.
What are you assessing?
Can it do the work — sustain a run, use tools, compute a result.
Can it reason — hold context, integrate evidence, build competing explanations.
Can it assess its own science and decide what should happen next.
Can a human trust, challenge, own and use the result.
what gets assessed
- 01Scientific objective definition
- 02Long-running execution
- 03Tool and method use
- 04Branching and parallel exploration
- 05Scientific memory
- 06Evidence integration
- 07Hypothesis generation
- 08Contradiction and falsification
- 09Scientific evaluation
- 10Scientific stopping criteria
- 11Human collaboration
- 12Provenance and reproducibility
- 13Decision-grade output
The model is written to be defensible on its own terms. Most credible scientific AI systems land at level 2 or 3, and no system we are aware of — including our own — fully occupies level 5.