Architecture
The scientific harness, opened up
- Layers
- 9
- Capabilities
- 45
- Reusable across loops
- 26
- Cross-cutting layers
- 3
Above — composition, exchange, coordination. Nine components, assembled as a pipeline, recomposed as a cycle so evaluation can send work back, given a gate that fires only when its inputs agree, then given the two spines that touch every stage. The last arrangement is the diagram below.
fig. 02 — the architecture, layer by layer
Open a layer and ask what happens without it
The nine layers are drawn as a cycle rather than a pipeline, because planning, execution, evidence and evaluation feed back into planning; memory and human judgment are vertical spines because they are not stages. Every capability states what it does, why it is necessary, and what a system does wrong when it is missing — and which class it belongs to, since that is what explains how a general AI scientist emerges from several validated vertical loops rather than one generic agent.
Four kinds of capability, deliberately distinguished
This distinction is the strategic core of the architecture. It explains how a general AI scientist can emerge from several validated vertical work loops rather than from one generic autonomous agent.
- Generic AI scientist capability
- Reusable infrastructure. Built once, used by every work loop — runtime, memory, provenance, orchestration.
- Scientific skill
- A validated scientific method, reusable across related scientific problems. Versioned and testable, like a library function for science.
- Loop-specific configuration
- Unique to one work loop. The recipe, thresholds and rubric that make omics-to-target different from safety-signal investigation.
- Organization-specific context
- Your therapeutic strategy, internal datasets, platform capabilities, prior failures and scientific preferences. Cannot be bought off the shelf.
Method selection is where the architecture is tested
Two layers make a claim that can be measured: planning selects the method, and evaluation scores how applicable that method is to the data in front of it. On one high-volume task the consensus method holds 92.5% of the literature while the context-appropriate choice rests on five verified papers — and a top-50 retrieval window contains none of them. The measurement, the instrument that flips the choice on a real dataset, and the capabilities that carry it are on their own page.
- head share
- 92.5%
- verified advocacy
- 5
- in the top 50
- 0
See the architecture doing real work
An architecture diagram is a claim. The work-loop explorer shows these layers applied to one scientific problem, stage by stage.