Skip to content
Scientific hypothesis generation

Omics data to target prioritization

Use a transcriptomic dataset and its experimental context to produce five ranked target hypotheses, each carrying supporting and contradicting evidence, an explicit uncertainty statement, and a recommended validation experiment.

dataset

GSE54456

stages

10

checkpoints

5

branch points

21

GSE54456NCBI GEO · Li et al., Journal of Investigative Dermatology, 2014

RNA-seq from lesional skin biopsies of 92 psoriasis patients and normal skin from 82 individuals. The associated study reported 3,577 differentially expressed genes (1,049 up, 2,528 down) and used pathway and weighted co-expression analyses to investigate immune and epidermal programs.

illustrative

The published study aimed to understand disease mechanisms. It did not perform target nomination. The five-target prioritization objective used here is an illustrative extension, not a result of the published work.

fig. 03 — the run, stage by stage

Walk the ten stages the way the system would

Every stage carries its inputs, its methods and the capability class each one belongs to, what it writes to scientific memory, and what it contributes to the work product. Watch for the places the run could have gone wrong: quality control naming a tissue-composition confounder before anything is interpreted, evaluation separating whether the signal is robust from whether it supports a mechanism, and a rubric whose weights change with the objective.

run trace

stage 01 · Interpretation

Define the biological and therapeutic objective

What are we trying to decide, and what would a good answer look like?

what this stage produces
5 ranked target hypotheseshuman geneticsmechanismspecificitysafety scancontradictionsREQUIRED BEFORE A CANDIDATE ADVANCES
The objective, with the evidence dimensions a candidate must satisfy before it can advance. Set before any compute is spent, so evaluation cannot be retrofitted to whatever was found.
Inputs
  • Disease area and therapeutic hypothesis
  • Decision the work feeds and its owner
  • Modality and platform constraints
  • Required output format and review standard
Methods

Objective structuring

Loop

Therapeutic strategy context

Org

Outputs
  • Structured scientific objective
  • Required evidence dimensions
  • Named reviewer and decision owner
  • Disqualifying conditions
Written to memory03
  • Objective statement
  • Success criteria and disqualifying conditions
  • Decision owner and reviewer
Can branch02
  • Genetics-anchored objective: prioritize causal human genetic support
  • Modality-anchored objective: prioritize extracellular accessibility for an antibody program
Disease-area lead01
  • Confirms the objective, the required evidence dimensions and the standard the memo must meet before any compute is spent.
How this stage is judged
  • 01Is the objective specific enough to constrain method selection?
  • 02Are success criteria stated before any analysis runs?
  • 03Is the decision the work feeds named, with an accountable owner?
What it contributes to the memo

The memo opens with the objective it was written to answer, so a reviewer can judge whether it answered that question.

Run memory3 entries
  • 01Objective statement
  • 01Success criteria and disqualifying conditions
  • 01Decision owner and reviewer

These are not notes. They are the premises that constrain what the AI scientist may conclude at later stages, and compression must preserve every one.

Target-prioritization memo with reproducibility pack1/10 sections
  • 01The memo opens with the objective it was written to answer, so a reviewer can judge whether it answered that question.

Assembled throughout the loop. That is what makes claim-level provenance possible.

How a general AI scientist emerges from one loop

Every method above is one of four kinds. The first two are reusable assets that compound across loops. The second two are what make this loop yours.

Generic

Generic AI scientist capability

Reusable infrastructure. Built once, used by every work loop — runtime, memory, provenance, orchestration.

Skill

Scientific skill

A validated scientific method, reusable across related scientific problems. Versioned and testable, like a library function for science.

Loopin use

Loop-specific configuration

Unique to one work loop. The recipe, thresholds and rubric that make omics-to-target different from safety-signal investigation.

Orgin use

Organization-specific context

Your therapeutic strategy, internal datasets, platform capabilities, prior failures and scientific preferences. Cannot be bought off the shelf.

How mature is the system you use?

Assess a product, an internal system or your organization against the capability model this loop depends on.

Assess a system