Pillar II · The evidence environment

Design what evidence systems acquire

Evidence Engineering and Closed-Loop Discovery & Design

Scientific progress can be constrained by evidence that is abundant but poorly specified, biased toward convenient measurements, or unable to distinguish competing explanations. More data does not automatically create a better scientific decision.

This Pillar treats evidence as a designed component of the discovery system. We examine what is known, identify competing explanations, and select the next dataset, perturbation, candidate or experiment for the uncertainty it can resolve.

Core question
What evidence should the system create next?
Research focus
Design the evidence environment surrounding the scientific learner.
Intended contribution
Datasets, evidence models, acquisition strategies and discovery loops that connect learning to the next scientific decision.

Try choosing the next experiment

Published example · Evidence engineering

Learning where to search next

The choice of the next calculation becomes part of the learning system. This published study connects evidence acquisition, model updates and molecular design.

Which molecule should we calculate next?

Bayesian optimisation and first-principles calculations form an active-learning loop, using new computed properties to improve predictions and guide the search for photosensitizers.

Molecules in the search space
>7 million
Potential high-performance candidates identified
5,357
Photosensitizers synthesized and evaluated
4

The shortlist contains computational candidates; experimental evaluation covered four synthesized photosensitizers.

A schematic search space separates labelled and uncertain regions. Blue squares select structures using uncertainty and low singlet–triplet energy gaps, while an arrow marks repeated active-learning cycles that improve the model.

Self-improving photosensitizer discovery: Bayesian search connects uncertainty, predicted properties and active learning.

Discovery-system diagram · Shidang Xu et al., J. Am. Chem. Soc. 143 (2021), 19769–19777. Existing author-website image; only display size and format are adjusted. · Published article · ACS author reuse policy

View full-size figure
Explore the wider discovery loop
Conceptual framework A shared way to reason about discovery. Projects may revisit or combine steps; new evidence can change any earlier choice.

Open a step for an example question and the research behind it.

  1. Observe Establish what is known and uncertain.

    Bring measurements into context, including how they were obtained, what they cover and where they may be unreliable.

    Example question
    Which observations are comparable, and what is still missing?
    Typical research output
    Context-aware observations with uncertainty and source information.

    Contributing research

  2. Question Make competing explanations explicit.

    Identify a scientific distinction that matters, then ask which alternative structures or mechanisms could explain the same observations.

    Example question
    Could a different interaction or mechanism produce the same pattern?
    Typical research output
    Competing explanations and their distinguishable consequences.

    Contributing research

  3. Design evidence Choose what could distinguish the alternatives.

    Select a measurement, perturbation or candidate for what it can reveal, within the available scientific and practical constraints.

    Example question
    Which next experiment would best separate these explanations?
    Typical research output
    An evidence-acquisition or experiment-design proposal.

    Contributing research

  4. Learn Update the system from the new evidence.

    Revise the learning system while examining stability, uncertainty and the alternatives that remain unresolved.

    Example question
    What should change in the learner after this observation?
    Typical research output
    A revised model with explicit assumptions and uncertainty.

    Contributing research

  5. Test Challenge predictions and explanations.

    Use controls, new conditions and independent evidence where available to probe whether an explanation survives plausible alternatives.

    Example question
    Does the explanation survive matched controls and changed conditions?
    Typical research output
    A comparison showing supported claims, failures and limits.

    Contributing research

  6. Explain or design State what is supported and what to test next.

    Form a testable explanation or a constrained design proposal. Keep the conditions and limits visible; a prediction alone does not establish a mechanism or a successful design.

    Example question
    What consequence or design choice could be tested next?
    Typical research output
    A testable explanation or design proposal with stated limits.

    Contributing research

New evidence changes the next question. Observations, failed tests and revised constraints return to the learner. Revisit the question, evidence choice or model when an explanation no longer holds.

Return to Observe

Solid arrows follow the reading sequence. The dashed path returns evidence to the next question.

Try the reasoning

Which evidence would tell them apart?

Two explanations can fit the same observations. Choose the next experiment, compare their predictions, and explore what a possible result would change.

Hypothetical molecular system · qualitative predictions, not experimental data or results from the study above.

What we have seen

  • Feature present · assembledHigh signal
  • Feature absent · dispersedLow signal

Explanation A

The molecular feature drives the signal.

Signal is high when the feature is present, regardless of assembly state.

Explanation B

Assembly drives the signal.

Signal is high when molecules are assembled, regardless of the feature.

Both simplified explanations fit these observations because the feature and assembly state vary together.

Read all three prediction comparisons

Repeat the original condition

Feature present · assembled

A: high · B: high

Both explanations predict a high signal. Replication can improve precision, but these predictions do not separate.

Change both factors together

Feature absent · dispersed

A: low · B: low

Both explanations predict a low signal. Changing two factors together leaves their contributions unresolved.

Change only the assembly state

Feature present · dispersed

A: high · B: low

A predicts high; B predicts low. This experiment can distinguish them if dispersion changes only assembly and the measurement resolves the difference.

A result matching both predictions leaves the comparison unresolved. A result matching neither calls for checks and revised explanations. An inconclusive result does not select a winner.

What makes this comparison valid?

These are two deliberately simplified explanations, not an exhaustive set of mechanisms. The independent perturbation assumes the feature, concentration, measurement conditions and other relevant factors remain unchanged. Real measurements have uncertainty; distinguishing high from low requires sufficient precision and replication. Experimental feasibility and cost also matter.

See the broader evidence-choice diagram

Choose evidence that separates explanations.

When different explanations fit what we already know, the next useful test is one that could tell them apart.

Two explanations fit current observations. A feasible perturbation leads to different predictions. Measurement with controls helps reweight the explanations, or leaves ambiguity that informs the next test.
Conceptual example, not experimental data. A difference between predictions is useful only if the measurement can resolve it under relevant controls and uncertainty. Positive, negative and inconclusive results all inform the next choice.

View full-size diagram: Wide layout Vertical layout

An amber candidate is selected from translucent forms for an illustrative measurement. An observation tile connects to a blue relational model; a return path leads toward the candidates, with a faint alternative model still visible.

A conceptual illustration of choosing an informative candidate, gathering an observation, and revising a model to guide the next question. Alternatives remain open as the cycle continues.

Closed-loop conceptual framework

From an evidence gap to a verified update

Decision context leads to evidence quality, competing hypotheses, acquisition and measurement. Results feed back into the next decision, closing the evidence loop.

View full-size framework: Wide layout Vertical layout

Read the framework step by step
  1. 01

    Decision context

    Define the scientific question, endpoint, boundary, and practical constraints.

  2. 02

    Evidence quality

    Understand measurements, their origins, coverage, bias and unresolved gaps.

  3. 03

    Competing hypotheses

    State alternatives and the observations that would distinguish them.

  4. 04

    Acquire or create

    Select a dataset, perturbation, candidate, or experiment for its expected information value.

  5. 05

    Measure and update

    Learn from positive, negative and inconclusive results when choosing the next action.

This diagram describes the decision logic of Pillar II. It is a conceptual framework, not a completed autonomous loop or an experimental result.

Scope

Evidence becomes a designed part of the system.

Evidence is evaluated by whether it is trustworthy, reusable, and capable of distinguishing explanations—not simply by how much of it is available.

01

Evidence quality

Evaluate measurement context, provenance, missingness, bias, comparability, and the decision boundary before model development.

02

Structured evidence

Build datasets that retain measurement context and uncertainty.

03

Active acquisition

Rank samples, perturbations, candidates, or experiments by the ambiguity they can resolve under real constraints.

04

Closed-loop update

Connect candidate generation, measurement and model updates so that each result informs the next choice.

Research contributions

Choose evidence that moves discovery forward.

We develop datasets, acquisition strategies and discovery loops that help distinguish explanations and guide the next experiment.

Method families under study

  • Data quality, measurement context and uncertainty analysis
  • Endpoint- and context-aware dataset construction
  • Mechanism-discriminating perturbation design
  • Active learning, information gain, disagreement, and out-of-distribution analysis
  • Constraint-aware de novo and inverse design
  • Prospective verification and closed-loop evidence integration

What we aim to develop

Evidence specification
A clear definition of measurements, their context, uncertainty and the decisions they can inform.
Structured evidence memory
A dataset that retains positive, negative, contradictory and inconclusive results, together with their measurement context.
Acquisition policy
A strategy for selecting informative samples, perturbations, candidates or experiments under scientific and practical constraints.
Discovery loop
A sequence of proposed candidates, measurements and model updates that guides the next scientific decision.

Representative testbeds

Evidence strategy must respect scientific context.

Measurements, feasible perturbations, constraints, and failure modes differ across testbeds. The shared contribution is the logic for deciding what evidence is valuable next.

  • 01Molecular interactions and drug discovery
  • 02Proteins, peptides, and sequences
  • 03Biomaterials and delivery systems
  • 04Complex scientific data and mechanisms

Research horizon

From informative measurements to adaptive discovery.

These are planning windows. Expansion will follow research progress and validation.

  1. Now

    Mechanistic questions in molecular and materials science

    We develop and test AI methods for mechanism understanding, discovery, and design through scientific questions and feedback from experiments and simulation.

  2. Over approximately five years

    Building reusable scientific capabilities

    We aim to develop reusable representation and model methods, scientific system architectures, research infrastructure, evidence-acquisition strategies, and evaluation systems that support mechanism understanding and reliable design.

  3. Over approximately five to ten years

    Testing principles across systems

    We aim to extract mathematical, computational, and system-design principles and test how these methods and capabilities transfer and combine across broader scientific fields.

Scientific principles

Does the next measurement change what can be learned?

Understand each measurement

Understand how measurements were obtained, what they represent and where uncertainty affects their use in learning and design.

Preserve negative and failure evidence

Negative, contradictory and inconclusive results can reveal limitations, challenge explanations and guide better experiments.

Compare acquisition logic

A proposed next experiment or candidate should be judged against transparent alternative selection strategies and practical constraints.

Separate proposal from verification

Test generated candidates and model predictions with measurements that can confirm, refine or challenge their expected effects.