FROM MODEL DISCOVERY TO TESTABLE INPUT-USE CLAIMS

Discover. Falsify. Revise.

Auditing Input-Use Claims from Source Code to Predictive Contribution
in Agent-Discovered Cell Models

Does a cell model use the perturbation it was given?
CellAudit connects the computation in source code to a fitted model’s behavior, then tests whether that behavior helps predict cellular responses.

Mengran Li1Bo Li2Chengyang Zhang3Yang Yan4Jinfeng Xu5Zhenchao Tang6

1 Sun Yat-sen University2 University of Macau3 Sichuan University4 Zhejiang University5 University of British Columbia6 Tencent AI Lab

01 / THE FRAMEWORK

A prediction score is the beginning
of the audit.

Three questions connect an input-use claim to the computation, the fitted predictor, and the observed response.

CellAudit framework: discover candidate models, falsify input-use claims with source inspection and replacement tests, and return evidence to model revision.
Discover → falsify → revise. Each claim is tied to a specific input, source path, checkpoint, and replacement protocol.
01

Source consumption

Can the input enter the computation?

Inspect the registered claim and its cited source path. An apparently conditional operation can be mathematically inactive.

02

Fitted dependence

Do predictions change with the input?

Replace one input while fixing the checkpoint and all other inputs. Measure the resulting prediction distance.

03

Predictive contribution

Does the original input help?

Compare target loss and predictive score under the correct and replacement inputs.

An executable path, prediction sensitivity, and predictive benefit answer different questions. Replacement effects are defined by the registered protocol.

THE COUNTEREXAMPLE

A strong score.
An unused input.

A score-selected model predicts cellular responses almost as well as its control-only counterpart. Yet replacing compound identity leaves its predictions exactly unchanged.

Source inspection explains why: the compound query attends to a single key–value pair. Its normalized attention weight is always one.

Read the source-to-behavior analysis ↗
BBBC047 · Held-out Fold 5Global PCC ↑
Score-selected predictor0.3153
Control-profile-only predictor0.3142
00.35
0Prediction change
under tested compound replacements

Five paired refits. Full − control-only PCC: +0.0011
95% paired-seed CI [−0.0005, +0.0027]. Compound invariance also persists with physically disjoint control wells.

02 / THE EVIDENCE

Sensitivity is common.
Benefit is more selective.

Across 48 sampled candidates, 47 change predictions under compound replacement on both folds. Twenty show a positive target-loss interval on both folds.

Explore the candidate audit

One square per generated source. Select a candidate to inspect its evidence.

Positive loss CI on both foldsSensitive; loss CI not positive on bothInvariant

48 sources · 20 positive on both folds · 27 other sensitive · 1 invariant

Dataset–policy-balanced, source-structure-stratified sample: 48 sources and 96 same-checkpoint fold evaluations. Target-loss intervals resample source scaffolds, conditional on checkpoints and registered donor maps. Counts describe this sample.

Candidate audit: prediction sensitivity, target-loss contribution, source-path outcomes, and replication across held-out folds.
From sensitivity to contribution. Prediction changes alone do not establish predictive benefit. The paper reports source-path findings alongside same-checkpoint replacement tests.

03 / FROM DIAGNOSIS TO REVISION

Revise the path.
Then test its contribution.

An audit can become feedback for model discovery: make the route executable, measure its effect, and learn a useful increment over a control-only predictor.

Prediction-score selection

A score without compound dependence

0.3153 PCC

Compound effect 0.0000

Path-constrained discovery

An explicit compound route

0.2922 PCC

Compound effect +0.0074

Falsification-guided discovery

A contribution-aware residual

0.3173 PCC

Compound effect +0.0048

BBBC047, Fold 5. These are three discovery settings. Prediction-score means summarize five refits; constrained and guided means summarize ten trajectory endpoints. Compound effects are factorial PCC effects averaged over the other input’s two states. Positive mean contributions are distinct from the paper’s reference-relative qualification criterion.

Return the audit to the search loop.

Five paired sci-Plex searches compare score feedback with an audit-enriched feedback package under matched training and selection rules.

Matched sci-Plex feedback: five paired search trajectories, predictive performance, input contributions, and uncertainty.
Audit-enriched search feedback. Higher mean held-out performance and larger mean compound and dose contributions on both folds; paired intervals retain the uncertainty at five trajectories.

04 / ACROSS ACQUISITIONS

How far does an
input-use claim travel?

LINCS-selected model designs are frozen and refit on the independently acquired LKCP cohort. Predictive gains and dose contribution persist; compound identity does not meet the same criterion.

LKCP REFERENCE-RELATIVE CRITERION

Same model designs.
A new acquisition.

50/50

Dose · passing refits

0/50

Compound identity · passing refits

Counts hold on each of the two LKCP evaluation folds. Ten selected endpoint instances contain six unique configurations, with five refits per instance. A positive mean compound allocation on Fold 5 does not imply that the reference-relative criterion is met.

Independent-acquisition transfer from LINCS to LKCP: predictive gains, compound and dose allocations, and criterion support.
Transporting designs, testing claims again. Effects and 95% intervals summarize ten endpoint instances after averaging refits within each instance. Compound and dose effects are Shapley allocations relative to both inputs replaced.
Look closer: BBBC047 audit and a revision trajectory +

05 / OPEN THE EVIDENCE

Inspect the claim.
Reproduce the evidence.

Follow the repository’s task contracts, frozen-model audit, and reproducibility guide. Extend the task registry, candidate language, or LLM provider through documented interfaces.

Start with the repository.

Install the package, then validate the registered configuration. The quickstart walks through data preparation, discovery, and model auditing.

Full setup instructions ↗

Download chart data ↓

Terminal
git clone https://github.com/limengran98/CellAudit.git
cd CellAudit
python -m venv .venv
# Activate .venv for your operating system.
python -m pip install -e .
cellaudit validate

Cite CellAudit

Discover, Falsify, Revise: Auditing Input-Use Claims from Source Code to Predictive Contribution in Agent-Discovered Cell Models

@misc{li2026cellaudit,
  title = {Discover, Falsify, Revise: Auditing Input-Use
    Claims from Source Code to Predictive Contribution
    in Agent-Discovered Cell Models},
  author = {Li, Mengran and Li, Bo and Zhang, Chengyang
    and Yan, Yang and Xu, Jinfeng and Tang, Zhenchao},
  year = {2026},
  url = {https://limengran98.github.io/CellAudit}
}
RELATED PROJECT

CellScientist

Model revision by diagnostic routing for morphological perturbation prediction.

Explore project ↗

Research figure