Morphology · measurement · scientific use

When is
morphology enough?

Predicting a cellular measurement is one question.
Knowing what you can do with it is another.

MorphoSuff evaluates whether inferred cellular readouts preserve the responses, priorities and biological conclusions that an experiment needs.

A measurement-sufficiency framework, evaluated with paired imaging and perturbation data.
The reporter landscape52 readouts · A549

Recovery is only part of the picture.

Loading the paired-measurement evidence…

Quantitative proxyRanking proxyOther decisions

One point per reporter. Dashed lines mark two reference gates; the full decision also uses rank, amplitude and reliability. Select a point to inspect its evidence.

Real measurements. Explicit uses. Evidence-led decisions.
The scientific question

If a measurement is inferred rather than acquired,
which properties survive — and which uses remain valid?

01 / The approach

From prediction to
measurement decisions.

Paired phase images and fluorescent reporters let us ask not only whether cellular states are predictable, but whether their perturbation responses remain useful.

The OPS atlas provides 9,996,286 paired phase–reporter observations across 1,000 gene knockouts, 52 reporters and 73 screens. One phase-imaged cell can contribute observations for multiple reporters.
01

Recover the response

Evaluate predictions for held-out fields, genes and screens. Distinguish cell-state information from generalization to new perturbations.

02

Check the measurement

Compare response ranking, amplitude, hit recovery and reproducibility of the measured reference.

03

Test the intended use

Replace measured responses with predictions and examine the perturbation priorities and functional results that remain.

02 / Explore the evidence

Different readouts.
Different decisions.

Browse the frozen, ten-method ensemble results. Each assignment belongs to a target–predictor–context combination, evaluated against the study’s reference operating points.

Reporter / biological system

Select a reporter to inspect its measurements.

What do the five decisions mean?

Quantitative proxy meets the study’s recovery, rank, amplitude and strong-hit criteria for a reproducible target.

Ranking proxy meets rank and strong-hit criteria for a reproducible target, but not all quantitative criteria.

Measurement required fails both proxy tiers despite a reproducible target, with recovery below the measurement-required cut-off.

Not identifiable has insufficient or unavailable reference-response reliability to establish sufficiency.

Unresolved has mixed or incomplete evidence under the reference decision rules.

What correlation can miss

A preserved ranking can hide
a compressed response.

Eleven reporters retain less than half the measured response-magnitude variance. A model can therefore preserve a useful ordering while losing the quantitative scale of a perturbation.

The right criterion follows the scientific use.

Ranking and quantitative amplitude are separate properties of the inferred response.
View the complete reporter decision ledger
The paper’s complete reporter ledger, including the distinct evidence used for each decision.
03 / An independent scientific use

Find the perturbations
worth measuring first.

In primary human hepatocytes, brightfield images are used to prioritize compounds with strong metabolic loss — a focused use distinct from preserving the full response profile.

OASIS / primary hepatocytes

Selecting the predicted top 5% retained 10 of the 11 compounds with the largest measured losses among 217 held-out compounds.

Signed, dose-averaged metabolic loss · compound-held-out evaluation
01

Measured priorities, inferred from images

Each point is a held-out compound, with signed metabolic loss relative to same-plate DMSO averaged over eight assayed concentrations. Dashed lines separate the measured and predicted top-5% selections.
02

Put the recovery in context

Ridge and histogram gradient boosting (HGB) compare brightfield features, fluorescence-derived cell counts and acquisition metadata. Filled cells mark recovered hits among the measured top 11; error bars show 95% compound-bootstrap intervals.

One dataset, more than one question. This focused analysis uses brightfield-DINO features to prioritize strong metabolic loss. The broader assessment uses CellProfiler-derived brightfield features to examine response ranking and amplitude. They evaluate different scientific uses with different representations.

The same questions, new measurement settings

A reusable assessment workflow.

Re-fit predictors for the new data, define the intended use, and assess its evidence with the same framework.

Each application yields an evidence profile, rather than a single correlation score. PERISCOPE uses fluorescent inputs; its low guide-partition reliability limits the sufficiency assignment even when predictive recovery is high.
04 / Use MorphoSuff

Your data.
The same questions.

A Python toolkit for paired-measurement assessment: prediction evaluation, response fidelity, reference reliability and use-specific decisions.

Bring a paired dataset.
Leave with an evidence profile.

Use aligned measured and predicted endpoints, perturbation identifiers and control annotations. Add replicate information to assess the reliability of the measured response.

Repository access, including code and documentation, is currently limited to collaborators.

Python ≥ 3.10MIT licenseDataset-adaptable

The paper & resources

When label-free morphology is sufficient for targeted cellular measurements

Mengran Li, Jianqing Zhu, Bo Li, and colleagues

Figures and interactive evidence on this page are drawn from the manuscript’s frozen source data. Each manuscript panel can be expanded and downloaded as a PDF.

Figure

Download PDF ↗

Original figure export · scroll to inspect on smaller screens