CELLULAR PHENOTYPES · MODEL DEVELOPMENT

CellScientist.

Model Revision by
Diagnostic Routing
for Morphological Perturbation Prediction

Turn execution and validation feedback into a decision about what to change next.

FROM FEEDBACK TO A BETTER CANDIDATEFIG. 01
Progress has a history. The retained model follows both improvements and regressions. This case reports search-used validation PCC.

Mengran Li1,2,*Bo Li3,*Jiaying Wang2Wenbin Xing2Chengyang Zhang4Jinlin Wu1Zhen Lei1Jiebo Luo1Stan Z. Li5Zelin Zang1,†

* Equal contribution · † Corresponding authorAffiliations +
  1. Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences
  2. Sun Yat-sen University
  3. University of Macau
  4. Sichuan University
  5. Westlake University

01 / THE IDEA

A score tells you how well.
Diagnostics guide what comes next.

Predicting how a chemical perturbation changes cellular morphology connects molecular interventions to measurable phenotypes. A useful model must coordinate chemical representation, dose, cellular context, and response modeling.

A disappointing score leaves a design decision unresolved. CellScientist connects diagnostics and earlier attempts to local model revisions, keeping the prediction task and evaluation protocol consistent throughout the search.

Every edit belongs to a design state. Every outcome becomes evidence for the next decision.
01

HRT

Keep the evidence.

History-Aware Revision Tracking connects design states to earlier attempts, including failures and regressions.

02

LCA

Make the edit executable.

Limited-Change Application checks candidate eligibility and applies bounded repair before evaluation.

03

PDR

Choose where to revise.

Performance-Discrepancy Refinement maps observed discrepancies to model components for local revision.

THE COMPLETE WORKFLOWFIG. 03
From a model-design state to an executable candidate, and from its outcome to the next edit.

02 / CONTROLLED EVIDENCE

Better revisions.
Fewer candidate evaluations.

A matched comparison isolates the revision policy within a shared 768-candidate language. Four task/split settings, five paired seeds, and a reporting fold held out from search.

Candidate-evaluation budget

CELLSCIENTIST − RANDOM ROUTING

+0.0373

mean held-out PCC advantage

95% confidence interval[0.0292, 0.0454]

A larger mean advantage appears at small evaluation budgets, where each revision has more influence on the search.

Paired-setting t-interval across four settings · Table 3

Held-out prediction quality

Global PCC ↑
CellScientistRandom routing
0.000.100.200.300.40

Bar length: five-seed mean. Whiskers: ±1 sample standard deviation.

The budget counts evaluated candidates, including the shared initial candidate. At ten evaluations, the mean advantage narrows and its confidence interval includes zero.

View all five policies and exact values
Held-out global PCC, mean ± sample standard deviation over five paired seeds. Budget: 5 evaluations.
Task / splitFixed initial modelFlat retryRandom routingAIDE-styleCellScientist
BBBC036 / Plate0.1664 ± 0.00050.2073 ± 0.00560.2277 ± 0.05020.2477 ± 0.02940.2674 ± 0.0003
BBBC036 / SMILES0.1547 ± 0.00050.1882 ± 0.01290.2049 ± 0.04170.2231 ± 0.01690.2346 ± 0.0009
BBBC047 / Plate0.2442 ± 0.00080.2630 ± 0.01300.3066 ± 0.05140.3153 ± 0.03990.3457 ± 0.0004
BBBC047 / SMILES0.2504 ± 0.00040.2916 ± 0.00030.3207 ± 0.05820.3371 ± 0.04170.3614 ± 0.0003
Download chart data (JSON) ↗

04 / INSIDE THE WORKFLOW

Inspect the components,
as well as the final score.

Registered component audits examine what history, candidate checks, and localized repair contribute to the workflow.

0%

Repeated proposals with HRT

170 unique proposals in 170 evaluated steps. Without persistent history: 82 unique proposals in 180 steps.

20 matched controlled runs
40/40

Registered violations handled

LCA repairs 15 cases and blocks 25 others. All 20 valid controls are accepted.

60 registered contract cases
15/15

Held-out faults repaired

Structured PDR repairs all registered held-out faults in one round on average, with no unrelated edits.

15 faults across five component addresses

05 / CODE, DATA & PAPER

Reproduce the workflow.
Build on the revision process.

The paper uses Gemini 3 Pro, with Qwen2.5-0.5B-Instruct for the open-weight controlled audit. Other LLM APIs can be connected through the OpenAI-compatible interface. LLM configuration guide ↗

CITATION

Reference this work

BIBTEX
@misc{li2026cellscientist,
  title = {CellScientist: Model Revision by Diagnostic Routing
           for Morphological Perturbation Prediction},
  author = {Li, Mengran and Li, Bo and Wang, Jiaying and
            Xing, Wenbin and Zhang, Chengyang and Wu, Jinlin and
            Lei, Zhen and Luo, Jiebo and Li, Stan Z. and Zang, Zelin},
  year = {2026},
  howpublished = {Preprint},
  url = {https://github.com/limengran98/CellScientist}
}

CellScientist / Research figure