HRT
Keep the evidence.
History-Aware Revision Tracking connects design states to earlier attempts, including failures and regressions.
CELLULAR PHENOTYPES · MODEL DEVELOPMENT
Turn execution and validation feedback into a decision about what to change next.
01 / THE IDEA
Predicting how a chemical perturbation changes cellular morphology connects molecular interventions to measurable phenotypes. A useful model must coordinate chemical representation, dose, cellular context, and response modeling.
A disappointing score leaves a design decision unresolved. CellScientist connects diagnostics and earlier attempts to local model revisions, keeping the prediction task and evaluation protocol consistent throughout the search.
HRT
History-Aware Revision Tracking connects design states to earlier attempts, including failures and regressions.
LCA
Limited-Change Application checks candidate eligibility and applies bounded repair before evaluation.
PDR
Performance-Discrepancy Refinement maps observed discrepancies to model components for local revision.
02 / CONTROLLED EVIDENCE
A matched comparison isolates the revision policy within a shared 768-candidate language. Four task/split settings, five paired seeds, and a reporting fold held out from search.
CELLSCIENTIST − RANDOM ROUTING
mean held-out PCC advantage
A larger mean advantage appears at small evaluation budgets, where each revision has more influence on the search.
Paired-setting t-interval across four settings · Table 3
The budget counts evaluated candidates, including the shared initial candidate. At ten evaluations, the mean advantage narrows and its confidence interval includes zero.
| Task / split | Fixed initial model | Flat retry | Random routing | AIDE-style | CellScientist |
|---|---|---|---|---|---|
| BBBC036 / Plate | 0.1664 ± 0.0005 | 0.2073 ± 0.0056 | 0.2277 ± 0.0502 | 0.2477 ± 0.0294 | 0.2674 ± 0.0003 |
| BBBC036 / SMILES | 0.1547 ± 0.0005 | 0.1882 ± 0.0129 | 0.2049 ± 0.0417 | 0.2231 ± 0.0169 | 0.2346 ± 0.0009 |
| BBBC047 / Plate | 0.2442 ± 0.0008 | 0.2630 ± 0.0130 | 0.3066 ± 0.0514 | 0.3153 ± 0.0399 | 0.3457 ± 0.0004 |
| BBBC047 / SMILES | 0.2504 ± 0.0004 | 0.2916 ± 0.0003 | 0.3207 ± 0.0582 | 0.3371 ± 0.0417 | 0.3614 ± 0.0003 |
03 / MODELS & TRAJECTORIES
Selected-model benchmarks and open search trajectories show how the workflow develops task-specific predictors. BBBC036, BBBC047, and CPG0016 form the morphology benchmark suite; transcriptomic and temporal tasks extend the evaluation.
The BBBC047 trajectory reports validation scores used during search. The controlled comparison above uses a separate held-out reporting fold.
04 / INSIDE THE WORKFLOW
Registered component audits examine what history, candidate checks, and localized repair contribute to the workflow.
170 unique proposals in 170 evaluated steps. Without persistent history: 82 unique proposals in 180 steps.
20 matched controlled runsLCA repairs 15 cases and blocks 25 others. All 20 valid controls are accepted.
60 registered contract casesStructured PDR repairs all registered held-out faults in one round on average, with no unrelated edits.
15 faults across five component addresses05 / CODE, DATA & PAPER
git clone https://github.com/limengran98/CellScientist.git
cd CellScientist
python -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[data]"
python -m cellscientist --helpFull setup, including Windows instructions ↗The paper uses Gemini 3 Pro, with Qwen2.5-0.5B-Instruct for the open-weight controlled audit. Other LLM APIs can be connected through the OpenAI-compatible interface. LLM configuration guide ↗
CITATION
@misc{li2026cellscientist,
title = {CellScientist: Model Revision by Diagnostic Routing
for Morphological Perturbation Prediction},
author = {Li, Mengran and Li, Bo and Wang, Jiaying and
Xing, Wenbin and Zhang, Chengyang and Wu, Jinlin and
Lei, Zhen and Luo, Jiebo and Li, Stan Z. and Zang, Zelin},
year = {2026},
howpublished = {Preprint},
url = {https://github.com/limengran98/CellScientist}
}