Benchmarking AI Agents in Long-Horizon Single-Cell Biology

Practical tasks in single-cell biology require complex analytical workflows, multiple sources of experimental data and creative thinking in the context of previous literature.
Biological research agents are developing rapidly. But integrating them into productive R&D workflows requires a strong balance of power.
We present scBench-Long, a long-horizon benchmark for single-cell biology where agents must derive scientific conclusions from raw or near-raw data without specified methods.

The benchmark consists of 21 tests including melanoma CD8 T-cell reactivity, CD8 RNA+ATAC regulatory inference, development of a human monkey chimera, KRAS-driven lung tumor aging, and lethal lung disease COVID-19. Of the 1,068 completed trajectories, powerful agents exceeded 16/63 runs (25.4%).
View results, including interactive agent trajectory viewers, at benchmarks.bio. Read the full manuscript.
Agentic AI-biology benchmarks have made rapid progress, but most focus on broad biological knowledge. Single cell benchmarks are available focus on local analytical steps and do not explore the practical science of the long horizon.
scBench-Long tests whether agents can go beyond accurately performing local analytical steps to thinking about the kinds of real-world scientific activities found in published studies or used in drug programs.

The benchmark builds on it building a benchmark with preconditions rather than introducing a new standard grading framework.
Each assessment includes:
-
Raw or near data
-
Task description: measured experimental context, anonymized labels where needed and a unified scientific question
-
The vocabulary of the final regulated response, usually including cell or genetic markers and regulatory signals
-
Production notes
-
Trajectory rubrics
The paper’s claims are used as sources of human evaluation, not as automatic ground truth. Candidate jobs are retained only when the target conclusion can be reproduced from staged data and is refined through review, bug design, and route testing from multiple model families.
We tested 17 wire model configurations for all 21 final tests, generating 1,068 completed trajectories. Scores use deterministic pass/fail grading of the final constructed response.

The strongest pair of harnesses was Claude Opus 4.8 and Claude Code, who won 16/63 runs, or 25.4%. Gemini 3.5 Flash with PI followed on 14/63, or 22.2%. GPT-5.5 reached 10/63 under both OpenAI Codex and PI, but both harnesses solved different tests.

Endpoint grading is stable and repeatable, but small. The model can fail the final answer while completing many useful intermediate steps. To understand this partial progress, scBench-Long also compiles the route-specific rubrics obtained by the multi-judge models.

Rubric scores are useful but not perfect. Endpoint-passing trajectories had higher rubric scores than endpoint-failing trajectories: 75.9% versus 58.7%. The highest decile of the scoring rubric was enriched with success, but contained many failures. Decisive final response grading is still our preferred method of reporting placement results, while rubric ratings are used as diagnostic tools.
The most interesting failures were the gaps in scientific thinking where the agents used to get the right intermediate results – the number of cells, the binding ligand-receptor interaction, the immune receptor patterns, or the regulatory systems – but did not integrate them into the correct biological application.
Some emerging patterns:
Standard biological material can extract evidence from real data. For melanoma CD8 activity, agents often relied on markers such as CD39, PD-1, TOX, LAG3, HAVCR2, and CXCL13 instead of combining tumor immune/APC interactions with clonal expansion of the paired TCR.
Immature abundance can be mistaken for biological importance. In the human-monkey chimera ligand-receptor chimera activity, agents tend to select the largest class of interactions even though no single signaling family dominates after accounting for guidance and database design.
Molecular correlation can be mistaken for a method. In tumor-T-cell conjugate activities, agents target tumor-rich systems but treat the association after antibody pairing as evidence of a targeted approach.
Agents fail to consult in various ways and rely on e.g. Only RNA. CD8 RNA+ATAC activities required testing whether regulatory patterns of RNA levels were consistent with chromatin accessibility data.
scBench-Long tests “agent systems”: model-harness pairs. The harness can have a dramatic effect on the performance of the base model.
Claude Opus 4.8 performed better with Claude Code than PI: 16/63 compared to 13/63. GPT-5.4 performed better with PI than OpenAI Codex: 6/63 versus 2/63. GPT-5.5 had the same combined pass rate under the OpenAI Codex and PI, but resolved different tests and had different job level strictness.

Harnesses affect how the model searches, writes code, checks intermediate files, recovers from errors, etc. We found open source harnesses with several tools that promote the most efficient (time + cost) trajectories in science.
We hope that scBench-Long serves as both a benchmarking tool and a diagnostic lens for developing agents that analyze single-cell data reliably, transparently, and reproducibly.
It is a focused offering within a broad measurement family that includes major classes of biological data and functional categories, including spatial biology, epigenomics, and medicine.
Broadly, we view benchmarks as emerging specifications of computational biology workflows, which support the test-driven development of agent systems whose behavior can be improved in both model training and network engineering.
Manuscript: latch.bio/scbench-long
Leaderboard: benchmarks.bio



