Benchmarking AI Agents in Pathogen Genomic Surveillance

Introducing BioSecBench-Surveillance, a validation benchmark to test whether AI agents can make the analytical decisions required for pathogen genomic testing.
The benchmark consists of 100 tests including seven task categories, six sample types, and short sequences and long readings. Agents receive real-time sequencing data and a negative surveillance context, then must choose the right tools, indicators, thresholds, and analysis methods.
Read the book full manuscript again interact with the leaderboard.
As sequence numbers increase, genomic observations are increasingly limited by analysis. Public health workflows depend on implicit and sometimes implicit choices about sequence indexes, databases, filters, normalizations, and limits.

AI agents are promising because they can inspect files, use tools, and iterate through workflows automatically. But surveillance is a challenging problem. Agents must integrate complex scientific decisions and analyzes appropriately from a complex biological context.
We tested the optimization of sixteen string models across 4,800 runs. Pass rates range from about 14% to 50%, with most border configurations averaging between 38% and 50%.

Rejection varies greatly by harness and provider, from zero to nearly one-third of jobs.
Performance varies more with the type of work and sequencing technology than with the type of sample or test. Most functional categories came in at between 35% and 50%, but unexplained detection dropped to 20%, with genetics next at 35%.

Long-read datasets also struggled, scoring 26% compared to 41% for short-read datasets. Sample type, nucleic-acid target, and assay type did not perform very well: clinical and isolation samples were very well treated, wastewater was very poor, DNA and RNA were only modestly different, and the gun and target samples were almost identical.
Agents often acquire rational tools, but have difficulty with scientific judgment. The failure came from the surrounding choices How persuasion of those tools in context, eg choosing the wrong reference, threshold, normalization method, or final definition of a biological signal.
The most difficult tasks were open-ended judgment calls where the agent had to decide what was important without being told which target to look for. Anomalous detection requires determining whether the weak signal was real or background. Genetic engineering simulations required determining whether the sequence pattern reflected intentional construction rather than native or homologous biology.
We’re building a future where agents analyze surveillance data as it comes in, fast enough to shape an outbreak response while it’s still relevant. Today’s agents may not be reliable enough to do so, but by measuring their capabilities, we are getting closer to this future.
Read the manuscript for more details: latch.bio/biosecbench-surveillance.
We are constantly updating our stand family with new models: benchmarks.bio.
This was a collaborative effort Aclidautomation platform for biosecurity and biosafety. Thanks to our scientific collaborators: Harmon Bhasin, Kevin Flyangolts, Dianzhuo Wang, Evan Seeyave, Arjun Banerjee, Amanda Darling, Joshua Stallings, David Stern, Shawn Higdon, Claire Duvallet, Bryan Tegomoh.



