Ai Health

An Agentic Benchmark for Genetic Variation Discovery and Interpretation

Analysis of data from gene sequencing and related experiments is a critical bottleneck in clinical and research settings. AI agents hold promise in automating and scaling such analyses.

Introducing VariantBench, a 118-test verifiable benchmark that tests agents on their ability to detect and interpret variants. GPT-5.6 Sol / Codex and Claude Opus 4.8 Max / Pi led with 42.1% pass rates, while no configuration passed more than half of all attempts.

Read the book manuscript. Look Results.

VarianBench contains functions derived from independently reproduced claims in peer-reviewed studies with publicly available data. Each function provides relevant files and scientific content without specifying the method, and then marks the agent’s systematic response using deterministic graders. Tests were reviewed for scientific relevance, robustness across logical workflows, redundancy, and resistance to assumptions or data contamination.

The benchmark includes differential discovery and quality control, personal and clinical genomics, and statistical and population genetics in 14 subfields.

We tested the configuration of 26 harness models across 9,204 trajectories. Performance varies with analytical background and sequencing technology, long-term studies and genotype quality control tasks that appear to be more difficult.

Agents often complete standard pipeline measures but fail when results depend on case-specific testing or scientific judgment. They struggle with repetitive circuits, structural variation, and integration of many types of biological evidence, even when their tracks show recognition of relevant concepts.

With permission from Sid Sijbrandij’s care management team, we tested the agents in a real-world use case where differential detection and interpretation of tumor sequencing data led to formation of a personal neoantigen vaccine.

Four experiments reproduced the successive stages of the candidate selection pipeline, in which agents were required to restore the genes used for neoantigen formation. Previous stages were memory crunched and could easily be mismanaged or run out of wall clock limitations. In all the later stages, GPT-5.6 Sol / Pi and Claude Opus 4.8 / Pi also produced the characteristics of the correct work flow but failed to recover several neoantigens that have advanced to the validation of tests and clinical use. Their leads used sound but overly restrictive filters, showing how scientific and clinical judgment still play a large role in returning clinically relevant candidates.

VariantBench shows that current agents can do a useful job in all genetic variant analysis but struggle when success depends on data-specific scientific judgment. The distribution of results across the newly released models, however, shows that the models are quickly becoming more reliable tools.

Future agents will be critical infrastructure for differential discovery and interpretation across clinical, statistical, and human genomics. A systematic assessment of current agent failure modes will be essential to get us there.

Read the book manuscript. Watch it open benchmarks[dot]history.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button