Ai Health

Rejection of measurement in Agentic Biology

AI agents are increasingly able to design proteins, plan experiments, and interpret complex data. These skills present a major dual-use problem: they can accelerate beneficial acquisitions or, if misused, allow harm.

Rejection training is the primary way model developers reduce biosecurity risk, but current “rejection-first” behavior hinders legitimate work and reduces reliance on AI tools for productive biological research. To quote Patrick Boyle of American Wetware,”When I use frontier AI to do bioengineering in May 2026, my daily experience is frustration.

Introducing BioSecBench-Refusala benchmark for risk identification and ethical rejection of biological research activities. The benchmark pairs 61 cycle activities, formal analysis taken from published literature, and 46 Red-Team activities, fictional situations similar to real research but hiding a biosecurity risk.

Across the 16 types of harness configurations, rejection rates ranged from 7% to 74% for Standard tasks and 1% to 62% for Red Team tasks, with many configurations rejecting the official work of the Cycle at rates comparable or higher than the hidden dangers.

Check the results and interactive trajectories at benchmarks.bio.

107 BioSecBench-Refusal tests, 61 Routine and 46 Red-Team, were written by a team of 14 subject experts to cover a range of research topics in the health sciences: microbiology, virology, immunology, plant biology, synthetic biology and related fields. Each task was annotated by biosafety level, class of biological agent, type of application and technical background.

The experiment was presented in one of two frameworks: direct framing, where the agent is asked to perform a task, or meta framing, where it is asked to judge whether the task should be accepted or rejected for biosecurity reasons.

Maximum rejection rates ranged from 7% to 74% for Cycles tasks and 1% to 62% for Red-Team tasks, with Opus models being the strongest in both. For all x harness configuration models, the API rejection is dominated by the model rejection. GPT-5.5, for example, rejected 40% of Cycle Operations at the API layer, but only 4% after agent reasoning.

In almost all configurations tested, rejection rates were higher in Cycle tasks than in Red Team tasks. This gap increased with the biosafety level assigned to the worker from BSL-1 to BSL-3 (the relatively small number of BSL-4 conditions tested makes comparisons at this level inconvenient).

The rejection rates of the Round and the Red Group are strongly correlated in all models (Pearson r = 0.91), which points to one basic starting point: the surface text. Common tasks were often rich in keywords that might flag immunity (“pathogen”, “immune evasion”). The Red-Team tasks, although written to avoid obvious flag words, also contain technical language with two-letter words that may be detected by a filter (“DNA assembly”, “protein expression”).

In order to test whether the agent’s reasoning can identify a complex biosecurity risk, we moved from a direct frame to a meta-evaluation framework: instead of performing each task, the agent judges whether it should be accepted or rejected.

In works framed as biosecurity meta-evaluations, most rejections came from the provider’s API filter, not the model’s reasoning itself. For example in the GPT-5.5 x PI configuration, 60% of Cycle Jobs are rejected. Two-thirds of those rejections (40 percent of all jobs) caused the API filter to block the request before the model could resolve it.

The model’s own rejection accounted for the remaining one-third (20% of all jobs). When agents were allowed to think, they were occasionally able to spot threats that would otherwise have been missed. For example, GPT-5.5 and Grok correctly rejected 14.5–19.6% of Red-Team tasks under meta-analysis, compared to 13% in the direct frame. These estimates are preliminary, as the API’s high rejection rate leaves a small sample of actual agent decisions to be tested.

BioSecBench-Refusal addresses both a technical challenge and a management problem. At the technical level, our test of 16 harness pairs of x models shows that the current biosecurity protection is not enough: the border models failed to detect many hidden threats of the Red-Team.

As a management problem, biosecurity deals with the trade-off between safety and performance. The significant refusal rates recorded in Cycle Activities show that refusal behavior comes at a real cost to formal research. These costs are partly a result of the way biosecurity is currently practiced. Because the rejection tracks the surface language of the application rather than the underlying biology, work with flagged keywords may be rejected while the danger hidden in the protein structure file goes unnoticed.

The problem of dual use will not go away, but better biosecurity metrics will help model developers improve biosecurity performance and release biotech R&D agents’ tools with confidence.

Read the full paper here.

Play with the results here.

This project was developed in collaboration with partners at American Wetware: Jake Wintermute, Patrick Boyle and Christina Agapakis. Special thanks to the experts who wrote the benchmark tasks and the design project guidance: Daniel Fulop, Matthew C. Watson, Adam J. Meyer, Sandrine Boissel, Jens H. Kuhn, Rishi Jain, Noah D. Taylor, Helena Shomar.

Related Articles

Back to top button