Tool Call Bench

Tool-Use Benchmark Design for API-Calling Agents

Benchmarks for API-calling agents need stateful environments, not mock stubs.

Staff Writer · · 11 min read
Cover illustration for “Tool-Use Benchmark Design for API-Calling Agents”
Agent Eval Benchmarks · October 7, 2026 · 11 min read · 2,436 words

Tool-use benchmarks for API-calling agents have not kept pace with what these agents are now asked to do, and that gap is the subject of this piece. The field once asked only whether a model could pick and format one correct function call, but now you need to know if it can carry out a long multi-step job across several services, where each step depends on what happened before it. The Berkeley Function Calling Leaderboard, known as BFCL, became the standard the community reaches for, and it has gone through four major versions: single-turn call evaluation in February 2024, then live data contributed by enterprises and open-source projects, then multi-turn interactions, then a broader agentic evaluation in July 2025. Each version exists because agents were failing in ways the one before it had missed. A structural problem runs through all four versions regardless: orchestrating multiple tools calls for picking the right subset of tools dynamically, tracking dependencies between them, scheduling calls in sequence or in parallel, recovering when something fails, and re-planning when the situation changes. Single-turn matching of a call's syntactic structure cannot see any of that. An agent can score well on BFCL and still fail badly once it hits a live system, because tasks tend to get selected for how easy they are to simulate and grade rather than for how well they represent what users actually need, which leaves multi-service dependencies and shifting environment states underrepresented exactly where failures are most likely.

AST Matching and Stateless Mocks Underspecify Real API Behavior

Diagram: BFCL's Four Versions and What Each One Missed. Visualizes: Show the four successive versions of the Berkeley Function Calling Leaderboard (BFCL) as a timeline with the failure mode each version was built to address.

Most tool-calling evaluation still runs on abstract syntax tree (AST) matching, a method that checks whether a generated function call is shaped correctly, but not whether it does the right thing. A model can write a function call with the correct name, the correct argument types, and the correct structure, and still call the wrong endpoint, pass units in the wrong scale, or supply an argument that is valid in form but wrong in meaning. BFCL's executable category runs the generated call against a real or simulated API and checks the returned value against a known answer, going beyond the call's shape. That is progress, but stateless mocks still leave a hole underneath it. A stateless mock is an endpoint-level stub standing in for a real system, not a system with memory, so it can only tell you whether the agent picked a plausible call against a simplified interface. It cannot tell you whether the agent can finish a job that depends on changes that persist.

AFT-Bench names the distinction between callability and operability. A tool is callable when an agent can put together a syntactically valid request. A tool is operable when its interface exposes enough state and enough meaning for the agent to keep acting safely once something goes wrong or the situation turns uncertain. Those are different properties, and a stateless stub cannot tell them apart, because a stub has no memory of what came before and nothing at stake in what happens next. Consider a case where an external effect actually goes through but its confirmation gets lost on the way back: the agent sees a timeout, and from where it sits, a committed change and an uncommitted one look identical. No amount of additional reasoning inside the model can fix that, because the missing information sits at the boundary between the agent and the tool, not inside the model's own logic. AFT-Bench demonstrated this directly by holding the task, the backend, the starting state, the injected failure, the agent, and the underlying language model all fixed, and varying only the interface the agent was given to work with, which confirmed that reliable tool use depends on both the model's policy and the operational semantics the interface exposes. Mock design compounds the problem in a smaller but common way: rate limiting gets left out of mocks routinely, so agents never meet a throttling error in testing and then meet one for the first time in production.

Why multi-turn trajectories require stateful environments, not stub chains

BFCL's third version moved to state-based verification for multi-turn tasks, and it had to make that move because convenience did not drive it. Instead of grading the structure of each function call on its own, the evaluation checks the actual state of the backend system, whether that's a file system, a set of booking records, or a database, after the model has run its whole sequence of calls. Once tool use stretches into long chains involving writes that change the world, keeping state consistent and managing race conditions under parallel execution become the central sources of instability, and you cannot get a chain of independent stub responses to reproduce either problem, because each stub answers without any memory of the one before it.

LiveClawBench puts this idea into practice at scale. It runs executable tasks across dozens of domains, and each task is built as a reproducible full-stack mock application, written in Bun and TypeScript, that keeps persistent backend state, logs activity on the service side, exposes a browser-facing interface, and runs inside a container for reproducibility. LiveClawBench keeps the execution semantics that actually determine success or failure, rather than swapping real services for simplified ones. Agent-Diff approaches the same problem from a different angle: it has agents work against the real API interfaces of enterprise productivity tools, among them Slack, Box, Linear, and Google Calendar, inside containers that can be spun up identically every time, which removes the instability that comes from testing against a live service while still keeping real API behavior in the loop.

Resumable invocation and durable execution state each raised recovery rates by 100 percentage points under their matched failure conditions in AFT-Bench's numbers, a result that is only possible against a backend that actually remembers whether an operation went through. FAMA's results on τ-bench, τ-trait, and ACEBench make the same point from the failure side. Even strong LLM-based agents struggle to finish realistic, multi-turn customer-service tasks that call for sustained reasoning and structured tool use, and the failure modes FAMA documents, domain policy violations, wrong retrieval from complicated tool outputs, misreading context, hallucination, incomplete fulfillment, and stopping early, only become visible once the environment is actually tracking state across turns. None of this works, though, unless the environment can be reset to an identical starting point run after run. If a stateful environment can't be reset that way, you can't compare scores from different runs, so containerized instantiation, not testing against live services directly, is the right architecture for a benchmark meant to be taken seriously.

Parallel Execution and Resource Scheduling

If benchmarks run tool calls one after another, they miss a failure mode that only occurs once calls run at the same time. An agent can plan a set of parallel tool calls that is logically sound, dispatch every one of them at once without any sense of the physical infrastructure behind them, and still cause an overflow of shared resources. Independent tasks sharing finite infrastructure produce sharp spikes in load, long queuing delays, and outages, and in production these look exactly like API errors even though a serial benchmark would never have exposed them. PeakBench splits the evaluation into two parts, one for logical planning and one for physical scheduling, each with its own metrics, and the results show that sound logical planning does not reliably carry over into execution that is safe or efficient once real resource limits are in play.

The programmatic tool calling study on BFCL v4 reveals a related limit. Native JSON tool calling drops calls entirely once the number of parallel calls crosses a model-specific threshold, somewhere between 70 and 72 for Claude Sonnet 5, but programmatic tool calling keeps full accuracy at a fan-out of 100. That ceiling in JSON calling only occurs under parallel fan-out conditions, so if a benchmark never tests fan-out, it will never catch it. A benchmark worth trusting needs workflows with dependency structure that's actually annotated and resource profiles that are actually measured, so that a failure caused by bad dependency planning can be told apart from one caused by scheduling that ignores resource limits. Without that separation, a single score mixes two unrelated kinds of error together, and you learn very little about which one happened.

Fault Injection and Operational Reality

The Failing Tools benchmark puts a number on this gap. Built on stateful APIs across multiple domains, it injects the kinds of failures real systems produce: denial of availability, stale data, operations that silently do nothing, corrupted state, mismatched schemas, cases where the right action is ambiguous, and cascades where several of these compound at once. Under its base scoring, which gates credit on actual recovery, no model clears even a low threshold across these scenarios. These are what integrating with real APIs actually involves: rate limiting gets left out of mock design often enough that agents never see a 429 or a 503 in testing, then meet both constantly once deployed.

If you run fault injection inside an environment that returns real error semantics instead of a stub's success response, you see what a benchmark built around the Stripe API shows. Some agents, faced with nonexistent Stripe data, got back a client error, read that error as proof the endpoint was working, and marked the task complete, rather than reading the error as what it was: evidence that the argument they'd supplied was wrong. A stub that always returns success would never reveal that failure mode. AFT-Bench's numbers show what the right interface design buys back: effect-aware interfaces cut duplicate effects by 56.9 percentage points, cut unsafe commits by a wide margin, and verification mechanisms reduce false claims that a task finished when it hadn't, and these gains appear once the benchmark injects the conditions, lost responses, state that committed without being acknowledged, that make those interface properties matter. FAMA's analysis of failure trajectories on τ-bench shows why this compounds over a long task: the most common errors, violating domain policy, pulling the wrong data from a complicated tool output, leaving a task incomplete, build on each other because an error early in a sequence corrupts the state the agent is reasoning from in every turn that follows. You only see that kind of cascade once fault injection runs against a backend that actually holds state, which is what the earlier sections have been building toward.

Why benchmark quality itself is now a contested problem

Diagram: BFCL Defect Rates by Task Category. Visualizes: Show defect rates across BFCL task categories as found by Epoch AI's review of a random 50-task sample: 48% of tasks overall contained defects affecting scoring accuracy; non-live AST tasks…

BFCL grades a model's output by checking it against a fixed answer key of acceptable function calls, so you have to keep that answer key accurate for the method to work. A review by Epoch AI looked at a random sample of BFCL tasks and found defects that could affect scoring accuracy in 24 of 50 tasks, 48 percent of the sample, with defect rates by category running from around a fifth for non-live AST tasks to roughly seven in ten for memory tasks. These numbers trace back to the same fidelity problems the rest of this piece has been describing, which produce them directly: ground truth that has gone stale, and scoring logic tied to live external state that can drift out from under it. The benchmark meant to catch unreliable API behavior suffers from the same kind of drift that makes testing against live APIs unreliable.

The high defect rate in the memory category is not a coincidence. Memory tasks require the benchmark to track what the agent has done across multiple turns, a stateful consistency property stateless mocks cannot provide. If a benchmark cannot reliably track its own state, it has no basis for grading agents on whether they tracked theirs. The programmatic tool calling study on BFCL v4 found that whether a model handles fan-out well tracks its generation rather than its family, so newer models across different lineages behave more alike than older and newer models within the same lineage do. That kind of finding would be masked rather than revealed if the benchmark's defects were concentrated in exactly the categories where generational differences are most visible. A benchmark score is only as trustworthy as the process used to keep that benchmark accurate. If you do not check a benchmark continuously against how the real API actually behaves, it will drift quietly, and developers relying on it will get false confidence about how reliable their agents actually are, the same failure stateless mocks produce in integration testing, by a different route.

What a benchmark environment built for production-predictive scores requires

The three problems this piece has traced, fidelity, statefulness, and fault injection, point to three design requirements that a serious benchmark cannot skip. The tool environment has to behave like the real API it stands in for. It has to be stateful, and it has to reset to an identical starting point across runs. And it has to support fault injection as a core part of evaluation, not an extra feature bolted on afterward.

Fidelity to the real API is not a one-time property to design in and forget. You have to check the benchmark's environment against the live API continuously, before each release, because ground truth drifts otherwise, and then the score ends up measuring how stale the benchmark has gotten rather than how capable the agent is, which is precisely the defect Epoch AI's review found running through BFCL. Statefulness means the environment has to carry the effects of one call into the next: add a record, list it, delete it, and the environment has to reflect all three actions as a connected sequence, not three isolated events. That is what lets an evaluation catch the cascading errors FAMA documented and the recovery failures AFT-Bench measured, so each turn can no longer hide the consequences of the one before it. An agent that performs well against an empty sandbox can still fail once it meets the kind of messy, realistic data that exposes retrieval mistakes, schema edge cases, and conflicting policies, so a sandbox built empty is a sandbox built to miss its own agent's weaknesses. Fault injection, finally, has to be available on demand: latency spikes, 429s, 503s, lost responses, and operations that silently fail to do anything all need to be things a benchmark can switch on at will, so that what gets measured is how an agent recovers, not just how it performs when nothing goes wrong. If you build a benchmark on all three of these foundations, its score still means something once the agent leaves the test environment and works against systems that were never built to make its job easy.

Sources

  1. The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
  2. FAMA: Failure-Aware Meta-Agentic Framework for Open-Source LLMs in Interactive Tool Use Environments
  3. The Bitter Lesson of Tool Calling
  4. PeakBench: Benchmarking Resource-AwareTool Invocation in LLM Agents
  5. Callability Is Not Operability: Controlled Interface Interventions for LLM Agents

More in Agent Eval Benchmarks