Tool Call Bench

Scoring Multi-Step Agent Tasks With External Side Effects

Agents that change real systems need scoring based on actual state, not just final answers.

Senior Technical Editor · · 10 min read
Cover illustration for “Scoring Multi-Step Agent Tasks With External Side Effects”
Agent Eval Benchmarks · October 8, 2026 · 10 min read · 2,305 words

A benchmark score of green tells you almost nothing about whether an agent is safe to put into production once that agent can act on the world. A multi-step agent does not just generate text. Evaluation here operates on at least two layers: final-answer scoring, which checks only the last output, and trajectory scoring, which checks whether the agent called the right tools in the right order. Neither layer, taken alone, catches what matters most for an agent that touches real systems.

The sharpest illustration of why both layers fall short is the tool misuse cascade. A single wrong tool call at step two of a six-step pipeline corrupts every step that follows it. The final answer can be entirely correct in form and still sit on top of a trajectory that has already caused harm no re-run can undo.

External side effects complicate scoring

External side effects are changes an agent makes in systems outside its own reasoning context, changes that outlive the agent's run and can be observed, accumulated, or built upon by whatever acts next. A Stripe charge that succeeds, a GitHub issue that gets closed, a Slack message that gets sent, a Kalshi order that gets filled: none of these are text outputs. They are world states, persisting in systems the agent does not control and cannot take back simply by generating a different final answer.

A side effect cannot be judged wrong or right in isolation the way a generated sentence can. The dependency runs forward through the whole trajectory, and a check confined to any single point in that chain sees only a fragment of what happened.

Research on multi-agent evaluation has started to map how wide this blind spot runs. The M-SAEA framework identifies four distinct layers where risk accumulates in agent systems: model, workflow, interaction, and system. None of these show up in a final-answer check, because none of them live in the final answer. They live in the sequence of external changes the agent made getting there. This is also the point where irreversibility starts to bite: catching a wrong tool call early costs little, but catching it after its downstream effects have compounded can be expensive, and in some cases nothing can undo it.

Step-level failures hiding inside trajectories that look correct at the surface

The step-level failures that matter most for side-effect scoring are not the ones that throw an error. An agent that fails loudly is, in a narrow sense, well-behaved: something flags the problem, a human or a retry loop gets a chance to intervene. The dangerous failures are the ones that return a normal-looking response while quietly doing the wrong thing, because nothing in a conventional trajectory check distinguishes that response from a correct one.

A 2025 arXiv analysis of multi-agent failures found two patterns dominating the landscape. Read in isolation, the final text answer looks correct, but the side effect underneath it is wrong, and no check confined to the text would ever surface that.

Planning defects compound the problem before execution even begins. Run the arithmetic on a six-step pipeline: a tool misuse at step two propagates through steps three, four, five, and six. The final-output scorer sees a confident, plausible answer at the end of that chain. The actual trajectory shows an error compounding at every step after the second, and no single step along the way raised a flag that would have caught it.

The downstream state of side effects as the primary scoring surface

Scoring a multi-step agent task that produces external side effects requires inspecting the state those side effects actually produced, treating that state not as a supplementary check performed after the fact but as the primary evidence of what happened. The final text answer is one agent's claim about its own performance. The state of the systems it touched is the record of what it actually did.

Defining task completion this way means verifying the resulting database or system state rather than trusting the final text alone. Tau-bench exists for precisely this reason: checking database state, alongside whatever text output a task requires, catches agents that report a task as complete without having completed it. The scoring surface under this approach is the post-task world. Did the Slack thread contain the right messages, arriving in the right order? These questions share a property that distinguishes them sharply from grading generated text: state-as-evidence is deterministic in a way text-as-evidence is not. A record either carries the right fields or it does not. These are binary checks, resolvable without an LLM judge rendering a subjective verdict on whether an answer sounds right. M-SAEA's system-layer probes are built to stress exactly this boundary, the point where an agent's actions meet real-world services, because that boundary is where side effects land and accumulate. Auditing inter-agent messages and tool invocations at that boundary is what converts an evaluation from outcome-only into evidence-based.

One objection deserves a direct answer: could the final state be correct even when the path that produced it was wrong? It could, and that possibility shows why state alone, checked only at the end, is not the whole answer. In systems where actions cannot be cleanly undone, a correct final state reached by a wrong path can mean that the wrong path already did damage along the way. The end state looks fine. The customer who saw two charges land in their account before the refund arrived experienced something the final-state check will never show a scorer that looks only at where things ended up.

What an environment needs to support state-as-scoring-surface

Treating downstream state as the primary scoring surface only works if the test environment can actually capture, hold, and reset that state reliably. That requires three properties that live API calls and stateless mocks both fail to provide on their own: state that persists and accumulates correctly across calls, full inspectability of that state between steps, and deterministic reset between separate runs.

Persistent state across calls means a record created in step 1 of a trajectory has to genuinely exist when step 3 goes looking for it. A stateless stub, lacking any memory of the cancellation, will let a broken implementation pass through every single time, since it never had the information needed to notice the implementation was wrong.

Full inspectability between steps means a scorer can read the environment's state at any point in a trajectory, not only once the run has finished. This matters because a wrong intermediate side effect needs to be caught at the step that produced it, not inferred later from downstream symptoms. Checking tool-call accuracy, for instance, requires validating both tool selection and argument correctness against a schema at each call, and that validation is impossible without direct access to the environment's own record of what was called and what state resulted from it.

Deterministic reset between runs is the third requirement, and it underwrites reproducibility itself. Pass^k scoring, the approach tau-bench uses to test an agent across k independent trials on the same task, only means something if the environment resets to an identical pre-seeded state before every trial. They accumulate state across test runs, and an artifact left behind by trial one contaminates the starting conditions for trial seven. That accumulation, combined with rate limits, the need for live credentials, and side effects that cannot be undone once triggered, makes a live API unsuitable as a scoring environment even when it is the very system under test. Stateless mocks fail differently: lacking any memory of prior calls, they cannot reflect the cumulative reality of sequential API interactions, so a side effect that depends on an earlier state transition simply does not exist as far as the mock is concerned.

Verified stateful simulators satisfy those requirements while unverified ones fail

A stateful simulator, one that accumulates effects across calls and resets deterministically between runs, satisfies all three requirements that live APIs and stateless mocks fail to meet. A cancellation changes a field, and every subsequent read of that record reflects the change. Reset becomes a first-class operation built into the environment itself: each evaluation run starts from a known, pre-seeded state, so trial 1 and trial 7 begin from identical conditions, which is what makes pass^k scoring meaningful. Inspectability comes from reading the simulator's internal state directly, rather than trusting only what comes back through API responses, so a scorer can confirm the right record was created with the right fields without relying on the agent's own account of what it did.

None of that holds if the simulator has quietly drifted away from the real API it is meant to represent. An agent's tests then pass against a contract that no longer describes production. A passing score becomes a score against a fiction.

This is not a hypothetical risk. Two days earlier, on October 3, 2026, a nightly scheduled spec drift detection run on Kalshi's SDK found that the contract suite no longer matched the upstream OpenAPI and AsyncAPI specifications, a mismatch that nightly contract test was built to surface. Both cases show drift as an ongoing condition of live APIs, not a one-time migration event that gets handled once and forgotten.

The mitigation is contract testing: checking the simulator against the real API before every release, hard-failing on drift in query parameters, request bodies, or WebSocket payloads.

Fault injection as a scoring dimension for side-effect resilience

A scoring framework built only around whether the right side effect occurred still misses half the picture. An agent's response to an environment failure matters just as much, because the most damaging real-world failures are not the ones that raise an error. The dangerous ones return a normal HTTP 200 while quietly producing the wrong side effect.

Silent failures are the hardest class to catch precisely because they generate no error signal for the agent to react to and no failure flag for a conventional scorer to notice. A tool call that returns 200 but charges the wrong amount, or writes a record missing a required field, looks identical to a successful call from the outside. Nothing about the response invites the agent to question what just happened, and nothing about a standard trajectory check would catch that the side effect produced was not the one intended.

Fault injection turns this into a scoring surface, measuring a risk that otherwise goes unmeasured. Running the same task with injected latency, injected 429 responses, injected 503s, and injected silent wrong-value responses lets a scorer check whether the agent notices the anomaly, retries in a way that actually resolves it, or halts rather than continuing forward on top of state it should no longer trust. This is formalized as one of three dimensions in ReliabilityBench, which evaluates agents on consistency under repeated runs, robustness to task perturbations that are semantically equivalent but phrased differently, and fault tolerance under injected tool and API failures. None of this is achievable against a live API, which cannot be instructed to return a 503 on step three of a given trajectory on demand. A stateful simulator with fault injection built in as a first-class feature can trigger exactly that failure, at exactly that step, every time the test runs. Scoring recovery then becomes a direct comparison: did the agent retry correctly, escalate appropriately, or abort cleanly, or did it proceed past a failure it never detected and produce a downstream side effect resting on a step that had silently broken underneath it?

A concrete scoring workflow across a multi-step agent task

Put these pieces together: a scoring workflow for a side-effect-producing task takes a specific, checkable shape. Consider an agent tasked with processing a refund request that spans a payment system, a support ticketing system, and a notification channel: a six-step trajectory touching a Stripe-like payment API, a GitHub-like issue tracker, and a Slack-like messaging system, run inside a stateful simulator that has been contract-tested against each of those real APIs before the evaluation begins.

Scoring starts before the agent takes a single action, with the simulator reset to a known, pre-seeded state: a customer record, an open support ticket, and no pending refund. As the agent executes, a scorer inspects state after each step rather than waiting for the final answer, checking whether the tool called at that step matches the one the task required and whether the arguments passed match the expected schema. When the agent issues a refund, the check is not whether the agent claims to have issued one. When the agent sends a Slack message confirming the refund, the check is whether that message exists in the right channel, in the right order relative to the ticket closure it is meant to follow.

Midway through this trajectory, fault injection adds a second layer of evidence. A separate run of the identical task can inject a 503 on the Slack call, testing whether the agent retries appropriately or reports a false success. Because the simulator resets deterministically, this same task can run across multiple independent trials, pass^k style, with each trial starting from identical conditions, turning a single pass-or-fail outcome into a distribution of outcomes under repeated runs.

The final text answer the agent produces still gets checked, since a correct side effect paired with a garbled or misleading summary is also a failure that scorers need to catch. But in this workflow, that text check is one input alongside several others, not the entire basis for the score. The state of the payment record, the ticket, and the message thread, inspected at each step and validated against a contract-tested simulator capable of injecting failure on demand, is what tells a scorer whether the agent actually did what it was asked to do, in the order it needed to happen, under conditions that resemble what production will eventually throw at it.

Sources

  1. From Tasks to Teams: A Risk-First Evaluation Framework for Multi-Agent
  2. When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation

More in Agent Eval Benchmarks