Tool Call Bench

Running Agent Evals in GitHub Actions Without Live API Credentials

Stateful simulators replace live API calls to make agent evals reliable and deterministic in CI.

Contributing Writer · · 10 min read
Cover illustration for “Running Agent Evals in GitHub Actions Without Live API Credentials”
CI Eval Pipelines · October 7, 2026 · 10 min read · 2,260 words

Teams that treat it as a secrets-management exercise, keeping keys out of YAML and rotating them on schedule, still end up calling live third-party APIs from inside CI, and that call is what makes the eval flaky, rate-limited, and non-deterministic. Secrets hygiene matters for security. It does nothing to fix the behavior of the pipeline running on top of it.

Agent evals differ from a typical unit test in a way that makes this worse than it first appears. A unit test checks one function against one expected output. An agent eval exercises a multi-step tool-calling sequence, where the output of each call becomes the input context for the next one, so any randomness or inconsistency introduced anywhere in that chain compounds across the whole trajectory. A single flaky response three calls into a ten-call sequence doesn't just fail that call. It changes what the agent sees, which changes what the agent does next, which changes whether the entire run passes or fails for reasons that have nothing to do with the agent's actual competence.

Live APIs bring three failure modes into CI, and no amount of credential rotation can touch them. Every open pull request that triggers an eval job hits the same third-party quota at once, so rate limits exhaust across concurrent runs. Costs multiply, because every eval run means another round of model invocations and API calls billed in real time. And because result variance is inherent to live services under load, a fixed pass/fail threshold means almost nothing: the same agent code can pass on one run and fail on the next for reasons entirely external to the code under test.

The fix is architectural. Credential-free evals come from replacing the live API surface itself with a locally hosted stand-in that behaves like the real thing across a full sequence of calls, not from better handling of the credentials that used to point at that surface. That distinction sets the direction for everything that follows: the question isn't how to keep a key safe, but how to remove the live dependency the key was protecting access to.

How stateless mocks break under multi-turn agent tool calls

The first instinct, once a team accepts that live APIs don't belong in CI, is to reach for mocking. A stateless mock returns the same hardcoded response regardless of what happened in any prior call, and that's precisely the property that makes it unfit for agent workloads: sequential tool calls depend on cumulative state, and a stateless mock has no memory to accumulate it.

Stripe's official stripe-mock is a useful concrete case. It accepts the same request shapes as the real Stripe API and rejects malformed parameters, and it does reflect some valid input parameters back into its responses. But its responses are drawn largely from fixed fixtures, and it makes no attempt to reproduce how the real Stripe API actually behaves over a sequence of operations. An agent that creates a customer, attaches a payment method to that customer, and then queries the customer back gets a fabricated response with no relationship to anything it just did, because the mock never recorded the first two calls.

Webhook-heavy integrations expose the same limitation more severely. About half the logic in a real Stripe integration sits in webhook handlers, the code that fires when a resource actually mutates. A stateless mock never fires those events, since nothing in it ever truly mutates, so entire branches of the agent's behavior, everything gated behind a webhook, go completely untested.

The dangerous part of this failure is how quietly it happens. The eval reports a pass because the mock returned a 200 status code, but the agent's actual behavior was never exercised: the mock absorbed whatever the agent sent and handed back a canned response no matter what that input was. So if these are the conditions, a passing eval tells a team nothing about whether the agent did the right thing. It only tells them the HTTP layer didn't throw an error. That gap between "returned 200" and "behaved correctly" is the same shape of failure that appears again later in fault injection testing, where the most damaging faults are the ones that look like success.

Stateless mocks still have a legitimate job: validating that a request has the right shape and that malformed parameters get rejected. They stop being useful the moment the eval needs to judge an agent's behavior across more than one sequential call, which for most agent architectures is most of the time.

What a stateful simulator gives an agent eval

A stateful simulator keeps a memory of what happened. Create a customer, list customers, then delete that customer, and each of those three calls updates a shared internal state the same way the real API's database would, so the second call sees the effect of the first and the third sees the effect of both.

The agent calls a tool, the simulator updates its internal state to reflect that call, and when the agent calls a second tool that reads from that state, it gets back something that depends on the first call. The eval result now reflects genuine sequential API behavior instead of a string of disconnected canned responses, and no stateless mock can give you that, no matter how many fixtures it ships with.

Pre-seeded, realistic data closes a second gap. If the sandbox starts empty, every eval run has to build up world state from nothing before the actual test logic can even begin, and that setup step is slow, brittle, and itself a fresh source of non-determinism. A simulator that ships already populated with realistic records removes that burden before the eval even starts, so the harness spends its time testing the agent rather than constructing a fixture to test it against.

None of this matters if the simulator's behavior doesn't match the real API it stands in for. A simulator that gets continuously checked against the live API before every release gives an eval the same behavioral guarantees as calling the real service, without the eval ever touching it, which is the dividing line between a simulator worth trusting and a hand-rolled fake that happens to run locally. If you skip that continuous check, the simulator drifts silently: the eval still reports green while the simulator's behavior quietly diverges from what production actually does, so it manufactures exactly the false confidence the eval existed to prevent.

The GitHub Actions workflow: wiring a credential-free eval job end to end

With the architecture settled, the workflow itself is straightforward to wire up.

The eval harness points at that simulator through the same environment variable that points at the live API in production, so you might set a base URL to localhost:<port> in CI and to the production API's URL elsewhere. Because the variable name doesn't change between environments, the harness code doesn't either: swapping environments is a configuration change, not a code change.

This is where the credential-free claim becomes concrete. You don't need a secrets: block anywhere for the third-party API itself, because the job never calls it. The only secrets that can legitimately appear are those for an LLM judge scoring the agent's output, and even those are avoidable: a deterministic trajectory-diff evaluator, one that checks the sequence and shape of tool calls against an expected pattern rather than asking a model to judge quality, removes the need for a judge credential on sequencing checks.

Trigger strategy keeps cost and feedback speed in balance. Running the full eval suite, including any LLM-judged checks, on every push to main and on a nightly schedule gives thorough coverage without gating every pull request behind it. Pull requests instead run only the deterministic trajectory-diff checks, which keeps PR feedback fast and avoids multiplying LLM judge invocations across every commit on every open branch.

The eval harness layer has multiple viable options, so pick one based on what you're testing rather than defaulting to the first name you encounter. AgentCheck, open source under Apache-2.0 and released in 2026, performs behavioral testing for tool-using agents by simulating declared tool execution while evaluating tool calls, failures, retries, confirmations, destructive actions, and the sequencing of actions; it integrates with the OpenAI Agents SDK, PydanticAI, and custom Python agents. Tools like this sit at the evaluation layer, consuming what the simulator produces.

Whatever harness is chosen, the job's exit code needs to reflect something real: whether the agent's trajectory score or behavioral diff fell below a configured threshold, not merely whether the harness executed without throwing an exception. Restricting full evals to push events on main and nightly runs and reserving pull requests for the fast deterministic checks keeps the cost of thorough evaluation from scaling with every commit.

Teams that want to avoid even running the agent live in CI have another option: capture representative OpenTelemetry spans from a staging environment, store them, and evaluate those stored spans in the pipeline. So you remove the simulator dependency entirely for pull-request checks this way, but you still exercise the evaluation logic itself. But it complements the simulator approach rather than replacing it, because stored traces evaluate behavior the agent already showed in staging, not the code sitting in the current pull request.

The offline case falls out of the architecture for free. Because the simulator runs locally inside the Actions runner itself, the eval job has no outbound network dependency on the third-party API at all, so it satisfies air-gapped CI requirements without any special configuration layered on top.

Fault injection against the simulator: testing agent resilience without touching live services

Once correctness evals run cleanly against a stateful simulator, the same simulator becomes the right place to test how the agent handles things going wrong. Faults can be triggered deterministically and repeated exactly the same way on every run, with no live service anywhere near the test.

Three fault classes matter most for agent evals. Rate-limit responses, a 429 status with retry headers, check whether the agent's retry logic actually backs off. Latency spikes and timeouts check whether the agent honestly reports a timeout instead of mistaking a slow response for a successful one. The third class, partial or empty success responses, is the most dangerous of the three: an exception from a timeout or an error gets caught by retry logic that was built to catch exactly that, but a response that returns successfully while carrying empty or truncated data looks like a normal success, so the agent proceeds to act on bad state without anything in its retry logic ever firing.

Research published under the name AgentChaos put a number on how much this matters. When you inject faults into LLM API responses, without changing a single line of agent source code, pass@1 falls by up to 49.66 points across five agent systems and four backbone models. The fault class responsible for the steepest drops wasn't the one that raised a visible exception. It was the one that returned a response looking superficially successful while carrying corrupted or incomplete data, because nothing in the agent's retry logic was built to notice a response that returned normally but lied.

The agent-chaos open-source toolkit puts this into practice as middleware sitting between the agent and its tools, injecting configured fault scenarios, latency spikes, timeouts, 429s with retry headers, leaving the agent's source code untouched. Paired with the simulator layer described earlier, the division of labor is clean: the simulator owns state, the fault-injection middleware owns the failure scenarios, and the agent experiences both exactly as it would in production, without either one requiring a live credential.

Simulator drift: how to keep your eval environment honest as the real API evolves

The real API keeps changing underneath the simulator it's built around. A field gets renamed, a status code changes, a new header becomes required, and the simulator can fall behind quietly, continuing to pass evals that no longer mean what they used to.

Drift appears in two distinct forms, each needing a different kind of check. Mock-to-real drift is when the simulator's behavior diverges from what the live API actually does today. Catching it means running the simulator's own test suite against the live API on a schedule and publishing the results as probe-level fidelity scores, numbers that state how many checks ran and how many passed instead of a general claim that the simulator is "accurate.

So if you want to answer mock-to-real drift, you need continuous verification built into the simulator's own release pipeline. If you test a simulator against the live API as part of its release process, it can't diverge without that divergence surfacing as a probe failure before the new version ships, so drift turns from a silent risk into a visible, blocking one. Published fidelity scores with an attached probe count give users something they can audit for themselves, not a marketing line they have to take on faith.

Spec-to-implementation drift is the second form: the real API's own documented spec falls out of sync with what the live provider actually does. The right cadence here is a small set of contract tests run against the live provider on a schedule, with credentials for that scheduled run stored in the CI secret store. This is the one place in the whole architecture where live credentials belong, scoped narrowly to a scheduled contract check. If you catch spec drift this way, you learn about real API changes before your simulator-based evals start producing results that look fine but no longer match what the provider does.

Sources

  1. Simulating Complex Multi-Turn Tool Calling Interactions in Stateless Execution Environments
  2. Stateful Inference for Low-Latency Multi-Agent Tool Calling
  3. EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation
  4. AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation
  5. Model-Free Assessment of Simulator Fidelity via Quantile Curves
  6. Benchmarking API Drift in LLM-Generated Quantum Code Across Successive SDK Versions