Hermetic replay

Hermetic replay

This is the seam that makes “the recorded run is the test” real: the same agent code runs live in production and deterministically offline in CI — no API keys, no cost, no flakiness.

The idea

Wrap each external call in call_tool / call_llm instead of calling it directly:

import tracely_sdk as tracely
 
def run(user_input: str) -> str:
    with tracely.agent("support-agent"):
        order = tracely.call_tool("get_order", lambda: get_order(order_id),
                                  args={"order_id": order_id})
        answer = tracely.call_llm("gpt-4o", lambda: chat(messages),
                                  input=messages, usage=(812, 96))
        return answer
  • In production (no fixtures active): call_tool / call_llm invoke your fn, record its output (and any error) on the span, and return it — a normal traced run.
  • In CI replay: tracely replay loads the regression case’s recorded fixture bundle and activates it; call_tool / call_llm then serve the recorded outputs (in order, or matched by args) and never call your fn.

A call that errored in production is reproduced on the replayed span and raised as tracely.ToolError — so your agent’s own try/except runs exactly as it would live, and the gate sees the same failure condition. Faithful error-handling replay, for free.

Recorded means recorded

fixtures(None) is “no replay requested” — everything runs live. Any other bundle, including {} (an explicitly empty recording), is recorded mode, strict by default: a call the recording cannot answer raises tracely.ReplayError instead of running your function, so a recorded execution can never quietly become a live one.

ReplayError.reasonWhen
missingnothing was recorded under that tool / model name
exhaustedevery recorded call for that name has already been served
mismatcha tool was called with args the recording doesn’t have — a different call is never consumed just because the name matches

Matching lives in one place and is deliberately simple: an entry recorded without args matches by order; otherwise both sides are canonicalised (JSON strings parsed, key order and tuple/list ignored) and compared for equality. LLM calls are looked up by model id, falling back to the next recorded LLM call of any name (auto-instrumentor span names are not model ids). An LLM call whose input differs from the recording is still served in order but stamped tracely.replay.divergence=input and reported: recorded-model replay tests your orchestration under the recorded conditions, and a changed prompt is exactly the divergence it must show, not hide. Recorded model outputs do not prove a changed prompt produces a better answer.

The block yields a ReplayReport:

with tracely.fixtures(bundle) as rep:
    run(case_input)
rep.mode      # "recorded" | "lenient" | "live"
rep.served    # ["tools:get_order", "llm:gpt-4o"]
rep.unused    # recorded calls the run never asked for — divergence evidence, not auto-fail
rep.diverged  # served despite a different input
rep.errors    # the strict misses that were raised
rep.providers # which SDK create-methods were patched in this process
rep.clean     # strict run that reproduced the recording exactly

fixtures(bundle, strict=False) keeps the pre-0.4.3 lenient behaviour (live fall-through on a miss, ordered serve on a mismatch) for callers migrating. Its report says lenient — it is never presented as a recorded run. tracely replay --lenient is the CLI form.

No rewrite needed: @observe tools and auto-instrumented LLM calls replay too

The call_tool/call_llm seam is the fully-portable path, but most agents don’t need a rewrite:

  • Tools — a function decorated @tracely.observe(as_type="tool") is transparently hermetic: under fixtures it serves the recorded output (or re-raises the recorded failure as tracely.ToolError) instead of running. Decorate your tools and the tool side replays as-is.
  • LLM calls — inside a fixtures() block Tracely class-patches the provider create-methods below, so code that calls the SDK directly (under instrument="auto", or via the drop-ins) is served the recorded completion — reconstructed into a provider-shaped response object — and never hits the network. A strict miss raises ReplayError at the call site, like call_llm.

What is actually intercepted

ProviderEntry pointSync / asyncStreamingServed shapeTested against
OpenAIchat.completions.createbothno — stream=True goes liveduck-typed ChatCompletion (choices[0].message.content / .tool_calls; no usage)real openai client
Anthropicmessages.createbothno — messages.stream goes liveduck-typed Message (text / tool_use blocks)real anthropic client
Google GenAImodels.generate_contentbothno.text, .function_calls, candidates[0].content.partsreal class when installed
Mistralchat.complete / complete_asyncbothnoOpenAI-shapedstand-in only
LiteLLMcompletion / acompletionbothnoOpenAI-shapedstand-in only
Anything else———live in replay—
⚠️

A surface outside this table — the OpenAI Responses API, Bedrock, Cohere, a framework’s own HTTP client, streaming on any provider — is live in replay; the SDK cannot see that call. The report catches it after the fact: recorded calls the run never asked for (rep.unused, also logged) mean something went to the network. The SDK does not sandbox Python or deny egress. If you need that guarantee, deny provider/tool egress in the CI executor itself, or use call_llm for a provider-agnostic seam.

Activating fixtures manually

tracely replay does this for you, but the primitive is public:

with tracely.fixtures(bundle):     # bundle = the case's recorded tool/LLM outputs
    run(case_input)                # call_tool/call_llm now serve recorded values
# outside the block, calls are live again

fixture(kind, name) peeks the next recorded output for a tool/model without consuming it.

After a fix, re-record. The recording is a recording of the bug. A fix that adds a call the bug never made (the silent-failure case: the fixed agent now calls the tool) cannot be served from it — strict replay reports INCOMPLETE — no recorded call for tools:get_weather, never a live call. Once the fixed code has run for real (CI, staging, production), re-record the case from that run: POST /api/cases/{id}/recapture with the trace id (same input required; bumps the case version). From then on the suite replays the fix hermetically and a regression (run_broken) fails on the missing tool. The same applies to a changed prompt: recorded-model replay serves the old answer back, so a quality fix is shown by re-recording from the faithful run, not by replaying the old recording.

tracely replay itself refuses to run a case whose recording cannot be loaded: it emits the case’s trace errored without running your agent, so the gate fails loudly instead of grading a live (or “nothing recorded”) run as the recorded case.

Why this matters

  • Deterministic — the agent sees the exact tool/LLM outputs the production trace saw.
  • Offline & free — no live keys or spend in CI; lambda: … is never called under fixtures.
  • Faithful — recorded errors replay as ToolError, so error-handling logic is gated, not bypassed.

The reference agent in sdk/examples/weather_agent.py wires this up (run, run_broken, run_handles, …) and is what tracely replay --entrypoint weather_agent:run executes.

Replay is single-turn and hermetic: it re-runs promoted regression cases through your own code against recorded fixtures. To gate multi-turn behaviour against a live service — in any language, with no agent code in CI — see Scenarios. Most teams run both.

Next: the CI gate CLI.

Comparing tools? Langfuse alternatives, compared honestly covers how hermetic replay differs from dataset-and-live-calls CI, including where the other approach wins.