Hermetic replay
This is the seam that makes “the recorded run is the test” real: the same agent code runs live in production and deterministically offline in CI — no API keys, no cost, no flakiness.
The idea
Wrap each external call in call_tool / call_llm instead of calling it directly:
import tracely_sdk as tracely
def run(user_input: str) -> str:
with tracely.agent("support-agent"):
order = tracely.call_tool("get_order", lambda: get_order(order_id),
args={"order_id": order_id})
answer = tracely.call_llm("gpt-4o", lambda: chat(messages),
input=messages, usage=(812, 96))
return answer- In production (no fixtures active):
call_tool/call_llminvoke yourfn, record its output (and any error) on the span, and return it — a normal traced run. - In CI replay:
tracely replayloads the regression case’s recorded fixture bundle and activates it;call_tool/call_llmthen serve the recorded outputs (in order, or matched byargs) and never call yourfn.
A call that errored in production is reproduced on the replayed span and raised as
tracely.ToolError — so your agent’s own try/except runs exactly as it would live, and the
gate sees the same failure condition. Faithful error-handling replay, for free.
Recorded means recorded
fixtures(None) is “no replay requested” — everything runs live. Any other bundle, including {}
(an explicitly empty recording), is recorded mode, strict by default: a call the recording
cannot answer raises tracely.ReplayError instead of running your function, so a recorded
execution can never quietly become a live one.
ReplayError.reason | When |
|---|---|
missing | nothing was recorded under that tool / model name |
exhausted | every recorded call for that name has already been served |
mismatch | a tool was called with args the recording doesn’t have — a different call is never consumed just because the name matches |
Matching lives in one place and is deliberately simple: an entry recorded without args matches by
order; otherwise both sides are canonicalised (JSON strings parsed, key order and tuple/list
ignored) and compared for equality. LLM calls are looked up by model id, falling back to the next
recorded LLM call of any name (auto-instrumentor span names are not model ids). An LLM call whose
input differs from the recording is still served in order but stamped
tracely.replay.divergence=input and reported: recorded-model replay tests your orchestration
under the recorded conditions, and a changed prompt is exactly the divergence it must show, not
hide. Recorded model outputs do not prove a changed prompt produces a better answer.
The block yields a ReplayReport:
with tracely.fixtures(bundle) as rep:
run(case_input)
rep.mode # "recorded" | "lenient" | "live"
rep.served # ["tools:get_order", "llm:gpt-4o"]
rep.unused # recorded calls the run never asked for — divergence evidence, not auto-fail
rep.diverged # served despite a different input
rep.errors # the strict misses that were raised
rep.providers # which SDK create-methods were patched in this process
rep.clean # strict run that reproduced the recording exactlyfixtures(bundle, strict=False) keeps the pre-0.4.3 lenient behaviour (live fall-through on a
miss, ordered serve on a mismatch) for callers migrating. Its report says lenient — it is never
presented as a recorded run. tracely replay --lenient is the CLI form.
No rewrite needed: @observe tools and auto-instrumented LLM calls replay too
The call_tool/call_llm seam is the fully-portable path, but most agents don’t need a rewrite:
- Tools — a function decorated
@tracely.observe(as_type="tool")is transparently hermetic: under fixtures it serves the recorded output (or re-raises the recorded failure astracely.ToolError) instead of running. Decorate your tools and the tool side replays as-is. - LLM calls — inside a
fixtures()block Tracely class-patches the provider create-methods below, so code that calls the SDK directly (underinstrument="auto", or via the drop-ins) is served the recorded completion — reconstructed into a provider-shaped response object — and never hits the network. A strict miss raisesReplayErrorat the call site, likecall_llm.
What is actually intercepted
| Provider | Entry point | Sync / async | Streaming | Served shape | Tested against |
|---|---|---|---|---|---|
| OpenAI | chat.completions.create | both | no — stream=True goes live | duck-typed ChatCompletion (choices[0].message.content / .tool_calls; no usage) | real openai client |
| Anthropic | messages.create | both | no — messages.stream goes live | duck-typed Message (text / tool_use blocks) | real anthropic client |
| Google GenAI | models.generate_content | both | no | .text, .function_calls, candidates[0].content.parts | real class when installed |
| Mistral | chat.complete / complete_async | both | no | OpenAI-shaped | stand-in only |
| LiteLLM | completion / acompletion | both | no | OpenAI-shaped | stand-in only |
| Anything else | — | — | — | live in replay | — |
A surface outside this table — the OpenAI Responses API, Bedrock, Cohere, a framework’s own
HTTP client, streaming on any provider — is live in replay; the SDK cannot see that call.
The report catches it after the fact: recorded calls the run never asked for (rep.unused,
also logged) mean something went to the network. The SDK does not sandbox Python or deny
egress. If you need that guarantee, deny provider/tool egress in the CI executor itself, or use
call_llm for a provider-agnostic seam.
Activating fixtures manually
tracely replay does this for you, but the primitive is public:
with tracely.fixtures(bundle): # bundle = the case's recorded tool/LLM outputs
run(case_input) # call_tool/call_llm now serve recorded values
# outside the block, calls are live againfixture(kind, name) peeks the next recorded output for a tool/model without consuming it.
After a fix, re-record. The recording is a recording of the bug. A fix that adds a call the
bug never made (the silent-failure case: the fixed agent now calls the tool) cannot be served from
it — strict replay reports INCOMPLETE — no recorded call for tools:get_weather, never a live
call. Once the fixed code has run for real (CI, staging, production), re-record the case from that
run: POST /api/cases/{id}/recapture with the trace id (same input required; bumps the case
version). From then on the suite replays the fix hermetically and a regression (run_broken)
fails on the missing tool. The same applies to a changed prompt: recorded-model replay serves the
old answer back, so a quality fix is shown by re-recording from the faithful run, not by replaying
the old recording.
tracely replay itself refuses to run a case whose recording cannot be loaded: it emits the
case’s trace errored without running your agent, so the gate fails loudly instead of grading a
live (or “nothing recorded”) run as the recorded case.
Why this matters
- Deterministic — the agent sees the exact tool/LLM outputs the production trace saw.
- Offline & free — no live keys or spend in CI;
lambda: …is never called under fixtures. - Faithful — recorded errors replay as
ToolError, so error-handling logic is gated, not bypassed.
The reference agent in
sdk/examples/weather_agent.py
wires this up (run, run_broken, run_handles, …) and is what tracely replay --entrypoint weather_agent:run executes.
Replay is single-turn and hermetic: it re-runs promoted regression cases through your own code against recorded fixtures. To gate multi-turn behaviour against a live service — in any language, with no agent code in CI — see Scenarios. Most teams run both.
Next: the CI gate CLI.
Comparing tools? Langfuse alternatives, compared honestly covers how hermetic replay differs from dataset-and-live-calls CI, including where the other approach wins.