Evaluations

Evaluations

Every evaluator you add is a column in the trace table. A column has three choices that decide what the judge reads and when it runs: a level, an execution mode, and a prompt style (basic rubric or advanced template).

Levels

LevelOne result perThe judge reads
Conversationthreadthe whole transcript, turn by turn
Messageturnthat turn’s user request and agent answer
Stepspan (tool call, generation, …)that step’s input and output

Batch vs sequential

Batch (default) grades every item independently — nothing carries over between items, so two grades of the same item are directly comparable.

Sequential grades items in order as one running conversation with the judge: the rubric is the system prompt, then item → verdict → item → verdict. A step column chains the steps within a message; a message column chains the turns of a thread. Each item is graded knowing what came before it — and the previous result of the same metric is available for continuity.

Sequential grading is incremental: when new turns arrive, only the new turns are graded and appended to the judge’s conversation. Re-running a column from the UI re-grades everything from the start. Conversation-level columns have a single item, so the mode makes no difference there.

Advanced templates and sequential mode

An advanced prompt (any prompt containing @VARIABLES like @HISTORY or @CURRENT_STEP) controls its own context: the resolved template is the whole prompt, and nothing is added around it.

Because of that, sequential works differently for advanced columns: there is no running conversation with the judge. Chaining happens only through @METRIC_PREVIOUS_RESULT, which resolves to the previous item’s result of the same metric — put it in your template where you want the judge to see it. If your template doesn’t reference it, an advanced sequential column behaves exactly like batch.

Output types

score (0–1 with a pass/fail threshold), number, boolean, text, or json — your own schema, built in the Add Column form. Enum fields are enforced on the model, and a numeric score field in your schema drives the pass/fail verdict via the threshold.

Decision models (TypeSafe Jev)

Pick TypeSafe Jev 1.13 as a column’s model and the column stops being a rubric. Jev is a classifier on OpenRouter’s Decisions API: it reads the graded item (the same request/answer, transcript or step I/O an LLM judge would get, or your @VARIABLE template in Advanced mode) and answers one question with calibrated probabilities. You write the question and its criteria:

Question typeYou defineCell showsVerdict
Binary (yes / no)what yes and no meanP(yes), 0–1PASS when P(yes) ≥ threshold (or below it, with Pass when: No)
Multi-class (pick one)2–255 labels, each with an optional descriptionthe chosen labelFAIL when the label is ticked fail; a label only when none are
Multi-label (pick any)2–50 labelsevery label with P ≥ thresholdFAIL when any ticked fail label applies

Multi-label asks one yes/no question per label, all in a single request, so each label is judged on its own and adding labels barely changes cost or latency.

Write the question literally, one judgment per column, and name the part of the item you mean (`Agent answer`). Ask “Did the agent hallucinate?” with Pass when: No rather than inverting the criteria. Jev bills input tokens only, at a fraction of an LLM judge’s cost, and needs a workspace OpenRouter key.

The built-in library (Browse Library in Add Column) is written for Jev: every template is a question with labels, and the old JSON sub-fields (severity, detected type) became the labels themselves. Each keeps its score name, so existing gates and alerts that reference it keep working.

Chaining columns: Depends On and Run only when

Depends On runs other columns first and hands their result for the same item to this one — as context in a basic prompt, or through @DEPENDENCIES in a template. Each result carries its label, value, verdict and reason.

Run only when grades an item only when conditions on those results hold (all of them). A condition reads a column’s label (a multi-class label, or any multi-label label), its verdict (PASS / FAIL / NONE) or its value (≥, >, ≤, <). Items that don’t match show Not run: no verdict, no model call.

The cheap-first pattern: let a Jev column classify every turn, and run the expensive LLM only where it flagged something — explain the turns labelled refund_cancellation, or diagnose only the turns where Answer quality FAILED:

{
  "model": "openai/gpt-6-luna",
  "output_type": "text",
  "prompt": "Explain how the agent handled this refund request.",
  "run_if": [{"column": "tracely.run.intent", "field": "label", "op": "in", "values": ["refund_cancellation"]}]
}

A condition’s column joins Depends On automatically. If it has no result for an item, the condition is false and the column doesn’t run. The in-app assistant knows this pattern — ask it for “an LLM explanation only when the intent is refund”.

Long items and model context

Every column is checked against its model’s context window before a call is made (Jev reads at most 32k tokens; an LLM’s limit comes from OpenRouter’s catalog). Nothing is truncated. An item that doesn’t fit is either graded by the column’s fallback model — an LLM with a larger context, set under If an item exceeds the context — or recorded as Skipped — input too long, which carries no verdict and never fails a gate. A Jev column’s fallback answers the same question with the same options, so the column keeps one score shape.

Further reading

New to this? LLM evaluation: metrics, methods and tools covers the concepts behind this page — offline versus online evaluation, when a deterministic check beats a judge, why the evaluation level decides what you can catch, and how to calibrate a judge before you let it gate a release.