# Trust Lab

_Evals that keep your agent honest in production_

Turn real production conversations into tests. Grade whether the agent did what the user actually wanted, catch regressions before users do, and ship changes with confidence.

## Ship agent changes without crossing your fingers

Trust Lab is the evaluation layer that measures whether your agent is actually improving. Change a prompt, a tool, or a model — and see what it did to quality before it reaches a customer.

- **Golden sets** — a saved battery of real prompts with expected outcomes, rerun on every change
- **Outcome matching** — grade whether the result matched user intent, not just that it ran
- **Shadow runs & scheduled checks** — catch regressions before users hit them
- **Model swaps, safely** — prove a faster or cheaper model loses no quality on your own traffic

Without evals, you're compounding blindly — the agent might get faster at delivering the wrong thing. Trust Lab turns raw iteration into directional improvement.
