Agent testing & quality
Test the change. Read the evidence.
Evaluate scripted cases, score conversations and compare agent variants.
Explore the workflowTest the change. Read the evidence.
A prompt change deserves more than a quick listen. Build repeatable evaluation cases, inspect the results and compare variants against a metric you choose in advance.
Interactive example No live actions
Evaluation example
Example case / A customer asks for a human.
Step 1 of 3 / Test case
Inside Agent testing & quality
Open a capability to read how it works.
Repeatable scenarios
Create an evaluation suite for an agent with caller personas, opening lines and success criteria. Include normal requests and difficult boundaries.
Inspect each result
Read the transcript and per-assertion result for each case instead of relying only on a summary pass count.
Score what matters
Use scorecards to define your quality criteria and inspect the breakdown for a conversation or evaluation run.
Compare agent variants
Configure experiment arms and traffic splits, assign callers through the experiment API and read results against a selected metric. Organize specialized agents into squads where supported.
Make your
first connection.
Read the documentation- 1
Choose an agent and write representative test cases with explicit success criteria.
- 2
Run the suite and inspect failures alongside their transcripts.
- 3
Test revisions again or configure an experiment with a control and a chosen metric.
Keep connecting.
Browse all products
AI agent builder
Give your agent instructions, knowledge and tools, then configure it for its channel.
Explore
Call operations
Manage numbers, menus, queues, outbound campaigns and live call oversight.
Explore
Tools & automations
Connect agents to your business tools, calendars, custom APIs and MCP servers.
Explore
