Measure your AI agent's projected performance and identify what to improve before going live.
Agent Evaluation analyzes your agent's knowledge base, chat history, and configuration to produce a comprehensive performance report with scored metrics, prioritized action items, and knowledge base gap analysis.
Running an Evaluation
- Go to Settings > AI > Agent Evaluation, or navigate directly to app.heymarket.com/admin/hey-ai/ai-evaluation/.
- Click Evaluate Again in the top right.
- In the dialog, choose your evaluation scope:
- Include chats and messages — uses real conversation history alongside your knowledge base for a more comprehensive evaluation (recommended).
- Knowledge base only — evaluates based solely on your knowledge base contents.
- Click Run Evaluation.
The evaluation runs in the background. Results appear on the same page once complete.
Note: You can re-run evaluations at any time, especially after adding knowledge base content or updating agent instructions. Each run reflects your current configuration.
Limit: Evaluations are currently capped at 2 per agent per day.
Reading Your Results
Overall Agent Performance Projection
A single percentage score summarizing projected agent performance across all test cases. Scores below 70% typically indicate significant gaps worth addressing before enabling autonomous mode.
Performance Metrics
Six dimensions are scored individually:
| Metric | What It Measures |
|---|---|
| Completeness | Whether responses fully address the customer's question |
| Tone & Empathy | Whether responses are appropriate and professional in tone |
| Helpfulness | Whether responses provide actionable, useful information |
| Domain Compliance | Whether the agent stays within its defined scope |
| Answer Relevancy | Whether responses are on-topic and directly relevant |
| Faithfulness | Whether responses accurately reflect knowledge base content |
Metrics marked with an orange dot are underperforming and should be prioritized.
Supervised Mode Recommendation
If your overall score is below a certain threshold, Agent Evaluation will recommend starting in Supervised Mode. In supervised mode, your team reviews and approves agent responses before they are sent. This is the default recommendation for new agents.
Priority Action Items
The evaluation surfaces up to three HIGH priority action items, each with:
- A description of the issue
- The expected improvement if resolved (e.g., "would fix failures across at least 15 test cases")
Common action item types include:
- Missing knowledge base content — topics the agent was asked about but had no information to draw from (e.g., 10DLC compliance, email configuration, call forwarding).
- Agent behavior issues — instruction-level problems such as over-routing to humans, sending bare "HUMAN ROUTING" outputs without a customer-facing message, or asking clarifying questions before leading with available answers.
Addressing high-priority action items typically produces the largest score improvements.
Knowledge Base Gaps
Below the action items, a Knowledge Base Gaps section lists topic areas where the agent lacked sufficient content to answer test questions. Each topic shows the number of affected test cases.
To resolve a gap: add or expand knowledge base articles covering that topic, then re-run the evaluation to confirm improvement.
Behavior & Response Issues
This section lists recurring agent behavior patterns that hurt performance, along with the number of test cases affected. Examples include:
- Agent over-routes to human for routine questions without attempting to answer first
- Agent asks unnecessary clarifying questions instead of leading with available information
- Agent sends bare routing labels instead of customer-facing messages
These issues are typically resolved by updating your Agent Behavior instructions (Settings > AI > Agent Behavior).
Frequently Asked Questions
How often should I run an evaluation? Run an evaluation whenever you make significant changes to your knowledge base or agent instructions. For new agents, run it before going live and again after your first week of real conversations.
Does the evaluation use real customer data? If you select "include chats and messages," the evaluation uses your existing conversation history to generate test cases. No data is sent externally beyond what already powers Heymarket's AI features.
What should I do if my score is low? Start with the highest-priority action items, which will have the most impact. Add missing knowledge base articles first, then review agent behavior instructions. Re-run the evaluation after each significant change.
When is Autonomous Mode appropriate? Once your overall score is consistently above 80% and high-priority behavior issues have been resolved, your agent is typically ready for autonomous mode. Consult your Heymarket contact before switching if you are unsure.