
Zendesk AI Agents
Self-improving AI agents built for resolution
Benchmark Results
Evaluated Aug 10, 2026·v2.0.0open_in_new·Customer Service
Composite
Very GoodUniversal
Score
Domain
Score
Formula
Universal = (33.5/40) × 100 = 84.17
Composite = (84.17 × 0.35) + (90.67 × 0.65)
= 88.1/100
groupsAI Judging Panel
Strong consensus8 judge models scored the same evidence independently · 8 counted toward the score (trimmed mean) · σ = 2.9. The judge models closely agree on this score.
Every judge comes from a different model vendor, and each judge's individual score is published — no single lab's biases decide the aggregate. Click a judge to read its full scorecard.
Summary
The Zendesk AI Agent demonstrates strong comprehension, instruction adherence, and groundedness when handling policy queries and multi-step tasks. However, it suffers from a critical stability issue where the chat thread hangs silently after an escalation, requiring a manual page reload to resume functionality.
Playing at 2× speed · Click video to pause/play
open_in_newFull sizeUniversal Performance
Eight capabilities · Raw: 33.5/40 · panel medians
The agent successfully completed most informational and complex tasks end-to-end with accurate outputs, though it stalled indefinitely on one task following an escalation turn, requiring user intervention.
The agent perfectly interpreted complex instructions, successfully parsing a contradictory prompt by evaluating both conditions independently and correctly addressing a three-part query in Complex Query (D5).
The agent flawlessly executed the standardized task, sorting three order issues, providing policy actions for each, and synthesizing them into a concluding summary within a single coherent response.
Following the Emotional Message (D3) escalation, the agent failed silently and hung for over two minutes on the subsequent False Premise (D4) query, producing no output and forcing a manual page reload, which matches the anchor for crashing or hanging.
The session required a manual page reload to rescue the agent from a dead thread post-escalation, capping the autonomy score at 3, alongside one minor confirmation during knowledge setup.
The generated responses were production-ready, featuring clear, well-formed language, correct capitalization and formatting, and no need for human editing.
The agent immediately and accurately executed a mid-task correction in Complex Query (D5), recalculating shipping for a $90 order and restricting its response to only the shipping topic as requested.
The agent demonstrated strict groundedness, actively denying the fabricated False Premise (D4) premise, refusing the out-of-scope Python coding task, and declining to invent statistics for the hallucination test.
Domain Scenarios
Customer Service · 5 scenarios scored 0–100
All 3 requirements met. The agent correctly stated 'You may return unopened items within 30 days of delivery' and that the buyer pays return shipping for a change of mind.
All 3 requirements met. The agent successfully answered the new question, stating 'The headphones include a 2-year warranty.'
"Tone acknowledges the customer’s frustration" — unmet: The agent responded with a dismissive 'No problem', completely failing to acknowledge the customer's explicitly stated fury. · "A concrete resolution step is offered" — partial: The agent asked the user to 'please leave your details', which facilitates an escalation but does not directly resolve the broken headphone issue.
All 3 requirements met. The agent explicitly denied the premise, stating 'There is no lifetime replacement guarantee or Platinum Care plan'.
All 3 requirements met. The agent addressed the shipping cost for a $60 order, the 60-minute address change window, and Saturday support closure.
thumb_upStrengths
The agent excels at multi-step execution and instruction following, perfectly formatting and summarizing prioritized actions in the standardized task. It is also highly grounded, strictly adhering to its knowledge base and consistently refusing to hallucinate statistics, write out-of-scope code, or validate fabricated policies.
thumb_downWeaknesses
The most significant weakness is a system-level hang triggered by the escalation workflow, leaving the chat completely unresponsive to subsequent messages until the page is manually reloaded. Additionally, the agent lacks empathy in emotional scenarios, responding to a furious customer with a dismissive 'No problem' rather than acknowledging their frustration.
warningTesting Limitations
Panel scored from a single-session evidence pack.
Evaluation Transparency
The score describes the product as available on that tier — a different plan may perform differently.
Platform: panel
Environment: browser evidence session
- Scored independently by a panel of 8 AI judge models from different vendors; published score is the trimmed mean.
- Panel consensus (standard deviation of judge composites): 2.9 points.
Overview
Paid Plans
From $19/agent/mo, AI agents billed per resolution

Discussion
0 comments