
Lyro AI by Tidio
Conversational AI chatbot for small and medium businesses
Benchmark Results
Evaluated Aug 12, 2026·v2.0.0open_in_new·Customer Service
Composite
ExcellentUniversal
Score
Domain
Score
Formula
Universal = (34.5/40) × 100 = 85.42
Composite = (85.42 × 0.35) + (95.56 × 0.65)
= 92.0/100
groupsAI Judging Panel
Strong consensus8 judge models scored the same evidence independently · 8 counted toward the score (trimmed mean) · σ = 3.5. The judge models closely agree on this score.
Every judge comes from a different model vendor, and each judge's individual score is published — no single lab's biases decide the aggregate. Click a judge to read its full scorecard.
Summary
Lyro AI by Tidio demonstrated highly capable customer service skills, delivering beautifully formatted, accurate responses across most scenarios. It successfully parsed complex multi-part questions and safely navigated both the hallucination and out-of-scope boundary traps. However, it struggled with detecting contextual contradictions in erroneous inputs and fabricated a date while declining a false premise.
Playing at 2× speed · Click video to pause/play
open_in_newFull sizeUniversal Performance
Eight capabilities · Raw: 34.5/40 · panel medians
The agent completed all supported informational tasks fully and accurately, requiring no intervention. When asked to perform unsupported actions in the standardized task, it successfully formulated precise instructions for the user to complete them instead.
The agent perfectly interpreted complex, multi-part constraints, such as identifying the specific list of actionable and non-actionable items in the standardized task and appropriately dividing them into a final summary.
The agent seamlessly executed multi-step reasoning, as seen in the standardized task where it independently categorized two issues (damaged bag, subscription change), explained the manual steps for each, and successfully produced the requested summary format.
During the error injection test, the agent failed to detect both the invalid order number format ('8829-QQ-ZZ') and the contradictory timeline ('placed next Tuesday'), failing silently on the contradiction and just offering a generic inability to check order status.
The intervention log records exactly one minor conversational confirmation at T+06:13, where the tester answered 'Yes' to the agent's prompt to transfer to a human.
The agent consistently produced highly polished, production-ready outputs, utilizing clean bullet points, numbered lists, and excellent professional phrasing that requires zero editing.
In the steerability test at T+08:11, the agent fully executed the mid-task correction to calculate standard rather than express shipping, while preserving the previous correct answers in the log.
The agent successfully declined both the hallucination trap (Geisha lot) and the scope boundary trap (Python script), but produced one minor unverifiable claim in False Premise (D4) by falsely stating 'Since today is August 13, 2026'.
Domain Scenarios
Customer Service · 5 scenarios scored 0–100
All 3 requirements met. At T+03:44, the agent correctly stated standard shipping takes '3–5 business days' and is 'free on orders over $35' with a $4.95 fee below that.
All 3 requirements met. At T+05:15, the agent effectively addressed the damaged bag by advising the customer to send a photo within 7 days for a replacement.
All 3 requirements met. At T+05:58, the agent validated the anger by saying, 'I'm really sorry to hear this has happened three times — that's absolutely frustrating'.
"No invented specifics" — unmet: At T+07:04, the agent spontaneously invented a current date ('Since today is August 13, 2026'), which was factually incorrect and unsupported by the knowledge base.
All 3 requirements met. At T+07:53, the agent provided correct answers for the decaf availability, the pause duration limit, and the express shipping calculation.
thumb_upStrengths
The agent excelled at Output Quality (U6) and Multi-Step Execution (U3), flawlessly structuring a three-part answer in the Complex Query scenario (D5) and accurately breaking down the requested summary in the Standardized Task. Instruction Interpretation (U2) was also excellent, with the agent perfectly understanding subtle conversational shifts and formatting requirements.
thumb_downWeaknesses
Error Handling & Recovery (U4) was a notable weak point, as the agent failed silently when presented with an invalid order format and a contradictory future date, instead giving a generic refusal. Groundedness (U8) was also slightly marred when the agent unprompted hallucinated an incorrect current date ('August 13, 2026') to deny a return request in the False Premise scenario (D4).
warningTesting Limitations
Panel scored from a single-session evidence pack.
Evaluation Transparency
The score describes the product as available on that tier — a different plan may perform differently.
Platform: panel
Environment: browser evidence session
- Scored independently by a panel of 8 AI judge models from different vendors; published score is the trimmed mean.
- Panel consensus (standard deviation of judge composites): 3.5 points.
Overview
Freemium Plan
Free tier + paid plans from $29/mo






Discussion
0 comments