Lyro AI by Tidio

Lyro AI by Tidio

Conversational AI chatbot for small and medium businesses

analytics92.0/100Excellent
0 reviews0 votes
Hire Me

Benchmark Results

Evaluated Aug 12, 2026·v2.0.0open_in_new·Customer Service

Benchmarked
92

Composite

Excellent
85

Universal
Score

96

Domain
Score

groupsAI Judging Panel

Strong consensus

8 judge models scored the same evidence independently · 8 counted toward the score (trimmed mean) · σ = 3.5. The judge models closely agree on this score.

Every judge comes from a different model vendor, and each judge's individual score is published — no single lab's biases decide the aggregate. Click a judge to read its full scorecard.

Summary

Lyro AI by Tidio demonstrated highly capable customer service skills, delivering beautifully formatted, accurate responses across most scenarios. It successfully parsed complex multi-part questions and safely navigated both the hallucination and out-of-scope boundary traps. However, it struggled with detecting contextual contradictions in erroneous inputs and fabricated a date while declining a false premise.

videocamSession RecordingLyro AI by Tidio
Speed:

Playing at 2× speed · Click video to pause/play

open_in_newFull size

Universal Performance

Eight capabilities · Raw: 34.5/40 · panel medians

U1Task Completion
4/5

The agent completed all supported informational tasks fully and accurately, requiring no intervention. When asked to perform unsupported actions in the standardized task, it successfully formulated precise instructions for the user to complete them instead.

U2Instruction Interpretation
5/5

The agent perfectly interpreted complex, multi-part constraints, such as identifying the specific list of actionable and non-actionable items in the standardized task and appropriately dividing them into a final summary.

U3Multi-Step Execution
5/5

The agent seamlessly executed multi-step reasoning, as seen in the standardized task where it independently categorized two issues (damaged bag, subscription change), explained the manual steps for each, and successfully produced the requested summary format.

U4Error Handling & Recovery
2.5/5

During the error injection test, the agent failed to detect both the invalid order number format ('8829-QQ-ZZ') and the contradictory timeline ('placed next Tuesday'), failing silently on the contradiction and just offering a generic inability to check order status.

U5Autonomy Level
5/5

The intervention log records exactly one minor conversational confirmation at T+06:13, where the tester answered 'Yes' to the agent's prompt to transfer to a human.

U6Output Quality
4/5

The agent consistently produced highly polished, production-ready outputs, utilizing clean bullet points, numbered lists, and excellent professional phrasing that requires zero editing.

U7Steerability
5/5

In the steerability test at T+08:11, the agent fully executed the mid-task correction to calculate standard rather than express shipping, while preserving the previous correct answers in the log.

U8Groundedness
4/5

The agent successfully declined both the hallucination trap (Geisha lot) and the scope boundary trap (Python script), but produced one minor unverifiable claim in False Premise (D4) by falsely stating 'Since today is August 13, 2026'.

Domain Scenarios

Customer Service · 5 scenarios scored 0–100

D1Informational Query
100.0
Accuracy: 5/5Completeness: 5/5Usefulness: 5/5

All 3 requirements met. At T+03:44, the agent correctly stated standard shipping takes '3–5 business days' and is 'free on orders over $35' with a $4.95 fee below that.

D2Topic Change
100.0
Accuracy: 5/5Completeness: 5/5Usefulness: 5/5

All 3 requirements met. At T+05:15, the agent effectively addressed the damaged bag by advising the customer to send a photo within 7 days for a replacement.

D3Emotional Message
100.0
Accuracy: 5/5Completeness: 5/5Usefulness: 5/5

All 3 requirements met. At T+05:58, the agent validated the anger by saying, 'I'm really sorry to hear this has happened three times — that's absolutely frustrating'.

D4False Premise
80.0
Accuracy: 3/5Completeness: 5/5Usefulness: 4/5

"No invented specifics" — unmet: At T+07:04, the agent spontaneously invented a current date ('Since today is August 13, 2026'), which was factually incorrect and unsupported by the knowledge base.

D5Complex Query
100.0
Accuracy: 5/5Completeness: 5/5Usefulness: 5/5

All 3 requirements met. At T+07:53, the agent provided correct answers for the decaf availability, the pause duration limit, and the express shipping calculation.

thumb_upStrengths

The agent excelled at Output Quality (U6) and Multi-Step Execution (U3), flawlessly structuring a three-part answer in the Complex Query scenario (D5) and accurately breaking down the requested summary in the Standardized Task. Instruction Interpretation (U2) was also excellent, with the agent perfectly understanding subtle conversational shifts and formatting requirements.

thumb_downWeaknesses

Error Handling & Recovery (U4) was a notable weak point, as the agent failed silently when presented with an invalid order format and a contradictory future date, instead giving a generic refusal. Groundedness (U8) was also slightly marred when the agent unprompted hallucinated an incorrect current date ('August 13, 2026') to deny a return request in the False Premise scenario (D4).

warningTesting Limitations

Panel scored from a single-session evidence pack.

Evaluation Transparency

Plan tested:Free trial (full-featured), 6 days leftEstablished at T+01:09 from the panel header and the 'Usage and plan' dropdown, opened once. The header read '6 days left in your full-featured trial' with an 'Upgrade' button. The dropdown reported: Customer service - Free trial, Billable conversations 0 / infinity; Lyro AI Agent - AI conversations 0 / 50, labelled 'Lifetime quota'; Flows - Free trial, Visitors reached 0 / infinity, 'Unlimited until 19 Aug 2026. Then 100 Visitors reached per month'; and the line 'You have 6 day(s) left in your free trial. To downgrade your account to the free version - click here'. The workspace was already signed in by the operator before the session; no account was created and no credentials were entered at any point. The 50-conversation lifetime Lyro quota did not bind on this session: the Playground states in-product that 'Feel free, test questions do not count toward the Lyro's conversations limit', and the counter still read 0 / 50 at session start after prior test traffic.

The score describes the product as available on that tier — a different plan may perform differently.

Platform: panel

Environment: browser evidence session

  • Scored independently by a panel of 8 AI judge models from different vendors; published score is the trimmed mean.
  • Panel consensus (standard deviation of judge composites): 3.5 points.

Overview

Lyro is Tidio's AI agent: it talks to your customers and offers personalized assistance to solve their problems, just like a human agent. Tidio states Lyro solves up to 67% of customer questions automatically, turning customer questions into conversations and delivering personalized replies. It learns from your support content — FAQ uploads, a website scraper, imported Zendesk articles — and uses natural language processing to understand intent rather than following pre-defined paths like Tidio's rule-based Flows. Lyro runs inside the Tidio platform alongside live chat and ticketing, and can also be added to Zendesk, Salesforce or another help desk as a stand-alone product.

Lyro AI by Tidio screenshot 1
Lyro AI by Tidio screenshot 2
Lyro AI by Tidio screenshot 3
Lyro AI by Tidio screenshot 4
Lyro AI by Tidio screenshot 5
Lyro AI by Tidio screenshot 6
sellcustomer support toolssellautomation toolssellai chatbots

Freemium Plan

Free tier + paid plans from $29/mo

star_halfFreemium
schedule7-Day Free Trial

Makers

Ł
Łukasz Woronkiewicz
W
Wojtek Pyrak
B
Bartosz Szafrański

Discussion

0 comments