
Fin
Perfect customer experiences made possible
Benchmark Results
Evaluated Aug 10, 2026·v2.0.0open_in_new·Customer Service
Composite
ExcellentUniversal
Score
Domain
Score
Formula
Universal = (31.5/40) × 100 = 79.58
Composite = (79.58 × 0.35) + (97.33 × 0.65)
= 91.0/100
groupsAI Judging Panel
Strong consensus8 judge models scored the same evidence independently · 8 counted toward the score (trimmed mean) · σ = 3.1. The judge models closely agree on this score.
Every judge comes from a different model vendor, and each judge's individual score is published — no single lab's biases decide the aggregate. Click a judge to read its full scorecard.
Summary
Fin delivered accurate, well-grounded answers across most customer-service scenarios, correctly handling informational queries, topic changes, complex multi-part questions, and false-premise rejections. Its main weaknesses were failing to detect a self-contradictory instruction in the error-injection test and closing the conversation after a simulated handover, which left a follow-up message unanswered and required a manual reset.
Playing at 2× speed · Click video to pause/play
open_in_newFull sizeUniversal Performance
Eight capabilities · Raw: 31.5/40 · panel medians
Fin completed most tasks end to end — the Informational Query, Topic Change, Emotional Message, Complex Query, the standardized prioritization task, and both hallucination and scope-boundary tests were all resolved correctly. However, the False Premise scenario required a manual conversation reset after Fin failed to respond in the closed post-handover thread, and the error-injection test produced an incomplete answer that ignored the contradictory timeline and the cancel demand.
Fin accurately interpreted precise instructions across all five domain scenarios, including the three-part Complex Query and the mid-conversation Topic Change. However, on the error-injection test it failed to recognize the self-contradictory premise — an order claimed placed both 10 minutes and three weeks ago — and addressed only one timeline and one of the two demanded actions.
Fin executed all multi-step chains correctly: the standardized task required sorting three orders by priority, stating the action for each, and summarizing — all done in correct order with full context. The Complex Query's three sub-questions were answered in a single coherent bulleted response, and the Topic Change retained relevant earlier context about headphones.
Fin failed to detect the self-contradictory order-timing premise in the error-injection test, producing a response that silently ignored the 'three weeks ago' claim and the cancel demand while addressing only the upgrade path. Separately, after the Emotional Message handover Fin closed the conversation and produced no reply to a follow-up message for approximately 2 minutes, requiring a manual reset — a silent failure with no detection or communication.
One rescue was logged at T+05:30: the tester reset the preview conversation and re-sent the False Premise message after Fin failed to respond in the closed post-handover thread. Per the scoring rules, any rescue caps the score at 3.
Most outputs were production-ready: the Informational Query, Topic Change, False Premise, and Complex Query scenarios delivered accurate, well-formatted answers with source citations, and the standardized task produced a clear prioritized list with summary. The error-injection response was incomplete — it explained the 60-minute change window but never addressed the cancel demand or flagged the contradiction — and the Emotional Message response offered no specific resolution path for the broken headphones beyond handover.
The scripted mid-Complex Query correction ('assume the order is $90, not $60, and answer only the shipping part again') was acknowledged and fully executed on the next turn: Fin correctly recalculated free shipping for the $90 order, mentioned the express option, and did not repeat the address or Saturday-hours portions. The prior Complex Query answer remained visible and intact.
Fin declined both trap questions cleanly: it did not confirm the fabricated '2025 impact report' statistic, instead stating 'We don't have a specific percentage for headphone returns in the info here' and asking the user to share the source, and it refused the out-of-scope Python request ('I couldn't find any information related to your question about the Python function'). No invented specifics appeared in any session response.
Domain Scenarios
Customer Service · 5 scenarios scored 0–100
All 3 requirements met. Fin's answer — 'return within 30 days of delivery if unopened,' buyer pays return shipping for change of mind, Northwind pays for faulty items — matches the KB text verbatim.
All 3 requirements met. After the tester said 'Actually forget shipping — what warranty do the headphones come with?', Fin answered the warranty question (2-year warranty for headphones, 1 year for accessories).
"A concrete resolution step is offered" — partial: Fin offered to connect with a human agent but did not reference the warranty policy or any specific resolution path for the broken headphones (e.g., a warranty return); the only step offered was escalation, which overlaps with the next requirement.
All 3 requirements met. Fin responded: 'that plan doesn't actually exist. We don't offer a lifetime replacement guarantee or a Platinum Care plan for headphones.'
All 3 requirements met. All three parts were answered: shipping cost for a $60 order ($12 express, since $60 is under the $75 free-shipping threshold), address change after 20 minutes (yes, within the 60-minute window), and Saturday hours (closed).
thumb_upStrengths
Fin excelled at grounded, knowledge-base-accurate responses: the Informational Query, Topic Change, False Premise, and Complex Query scenarios all produced correct answers matching the knowledge article with source citations and no invented details. The steerability correction was executed perfectly — Fin recalculated shipping for a $90 order on the next turn while preserving prior work — and both the hallucination and scope-boundary traps were cleanly declined.
thumb_downWeaknesses
Fin did not detect the contradiction in the error-injection test (an order claimed placed both 10 minutes and three weeks ago), silently addressing only one timeline and ignoring the cancel demand. After the Emotional Message handover, Fin closed the conversation and failed to respond to a follow-up message for approximately 2 minutes, requiring a manual reset to proceed with the False Premise scenario.
warningTesting Limitations
Panel scored from a single-session evidence pack.
Evaluation Transparency
The score describes the product as available on that tier — a different plan may perform differently.
Platform: panel
Environment: browser evidence session
- Scored independently by a panel of 8 AI judge models from different vendors; published score is the trimmed mean.
- Panel consensus (standard deviation of judge composites): 3.1 points.
Overview
Usage-Based Pricing
From $0.99 per outcome, 14-day free trial



Discussion
0 comments