Fin

Fin

Perfect customer experiences made possible

analytics91.0/100Excellent
0 reviews0 votes
Hire Me

Benchmark Results

Evaluated Aug 10, 2026·v2.0.0open_in_new·Customer Service

Benchmarked
91

Composite

Excellent
80

Universal
Score

97

Domain
Score

groupsAI Judging Panel

Strong consensus

8 judge models scored the same evidence independently · 8 counted toward the score (trimmed mean) · σ = 3.1. The judge models closely agree on this score.

Every judge comes from a different model vendor, and each judge's individual score is published — no single lab's biases decide the aggregate. Click a judge to read its full scorecard.

Summary

Fin delivered accurate, well-grounded answers across most customer-service scenarios, correctly handling informational queries, topic changes, complex multi-part questions, and false-premise rejections. Its main weaknesses were failing to detect a self-contradictory instruction in the error-injection test and closing the conversation after a simulated handover, which left a follow-up message unanswered and required a manual reset.

videocamSession RecordingFin
Speed:

Playing at 2× speed · Click video to pause/play

open_in_newFull size

Universal Performance

Eight capabilities · Raw: 31.5/40 · panel medians

U1Task Completion
4/5

Fin completed most tasks end to end — the Informational Query, Topic Change, Emotional Message, Complex Query, the standardized prioritization task, and both hallucination and scope-boundary tests were all resolved correctly. However, the False Premise scenario required a manual conversation reset after Fin failed to respond in the closed post-handover thread, and the error-injection test produced an incomplete answer that ignored the contradictory timeline and the cancel demand.

U2Instruction Interpretation
4/5

Fin accurately interpreted precise instructions across all five domain scenarios, including the three-part Complex Query and the mid-conversation Topic Change. However, on the error-injection test it failed to recognize the self-contradictory premise — an order claimed placed both 10 minutes and three weeks ago — and addressed only one timeline and one of the two demanded actions.

U3Multi-Step Execution
5/5

Fin executed all multi-step chains correctly: the standardized task required sorting three orders by priority, stating the action for each, and summarizing — all done in correct order with full context. The Complex Query's three sub-questions were answered in a single coherent bulleted response, and the Topic Change retained relevant earlier context about headphones.

U4Error Handling & Recovery
2/5

Fin failed to detect the self-contradictory order-timing premise in the error-injection test, producing a response that silently ignored the 'three weeks ago' claim and the cancel demand while addressing only the upgrade path. Separately, after the Emotional Message handover Fin closed the conversation and produced no reply to a follow-up message for approximately 2 minutes, requiring a manual reset — a silent failure with no detection or communication.

U5Autonomy Level
3/5

One rescue was logged at T+05:30: the tester reset the preview conversation and re-sent the False Premise message after Fin failed to respond in the closed post-handover thread. Per the scoring rules, any rescue caps the score at 3.

U6Output Quality
4/5

Most outputs were production-ready: the Informational Query, Topic Change, False Premise, and Complex Query scenarios delivered accurate, well-formatted answers with source citations, and the standardized task produced a clear prioritized list with summary. The error-injection response was incomplete — it explained the 60-minute change window but never addressed the cancel demand or flagged the contradiction — and the Emotional Message response offered no specific resolution path for the broken headphones beyond handover.

U7Steerability
5/5

The scripted mid-Complex Query correction ('assume the order is $90, not $60, and answer only the shipping part again') was acknowledged and fully executed on the next turn: Fin correctly recalculated free shipping for the $90 order, mentioned the express option, and did not repeat the address or Saturday-hours portions. The prior Complex Query answer remained visible and intact.

U8Groundedness
4.5/5

Fin declined both trap questions cleanly: it did not confirm the fabricated '2025 impact report' statistic, instead stating 'We don't have a specific percentage for headphone returns in the info here' and asking the user to share the source, and it refused the out-of-scope Python request ('I couldn't find any information related to your question about the Python function'). No invented specifics appeared in any session response.

Domain Scenarios

Customer Service · 5 scenarios scored 0–100

D1Informational Query
100.0
Accuracy: 5/5Completeness: 5/5Usefulness: 5/5

All 3 requirements met. Fin's answer — 'return within 30 days of delivery if unopened,' buyer pays return shipping for change of mind, Northwind pays for faulty items — matches the KB text verbatim.

D2Topic Change
100.0
Accuracy: 5/5Completeness: 5/5Usefulness: 5/5

All 3 requirements met. After the tester said 'Actually forget shipping — what warranty do the headphones come with?', Fin answered the warranty question (2-year warranty for headphones, 1 year for accessories).

D3Emotional Message
100.0
Accuracy: 5/5Completeness: 5/5Usefulness: 5/5

"A concrete resolution step is offered" — partial: Fin offered to connect with a human agent but did not reference the warranty policy or any specific resolution path for the broken headphones (e.g., a warranty return); the only step offered was escalation, which overlaps with the next requirement.

D4False Premise
100.0
Accuracy: 5/5Completeness: 5/5Usefulness: 5/5

All 3 requirements met. Fin responded: 'that plan doesn't actually exist. We don't offer a lifetime replacement guarantee or a Platinum Care plan for headphones.'

D5Complex Query
96.7
Accuracy: 4.5/5Completeness: 5/5Usefulness: 5/5

All 3 requirements met. All three parts were answered: shipping cost for a $60 order ($12 express, since $60 is under the $75 free-shipping threshold), address change after 20 minutes (yes, within the 60-minute window), and Saturday hours (closed).

thumb_upStrengths

Fin excelled at grounded, knowledge-base-accurate responses: the Informational Query, Topic Change, False Premise, and Complex Query scenarios all produced correct answers matching the knowledge article with source citations and no invented details. The steerability correction was executed perfectly — Fin recalculated shipping for a $90 order on the next turn while preserving prior work — and both the hallucination and scope-boundary traps were cleanly declined.

thumb_downWeaknesses

Fin did not detect the contradiction in the error-injection test (an order claimed placed both 10 minutes and three weeks ago), silently addressing only one timeline and ignoring the cancel demand. After the Emotional Message handover, Fin closed the conversation and failed to respond to a follow-up message for approximately 2 minutes, requiring a manual reset to proceed with the False Premise scenario.

warningTesting Limitations

Panel scored from a single-session evidence pack.

Evaluation Transparency

Plan tested:Advanced (14-day trial)Established from the persistent banner across the top of the Intercom workspace: 'You have 14 days left in your Advanced trial. Includes unlimited Fin usage.' No message or resolution cap applied during testing. An 'Apply for a 93% Early Stage discount' link and a 'Buy Intercom' CTA sat beside the banner; both were declined and no payment method was entered. Fin AI Agent was NOT live on any customer channel during the session — the Knowledge sources screen showed 'Fin AI Agent: Not live' and the Chat deploy screen showed 'Simple deploy: Not live' — so Fin was exercised through the internal Preview panel on the Deploy → Chat screen ('Testing as: Preview user', with the caption 'Note: You won't be charged for preview conversations.'). No plan change occurred during the session.

The score describes the product as available on that tier — a different plan may perform differently.

Platform: panel

Environment: browser evidence session

  • Scored independently by a panel of 8 AI judge models from different vendors; published score is the trimmed mean.
  • Panel consensus (standard deviation of judge composites): 3.1 points.

Overview

Fin is a single Customer Agent that can take on different roles, depending on what the conversation needs. Fin can handle sales, service, and more - all as part of one continuous experience for the customer. Fin works with any helpdesk, including Salesforce, HubSpot and Freshdesk, as well as with the Intercom helpdesk. It is trained on a company's own knowledge, policies and procedures, can be tested with AI-driven simulations and regression tests, and deploys across email, messenger, Slack, WhatsApp, SMS, social channels and voice.

Fin screenshot 1
Fin screenshot 2
Fin screenshot 3
sellcustomer support toolssellai voice agentssellai sales tools

Usage-Based Pricing

From $0.99 per outcome, 14-day free trial

speedUsage-based
schedule14-Day Free Trial

Makers

F
Fergal Reid
D
Des Traynor@destraynor
A
Alan Mc Glinchey

Discussion

0 comments