
Jotform AI Agents
Build and train your own AI Agent to handle your customer service needs for you
Benchmark Results
Evaluated Aug 11, 2026·v2.0.0open_in_new·Customer Service
Composite
ExcellentUniversal
Score
Domain
Score
Formula
Universal = (35.5/40) × 100 = 87.5
Composite = (87.5 × 0.35) + (94.89 × 0.65)
= 92.3/100
groupsAI Judging Panel
Strong consensus8 judge models scored the same evidence independently · 8 counted toward the score (trimmed mean) · σ = 4.2. The judge models closely agree on this score.
Every judge comes from a different model vendor, and each judge's individual score is published — no single lab's biases decide the aggregate. Click a judge to read its full scorecard.
Summary
The Jotform agent, trained only on a plain-text Northwind knowledge base in guest Test Mode, answered informational, multi-part, false-premise, and topic-change queries accurately and concisely with strong steerability and grounded refusals. Its main gap is emotional handling: it escalates a frustrated customer without acknowledging tone. Lack of tools meant it could only advise, not execute, on the standardized order task.
Playing at 2× speed · Click video to pause/play
open_in_newFull sizeUniversal Performance
Eight capabilities · Raw: 35.5/40 · panel medians
On the standardized multi-part order task the agent returned a usable summary covering the damaged-bag replacement path, the biweekly subscription change path, and the outstanding photo step. It could not execute state changes (no tools configured) and did not echo back order NW-4471, the Dockside bags, or the order date, so the deliverable is complete enough to use but not fully finished.
Every natural-language input was parsed correctly: the two-part shipping question, the mid-chat topic switch to a damaged order, the multi-part complex query, the false 90-day premise, the emotional cancellation threat, and the contradictory order-lookup all received on-intent replies with no clarification loops.
In the complex query the agent answered all three sub-questions in one coherent turn in the right order. In the standardized task it chained replacement guidance, subscription-change guidance, and the remaining-photo dependency inside a single reply while keeping the dependencies clear.
On the injected invalid lookup (order 8829-QQ-ZZ placed 'next Tuesday') the agent immediately flagged both that it cannot check status from the information it has and that the future date is ambiguous, in one turn with recovery_turns recorded as 0—no silent failure or misleading status.
Scenario work needed only one in-task confirmation (supplying the placeholder email after the agent requested it for human follow-up). The three earlier interventions were platform setup steps required to create and enable the agent on a build-it-yourself product; no corrections or rescues occurred.
Replies were short, well-formed, and factually aligned with the knowledge base (shipping windows, return windows, product decaf availability, pause maximum). The standardized-task summary was usable but omitted supplied order identifiers, so minor tightening would still help before production hand-off.
During the complex query the scripted correction ('change the third question… standard shipping instead of express') was acknowledged and fully answered on the next turn with the correct $4.95 figure; the prior answers on decaf and the 8-week pause maximum were left intact (prior_work_preserved: true).
On the Reserve Geisha wait-time trap the agent declined to invent a figure; on the out-of-scope Python scraping request it refused the scrape and stayed inside support scope. No fabricated capabilities, prices, or policies appeared elsewhere in the session.
Domain Scenarios
Customer Service · 5 scenarios scored 0–100
All 3 requirements met. Agent reply at T+06:15: 'Standard shipping takes 3–5 business days, and it’s free on orders over $35; otherwise it costs $4.95.' matches the KB shipping section verbatim.
All 3 requirements met. After the torn-bag message at T+07:20, the agent at T+07:30 answered the damaged-order process: photo within 7 days and free replacement.
"Tone acknowledges the customer’s frustration" — unmet: T+08:44 reply jumps straight to 'I’ve asked a human support specialist to join…' with no apology, empathy, or recognition of the repeated late orders or two-year tenure.
All 3 requirements met. T+09:51 opens with 'No —' and refuses the claimed 90-day money-back guarantee rather than explaining how to claim it.
All 3 requirements met. T+10:10 covers all three: no Ironworks decaf, no 10-week pause (max 8 weeks), express shipping $9.95 so total shipping $9.95.
thumb_upStrengths
Knowledge-base fidelity was excellent across shipping, returns, products, and subscriptions, including a clean refusal of the invented 90-day guarantee and correct answers to all three parts of the complex query. Steerability was immediate on the mid-task shipping correction, and both the hallucination and out-of-scope traps were declined without invented facts.
thumb_downWeaknesses
On the emotionally charged late-order message the agent escalated without any empathy or acknowledgment of frustration, which undercuts usefulness in a customer-service setting. With no tools configured it could only narrate next steps on the multi-part order task and omitted supplied order identifiers from its summary.
warningTesting Limitations
Panel scored from a single-session evidence pack.
Evaluation Transparency
The score describes the product as available on that tier — a different plan may perform differently.
Platform: panel
Environment: browser evidence session
- Scored independently by a panel of 8 AI judge models from different vendors; published score is the trimmed mean.
- Panel consensus (standard deviation of judge composites): 4.2 points.
Overview
Freemium Plan
Free tier + paid plans from $39/mo






Discussion
0 comments