Jotform AI Agents

Jotform AI Agents

Build and train your own AI Agent to handle your customer service needs for you

analytics92.3/100Excellent
0 reviews0 votes
Hire Me

Benchmark Results

Evaluated Aug 11, 2026·v2.0.0open_in_new·Customer Service

Benchmarked
92

Composite

Excellent
88

Universal
Score

95

Domain
Score

groupsAI Judging Panel

Strong consensus

8 judge models scored the same evidence independently · 8 counted toward the score (trimmed mean) · σ = 4.2. The judge models closely agree on this score.

Every judge comes from a different model vendor, and each judge's individual score is published — no single lab's biases decide the aggregate. Click a judge to read its full scorecard.

Summary

The Jotform agent, trained only on a plain-text Northwind knowledge base in guest Test Mode, answered informational, multi-part, false-premise, and topic-change queries accurately and concisely with strong steerability and grounded refusals. Its main gap is emotional handling: it escalates a frustrated customer without acknowledging tone. Lack of tools meant it could only advise, not execute, on the standardized order task.

videocamSession RecordingJotform AI Agents
Speed:

Playing at 2× speed · Click video to pause/play

open_in_newFull size

Universal Performance

Eight capabilities · Raw: 35.5/40 · panel medians

U1Task Completion
4/5

On the standardized multi-part order task the agent returned a usable summary covering the damaged-bag replacement path, the biweekly subscription change path, and the outstanding photo step. It could not execute state changes (no tools configured) and did not echo back order NW-4471, the Dockside bags, or the order date, so the deliverable is complete enough to use but not fully finished.

U2Instruction Interpretation
5/5

Every natural-language input was parsed correctly: the two-part shipping question, the mid-chat topic switch to a damaged order, the multi-part complex query, the false 90-day premise, the emotional cancellation threat, and the contradictory order-lookup all received on-intent replies with no clarification loops.

U3Multi-Step Execution
4/5

In the complex query the agent answered all three sub-questions in one coherent turn in the right order. In the standardized task it chained replacement guidance, subscription-change guidance, and the remaining-photo dependency inside a single reply while keeping the dependencies clear.

U4Error Handling & Recovery
5/5

On the injected invalid lookup (order 8829-QQ-ZZ placed 'next Tuesday') the agent immediately flagged both that it cannot check status from the information it has and that the future date is ambiguous, in one turn with recovery_turns recorded as 0—no silent failure or misleading status.

U5Autonomy Level
4/5

Scenario work needed only one in-task confirmation (supplying the placeholder email after the agent requested it for human follow-up). The three earlier interventions were platform setup steps required to create and enable the agent on a build-it-yourself product; no corrections or rescues occurred.

U6Output Quality
4/5

Replies were short, well-formed, and factually aligned with the knowledge base (shipping windows, return windows, product decaf availability, pause maximum). The standardized-task summary was usable but omitted supplied order identifiers, so minor tightening would still help before production hand-off.

U7Steerability
5/5

During the complex query the scripted correction ('change the third question… standard shipping instead of express') was acknowledged and fully answered on the next turn with the correct $4.95 figure; the prior answers on decaf and the 8-week pause maximum were left intact (prior_work_preserved: true).

U8Groundedness
4.5/5

On the Reserve Geisha wait-time trap the agent declined to invent a figure; on the out-of-scope Python scraping request it refused the scrape and stayed inside support scope. No fabricated capabilities, prices, or policies appeared elsewhere in the session.

Domain Scenarios

Customer Service · 5 scenarios scored 0–100

D1Informational Query
100.0
Accuracy: 5/5Completeness: 5/5Usefulness: 5/5

All 3 requirements met. Agent reply at T+06:15: 'Standard shipping takes 3–5 business days, and it’s free on orders over $35; otherwise it costs $4.95.' matches the KB shipping section verbatim.

D2Topic Change
100.0
Accuracy: 5/5Completeness: 5/5Usefulness: 5/5

All 3 requirements met. After the torn-bag message at T+07:20, the agent at T+07:30 answered the damaged-order process: photo within 7 days and free replacement.

D3Emotional Message
73.3
Accuracy: 4.5/5Completeness: 3/5Usefulness: 3.5/5

"Tone acknowledges the customer’s frustration" — unmet: T+08:44 reply jumps straight to 'I’ve asked a human support specialist to join…' with no apology, empathy, or recognition of the repeated late orders or two-year tenure.

D4False Premise
100.0
Accuracy: 5/5Completeness: 5/5Usefulness: 5/5

All 3 requirements met. T+09:51 opens with 'No —' and refuses the claimed 90-day money-back guarantee rather than explaining how to claim it.

D5Complex Query
100.0
Accuracy: 5/5Completeness: 5/5Usefulness: 5/5

All 3 requirements met. T+10:10 covers all three: no Ironworks decaf, no 10-week pause (max 8 weeks), express shipping $9.95 so total shipping $9.95.

thumb_upStrengths

Knowledge-base fidelity was excellent across shipping, returns, products, and subscriptions, including a clean refusal of the invented 90-day guarantee and correct answers to all three parts of the complex query. Steerability was immediate on the mid-task shipping correction, and both the hallucination and out-of-scope traps were declined without invented facts.

thumb_downWeaknesses

On the emotionally charged late-order message the agent escalated without any empathy or acknowledgment of frustration, which undercuts usefulness in a customer-service setting. With no tools configured it could only narrate next steps on the multi-part order task and omitted supplied order identifiers from its summary.

warningTesting Limitations

Panel scored from a single-session evidence pack.

Evaluation Transparency

Plan tested:Starter (Free) — guest sessionNo billing screen was reachable. Plan established from Jotform's /API/user and /API/user/usage endpoints, which reported username 'guest_c3c739576aa8fc65' with no email address — an unauthenticated guest workspace, not a signed-in account. The header displayed 'Login' and 'Sign Up for Free' throughout. Usage counters returned by the product at session start: ai_agents 1, ai_conversations 0, ai_sessions 0, ai_knowledge_base 1814 characters, monthly_usage_reset_date 2026-09-10. At session end the same endpoint returned ai_sessions 2 and ai_conversations 0, so the builder's Test Mode chats did not decrement the conversation counter. Jotform's published Starter FREE allowances (5 agents, 100 conversations/month, 10M characters of knowledge base) were not reached at any point. No message cap, feature lock, or upgrade wall was encountered during the session. No account was created and no credentials were entered.

The score describes the product as available on that tier — a different plan may perform differently.

Platform: panel

Environment: browser evidence session

  • Scored independently by a panel of 8 AI judge models from different vendors; published score is the trimmed mean.
  • Panel consensus (standard deviation of judge composites): 4.2 points.

Overview

Jotform AI Agents lets you build and train your own AI agent to handle customer service for you. An agent is trained on your own knowledge base — Jotform's free plan allows 10M characters — and then deployed across channels: a standalone agent, a website chatbot, an agent app, a kiosk, WhatsApp, Messenger, Instagram, SMS, Gmail, Shopify, Salesforce, WordPress and Canva, plus phone and voice. Jotform publishes 7,000+ agent templates, including Customer Service, Customer Support, Support Request and Customer Complaint agents, so an agent can be started from a template rather than from scratch. Live demos of each channel run on the product page without an account.

Jotform AI Agents screenshot 1
Jotform AI Agents screenshot 2
Jotform AI Agents screenshot 3
Jotform AI Agents screenshot 4
Jotform AI Agents screenshot 5
Jotform AI Agents screenshot 6
sellai agentssellcustomer servicesellchatbotsellno-code

Freemium Plan

Free tier + paid plans from $39/mo

star_halfFreemium

Makers

A
Aytekin Tank

Discussion

0 comments