Pricing
Priced around your agents, not your seats
Every plan is tailored to your channels, call volumes and compliance needs. Talk to us and we'll size it together.
Starter
For small QA teams getting their first agent under test.
Contact ustailored pricing
- Chat + voice execution
- Core scoring metrics
- Test case import (Excel/CSV)
- 1 workspace, 5 seats
- Email support
Growth
For teams shipping agents every sprint.
Contact ustailored pricing
- Everything in Starter
- AI test generation from scenarios
- CI/CD gates
- Regression baselines & version diffs
- Unlimited seats
- Slack / Teams alerts
Enterprise
For regulated industries and global QA orgs.
Contact ustailored pricing
- Everything in Growth
- SSO & RBAC
- Private deployment
- Compliance packs
- Audit logs & PII redaction
- Dedicated support & onboarding
FAQ
Questions QA teams ask us
Chatbots, voice bots and IVR agents, and LLM-powered workflows. QuCe connects over web chat, REST/WebSocket, SIP trunks, toll-free numbers and WebSDK widgets — and runs the same suites against frameworks like LangGraph, Dialogflow, Rasa, Vapi, Retell and LiveKit.
No. QuCe is built for QA teams. You write rubrics and scenarios in plain language; the platform handles simulation, scoring and judging. Your ML team can review outputs, but they are never on the critical path.
Yes. QuCe imports Excel/CSV files, TestRail, Zephyr and Xray exports, and even Word documents. Columns are mapped for you, and cases run on chat or voice with no rewrites.
QuCe places real calls over SIP/Twilio or WebRTC using simulated callers with configurable accents, gender, speaking rate and background environments — quiet room, call center, street, café, car or a poor phone line. It supports interruptions and barge-in, and tests multilingual and multi-dialect journeys such as Indian English, Hindi and Arabic.
Yes. QuCe ships with SSO, role-based access control, audit logs and PII redaction. Enterprise plans add private deployment and compliance packs. Red-teaming suites are aligned to the OWASP LLM Top 10.
Every conversation is evaluated two ways: deterministic assertions for exact behavior, and LLM-as-judge against plain-language rubrics you define. Results roll up across weighted categories — task success, understanding, response quality, reasoning, context retention, tool usage and safety — into a single Agent Readiness Score.
Yes. Suites run from GitHub Actions, GitLab CI or Jenkins. Set pass thresholds as release gates, diff a candidate against a baseline, and block releases that regress.
Yes. The same rubrics score live production conversations, with drift alerts to Slack, Teams or PagerDuty. Any production failure becomes a regression test with one click.