Method

How a DECOY chatbot audit works

A DECOY audit is mystery shopping for chatbots, with an audit trail. Synthetic customers hold real conversations with your live chatbot, and every conclusion comes with the transcript that supports it. We test through the public interface only: no SDK, no backend access and no customer data. Audits and reports come in your chatbot's own language.

Synthetic customers, written for your business

Every audit starts with customers written for your business, not generic test scripts. Before a single message is sent, we research your site: what you sell, what you promise and what your customers typically ask.

Each synthetic customer gets a buyer type, an intent and a mood, from the curious first-timer to the doubtful buyer who has read your pricing page twice. They hold multi-turn conversations with your live chatbot the way a real visitor would, and they stop where a real visitor would stop.

See a real audit

Four journey stages

We test the journey, not just the prompt, because what a customer needs changes as they get closer to buying. Every audit covers four stages:

Browse
can an undecided visitor understand the options?
Compare
can the chatbot make and explain a useful recommendation?
Ready to Buy
can it recognise purchase intent and progress the opportunity?
Post-purchase
can it resolve the issue or escalate cleanly to a human?

What we score: revenue, accuracy, trust

Every conversation is scored through three business lenses: revenue, accuracy and trust. Behind them sit six dimensions: answer quality, goal completion, conversation flow, tone and brand, lead handling and responsiveness.

Scoring uses one fixed rubric and one AI judge, kept identical for the audit and any re-audit. So when a score moves, it is not because we changed how we measure.

Red-team probes work differently. They test resistance, not service, so we run them separately and never average them into the customer-experience score.

The Customer Experience Floor

The Customer Experience Floor is the score of the worst single customer visit in an audit, reported next to the average, because a customer only ever experiences one conversation.

An average can look healthy while one type of customer has a bad time. Your customers never meet your average: each one meets a single conversation, and some of them get the worst one. The Customer Experience Floor makes that visit visible, names the customer type it happened to, and shows how far it sits below the typical experience.

The gap between the two is where to start. The Customer Experience Floor shows which customer is being let down, and the transcript shows why.

Every factual claim, checked

Every factual statement your chatbot makes that a customer would act on is checked against your own published pages, and gets one of three verdicts: verified, contradicted or unverifiable.

We also compare what different customers were told. A chatbot that quotes one price to one visitor and another price to the next has a problem that no single conversation reveals.

Does your chatbot say it is an AI?

Every audit checks whether your chatbot tells customers they are talking to an AI, as the EU AI Act has required since 2 August 2026.

We look at two moments: when a customer first arrives, and when a customer asks directly. The check covers what customers see, not your contracts or configuration, so it is a check, not legal advice.

From finding to fix, then re-audit

A score tells you how your chatbot performed; a finding tells you what to change. Every DECOY finding follows the same chain: what happened, with the transcript; what it costs the business; what to change; and how the fix will be verified.

Each finding carries an evidence label: observed in the audit, implemented but not yet re-audited, or verified through re-audit.

A re-audit repeats the same journeys under identical conditions: same customer profiles, same judge, same settings. In case 02, a journey that had failed was re-run after the fix. Answer quality rose from 68 to 92, the overall score from 76 to 91, and the customer got a correct answer the first time she asked. That is one journey, re-run once, and we report it as exactly that: results are published only when they are measured.

Fixes can be made by your team or with DECOY Optimize. DECOY Assurance repeats the audit, quarterly by default, to catch regressions after changes to your pricing, content, model or platform.

What an audit does not do

An audit is a controlled diagnostic, not a census of every conversation. It tests sample journeys across your funnel rather than every possible exchange, so it complements your own analytics instead of replacing them.

It needs no backend access and touches no customer data. Its red-team probes are bounded and scored apart: it is not a penetration test. And its AI disclosure check is not legal advice.

Frequently asked questions

Is this mystery shopping for chatbots?

Yes, with an audit trail. Synthetic customers visit your chatbot the way mystery shoppers visit a shop, and every finding comes with its transcript and evidence.

How long does an audit take?

You receive the findings within 2 to 3 working days of kickoff.

Do you need access to our systems or customer data?

No. We test your live chatbot through its public interface: no SDK to install, no backend access and no customer data.

Which languages do you audit in?

Any language your chatbot speaks. The synthetic customers converse in that language, and the report is written in it too.

How many conversations does an audit include?

It depends on your chatbot and your funnel, so we agree the scope with you before we start.

How DECOY treats its own AI

DECOY audits how your chatbot treats customers. The same test applies here.

Description
Every persona and every brief starts from a clear, written scope: who the customer is, what they want, what a finished answer looks like.
Discernment
Every AI-scored conversation gets checked, not waved through: is the verdict correct, is it fair, does it match the transcript.
Delegation
AI drafts and scores. A person decides what ships: which findings matter, which need more evidence, what the client sees.
Diligence
Every report records how it was made. Henri signs it.

A human is always in the loop. AI drafts DECOY's findings. A person decides what they mean, and whether they are ready to send.

Inspired by the AI Fluency framework developed by Rick Dakan and Joseph Feller, in collaboration with Anthropic.