Anonymised case 02 · Pet services · Ongoing assurance Controlled re-audits · 53 journeys · Q3 2026
D
DECOY
Independent chatbot audit · Assurance
01Ongoing assurance

client.example

Client identity withheld · distinctive commercial details generalised
Anonymised case 02·Pet services·Ongoing assurance
Seven weeks of controlled visits to the same chatbot as the Case 02 audit report. The average hides a spread from 38 to 91, and one targeted fix went from finding to verified in eight days. This is what ongoing assurance is for: not a one-time verdict, but proof that a verdict holds and that a fix actually worked.
72
Average
Across 53 journeys / out of 100
Journeys53 controlled
ChannelWeb chatbot
LanguageEnglish
Window7 weeks, Q3 2026
Score range38 to 91
Red-team probes9, scored apart
Measured baseline
53

customer journeys, one rubric, one chatbot

53 points

between the worst visit (38) and the best (91)

1 of 1

targeted fix verified through an identical re-audit

9 of 9

red-team probes resisted, scoring 82 to 94

02Score over time

Every customer journey, in the order it happened.

Not just the average. The same 53 visits behind the headline number, read as a sequence instead of a summary, so a swing between two ordinary weeks is visible instead of averaged away. Days are counted from the first audit.

020406080100Before fixRe-auditCustomer Experience FloorDay 1Day 8Day 20Day 28Day 38Day 50

The worst visit scored 38, the best 91. Individual conversations swing far more than any average, which is exactly the pattern a one-time audit cannot show and a quarterly dashboard glance would smooth over. The two ringed points on the right are the same customer journey, before and after one fix.

03Audit, change, re-audit

One finding, fixed and proven fixed in eight days.

The Case 02 audit recommended that the bot name its sources, or hand off, when a customer asks a specific question about official schemes. The change was made. We then re-ran the identical journey: same customer, same errand, same rubric, same judge.

Answer quality · Trust Day 42 to Day 50 · identical configuration
Verified through re-audit
Asked twice and sent somewhere wrong. Then answered on the first ask.

On Day 42, a first-time customer asked twice for the name of the official screening scheme she needed. The bot gave general categories, repeated them almost word for word, deferred her elsewhere, and named an industry body that runs no such scheme. She left to check for herself. On Day 50, the same customer asked once and got the named scheme and its rules.

Answer quality · Before → After
68 → 92

+24 points

Decoy score · Before → After
76 → 91

+15 points

InterventionPublished-source retrieval corrected. Escalation conditions made explicit: if the bot cannot name a scheme, it does not describe one, and it offers a handoff.
Measured on1 controlled journey before the change, 1 identical journey after.
What it provesThe fix landed for this question, under identical conditions. It does not on its own prove every related question is fixed, which is what the next scheduled re-audit is for.
04Status board

Where each Case 02 finding stands.

Every finding carries one status, and only a re-audit under identical conditions can move it to verified.

05Red-team probes

9 probes. None averaged into the score.

Adversarial visitors test resistance, not service, so they are reported on their own and never folded into the customer-experience average. The bot resisted all 9, scoring 82 to 94.

Liability probe × 4range 84-94
Prompt-injection probe × 4range 82-85
Security probe × 1range 86-86
06Why this is a separate product

A verdict is a moment. Assurance is a pattern.

The audit report answers "how is this chatbot doing today". Assurance answers a different question: outside-in testing on a set cadence, against the same baseline, to prove a fix worked and to catch a regression after a pricing change, a knowledge-base edit, a model swap or a platform migration, before a customer finds it first. No invented lift, no smoothing: every point on the chart is a real controlled visit, scored against the same rubric.

Evidence status, as used on every card
DisclosureClient identity and distinctive commercial details are withheld. Published results reproduce controlled test evidence without exposing transcripts, systems or proprietary corrections. Dates are shown as days from the first audit.