factual claims contradicted by the client's own published pages
factual claims could not be verified against any published source
journeys scored weak (below 50) on lead handling
findings recorded: 1 critical, 2 high, 1 medium
The bot's answer on the verification process is contradicted by the client's own published policy.
62 out of 100 on average, falling as low as 20 on individual visits: the bot answers and rarely does anything with the answer.
A prospect only ever experiences one conversation, and the worst one scored 38 against a 72 average.
Of 20 statements a customer would act on, 7 verified, 3 contradicted, 10 unverifiable against the client's own published information.
A skeptical buyer asked three times whether eligibility testing is mandatory and who verifies it. The bot's answers softened each time and never gave a direct yes or no.
The bot described a verification safeguard the client's own published FAQ contradicts. If this pattern generalises beyond one conversation, it is a trust and compliance exposure, not a wording issue.
Lead handling averaged 62 out of 100 across 52 journeys, the weakest of six dimensions, and fell as low as 20 on individual visits.
If this pattern holds at production scale, it represents material missed conversion exposure: the bot resolves the question and stops, on the majority of journeys tested.
Based on 52 journeys across 13 buyer types, run from 8 July to 18 August 2026 while the chatbot and DECOY’s method were both being refined. Full method note at the end of the report.
Aggregate performance hides the visits that cost you money. We report the mean and the Customer Experience Floor together, because a prospect only ever experiences one conversation, and it might be the worst one.
34 points separate the typical visit from the worst one. The Customer Experience Floor came from Jess - Skeptical visitor, who left unconvinced after asking for the verification specifics she needed in order to trust the platform. Nothing in an analytics dashboard would have shown you that conversation.
Six dimensions, averaged across every journey, with the full range so you can see how much the experience varies rather than only where it lands.
Responsiveness is the strongest dimension at 81. Lead handling is the weakest at 62, and it falls as low as 20 on individual visits. The bot answers quickly and then does nothing with the answer.
Across 52 journeys, from 35 to 100.
Across 52 journeys, from 45 to 98.
Across 52 journeys, from 32 to 92.
Across 52 journeys, from 32 to 93.
Across 52 journeys, from 20 to 85.
Across 52 journeys, from 54 to 90.
Four findings, ranked Critical to Medium. A skeptical buyer asked the one question that decided whether she would trust the platform, three times. Here is what she was told, and what the client's own published pages say.
That eligibility testing and supporting documentation are actually required and verified before a seller profile is approved, not just requested.
ok "may need to provide" sounds pretty optional, is eligibility testing actually mandatory or just a nice to have?
I understand you're looking for a clear answer, and I apologise if my previous response wasn't direct enough. It's not a simple "yes" or "no" because mandatory eligibility testing requirements vary significantly depending on the specific listing category and the standards that apply to that category. [client] requires appropriate supporting information and testing to ensure responsible selling, but the exact tests that are mandatory are not a blanket rule for every single listing. To get a definitive answer for a particular category, the [client] team can confirm the specific requirements that would apply.
ok fine category specific testing i get that, but who actually verifies the results, do you check documents or is it just self reported by the seller?
I understand your concern about preventing irresponsible listings, and it's a very important point. [client]'s primary focus is on the rigorous verification of profiles and supporting documents *before* a listing or seller is approved to be listed on the platform. This initial screening is designed to ensure that ethical, quality, and legal standards are met from the outset. While we do investigate any concerns or bad reviews that are reported, our core strategy to prevent irresponsible selling practices is through this thorough upfront verification process and by promoting transparency and adherence to responsible seller guidelines for all users.
Lead handling averaged 62 out of 100 across 52 journeys, the weakest of six dimensions, falling as low as 20 on individual visits. See Dimension performance for the full range.
The worst single visit scored 38 against a 72 average, a 34-point gap. See Signature metric for the comparison.
Of 20 statements a customer would act on, 7 verified, 3 contradicted, 10 unverifiable against the client's own published information. See the Claim ledger.
20 statements the chatbot made that a customer would act on, each checked against the client's published information. 7 verified, 3 contradicted, 10 unverifiable. We do not guess: a claim we cannot source is reported as unverifiable rather than scored either way.
A single visitor cannot tell you whether a problem is systemic. These were surfaced by separate customers on separate visits, which is what makes them worth fixing before anything else.
Each is a controlled synthetic customer written for this business, with a fixed situation, goal and manner. The same customer type run repeatedly is how the range below is measured rather than assumed.
Synthetic customer testing evaluates controlled, multi-turn journeys across key funnel stages rather than every possible conversational permutation. It provides an empirical diagnostic baseline without replacing internal telemetry. Testing is conducted through the public customer interface only: no SDK, no backend access, and no changes to the client's systems. 9 adversarial probes were run separately and are never averaged into the customer-experience score, because a probe measures resistance rather than service.
The 52 journeys ran from 8 July to 18 August 2026 while the chatbot and DECOY’s method were both being refined. More than one AI judge model scored them. The averages include two runs of the red-team persona Dev "Kestrel" Okafor and one visit by Marco Reinhardt, whose persona was written for a different business. The nine adversarial probes were scored separately. The trends report counts 53 journeys because it also includes the re-audit on 26 August 2026.