An abstract shape in which several branching lines enter a testing space, some stop at gaps in the middle, and only one smooth path continues forward.
AI and LLMsthrough Working Back From Failure

Catch customer-service automation failures with simulated customers first

A service that runs failure scenarios with simulated customers before launch could be a smaller, more specific business than building the chat-support system itself.

Published 2026. 10. 5.

A service that runs failure scenarios with simulated customers before launch could be a smaller, more specific business than building the chat-support system itself.

More than 16,000 tests before showing it to customers

Researchers at Brazilian financial company Nubank had an AI customer-support assistant talk to simulated customers before putting it in front of real customers. They tested workflows such as checking card-delivery status, confirming an address, and resending a card. The conversations continued with test lookup results rather than results from the live operating system.

The researchers changed response approaches and settings, then ran more than 16,000 simulated conversations. They applied the selected approach across 8,400 real customer conversations and said the share of conversations completed without a request for a human agent rose by 8.82 percentage points. If the original rate had been 50%, that would mean 58.82%, but the actual before-and-after rates were not disclosed.

A score asking whether customers would recommend the support experience fell by 1.21 points, which the researchers judged not to be a meaningful change. More conversations avoided handoff to a person, but this does not show that customer satisfaction improved. The result is also a self-reported finding from Nubank and partner researchers that has not completed peer review.

At first, Nubank let the AI carry out three steps on its own: checking delivery status, validating an address, and resending a card. When it changed the order or skipped an intermediate step, the company grouped the three procedures into one fixed process. The AI could still understand the language, but people defined the order for a workflow where mistakes carry a high cost.

The weaknesses of automated support are also clear in Korea. In a Seoul Metropolitan Government survey of 1,000 online shoppers, 39.4% named irrelevant, standardized answers as their biggest frustration, while 4.1% said they preferred AI chat support. That is why finding wrong answers and blocked moments can matter more than adding more automated replies.

Where the day changes for a 12-person online store

Consider the customer-support lead at a custom-furniture online store with 12 employees. To handle one inquiry, the team moves between the store's order screen, a courier tracking screen, and an exchange-policy document, copying an order number and delivery status two or three times. Exceptional details received by phone are recorded again in a shared document.

The problem is that questions do not arrive in policy-shaped form. “The delivery person arrived, but the elevator is broken,” “Only the tabletop is scratched—do I need to return the whole item?” and “Please cancel an order my parents placed” each combine delivery, cost, and identity verification in one sentence. A new support assistant may answer an ordinary delivery-status request but promise a refund or delay handoff to a person in these boundary cases.

Simulated-customer testing starts here. Using redacted past conversations and current policies, it creates customers with frequent typos, customers who repeat the same question, customers who refuse identity checks, and angry customers. It also varies order lookup results beyond normal delivery, including address mismatches, suspected loss, and an unconfirmed installation date.

The scoring criteria should not stop at whether an answer sounds natural. They should check whether the current policy was cited, whether an uncertain cost was promised, whether required information was omitted, and whether the conversation was handed to a person after two failures. Results become distorted if a customer abandons the conversation but “no agent handoff” is counted as a success.

The work changes on the day an exchange policy changes. Previously, a staff member might have tried a few questions manually before updating the support assistant. Now, the team reruns all saved delivery, exchange, and complaint scenarios together. A person can review only the cases where an answer that was correct yesterday now conflicts with the new policy.

Some work still belongs to people. Staff need to judge product defects, approve exceptional compensation, and persuade customers disputing responsibility. Passing a test should not immediately grant authority to change real orders. Start with lookup and guidance, then apply the assistant to a limited group of customers.

This approach can fail if simulated customers are much more cooperative than real ones, if the share of conversations not handed to people is treated as the only success measure, or if test order statuses differ from the live operating screen. An independent study found that the assessed support-success rate could vary by as much as 9 percentage points depending on how simulated customers were constructed, and that some speaking styles and groups were not represented well. Simulated testing is therefore a filter for candidates, not a replacement for tests with real customers.

Support platforms abroad are beginning to offer testing as a separate capability

Intercom, which operates from the United States and Ireland, offers a “Fin Bulk Testing” feature for customer-support teams. A company can provide questions it has actually received from customers, then review in bulk how its AI support assistant would answer and which documents it used as a basis. Test responses are not charged for; the company charges based on results from handling issues in live operation.

Instead of testing a new support assistant from a blank screen, the service starts with questions that have already arrived. It can now check answers in each language, knowledge-document selection, and whether predefined automated actions work before launch. But companies must separately add new fraud methods or rare exceptions that do not appear in past questions.

Salesforce in the United States offers the “Agentforce Testing Center,” where companies can upload test scenarios or automatically create test questions from business procedures. It evaluates whether the support assistant found the right information, performed required actions, followed instructions, and how long it took to respond. Salesforce advises running this in a separate test environment so real customers and order data are not changed.

Salesforce also said it set limited rules and tested its own customer-support site for about two months before releasing its support assistant. The company says it later recorded a 76% rate of handling without human intervention across more than 1.7 million conversations, but this figure has not been independently verified. The important shift is that support-automation providers are beginning to treat pre-launch testing and scoring as a separate capability, not just an answer-generation feature.

Four small ways to start here

1. A simulated-customer pack for exchanges and returns

  • A service that takes a store's policies and past conversations, then automatically creates scenarios for exchanges, partial refunds, and delivery delays.
  • It is for a side-dish shop newly starting delivery sales or a household-goods business that has opened its own online store.
  • Support automation is growing, but small sellers often lack someone who can write test questions themselves.
  • The first screen should have industry selection, policy-document upload, and a “Create 20 simulated inquiries” button.

2. Retest alerts after a policy change

  • A service that reruns existing answers when refund criteria or operating hours change and identifies only the answers that have gone wrong.
  • It is for small academies with around five employees, where make-up classes, refunds, and vehicle-operation rules change often.
  • A support assistant can keep giving outdated answers after a change, creating a bigger problem than the initial setup.
  • The first screen should place the old and new policies side by side, with a button to show only changed answers.

3. A handoff-moment checker

  • A service that finds the moments when a support assistant should stop answering and hand off to a staff member, then tests whether the handoff actually happens.
  • It is for appointment coordinators at neighborhood clinics handling appointment changes, treatment-cost questions, and requests from family representatives.
  • For sensitive inquiries, stopping at the right time can matter more than answering well.
  • The first screen should offer three tests: “immediate handoff,” “handoff after two failures,” and “intake outside operating hours.”

4. A simulated-customer practice room for new support staff

  • A service where AI plays difficult customers and reviews a new staff member's explanations, clarification questions, and policy compliance.
  • It is for travel-booking businesses or local-event registration agencies that hire many short-term support staff during busy seasons.
  • It enables repeated practice and decisions about readiness for live work without using real customers as training material.
  • The first screen should contain only one scenario for today's practice, three passing criteria, and a button to start the conversation.

Why this matters where you are

You can check whether your own support workflow has saved conversations, policy documents, and safe test data that can be used to recreate failure cases. The policies, customer expectations, and sensitive situations will differ by market and sector. Start by testing the moments where an automated assistant must follow a fixed process or hand work to a person.

What to check today

Choose 20 recent support cases that were handed to a person or led a customer to contact you again, and write down only the reason they failed. Remove personal information and classify each case in one line, such as “ambiguous policy,” “lookup failure,” or “delayed handoff to a person.” This takes 30 minutes. If the same reason appears five or more times, that one scenario may be worth turning into a simulated-customer testing service.

Sources

6 sources

Every fact in this article came from the pages below. Check them yourself.

Catch customer-service automation failures with simulated customers first | Prometheon