<- Blog

Revenue Agent vs. Deflection Bot: How to Tell Them Apart

Revenue Agent vs. Deflection Bot: How to Tell Them Apart

Holden Lewis

Share

Five tests for evaluating a voice AI vendor

“Revenue agent” and “deflection bot” are useful sales shorthand. They are poor product categories. The same platform may handle either role; the workflow, permissions, and scorecard determine what it actually does. Buyers should ask whether the agent can finish valuable work safely in their systems and whether a production pilot proves the economics.

Customer-care priorities have already widened beyond cost reduction. In McKinsey’s 2024 survey of more than 340 customer-care leaders, one-third named revenue generation as a priority, while 11% named reducing contact volume. An evaluation centered on containment can miss the work closest to growth.

1. What does the scorecard reward?

A vendor’s default dashboard shows what its product was designed to optimize. Extra charts prove little. Ask which outcomes the vendor can calculate from your order, payment, contact-center, and CRM data, and which events sit behind each metric.

For service work, useful measures include first-contact resolution, repeat-contact rate, customer satisfaction, and cost per completed task. For sales work, use completed orders from eligible calls, average order value, gross margin, upsell acceptance, and revenue recovered from callers who would otherwise abandon.

Containment still matters because it shows how often the agent finishes without a person. On its own, it says little about customer value or additional revenue. A contained call may lead to a repeat contact, refund, cancellation, or lost sale. Track those outcomes for long enough to see whether the original call was truly successful.

Choose one number that will decide the pilot before it starts. Then name the warning signs that could outweigh a gain in that number. A sales pilot might optimize profit per eligible call while limiting increases in refunds, repeat contacts, transfers, complaints, and abandonment. Define an “eligible call” in plain terms so nobody can change the denominator after seeing the results.

Revenue handled by the agent is different from revenue the agent added. To measure the difference, keep some comparable callers on the current call path. The cleanest setup assigns eligible callers at random and keeps each caller on the same path throughout the test. That split gives both groups a similar mix of customers and demand; NIST explains why random assignment produces more dependable comparisons. If caller-level assignment is impractical, split comparable time blocks or customer groups and document the compromises.

Build the business case on profit from retained orders. The calculation can stay simple:

Eligible calls × the difference in retained-order rate × average profit per retained order, less the added platform, telephony, implementation, and service costs.

Finance should be able to reproduce every term from company data. The vendor’s dashboard can help explain performance, but it should never be the only place where the result exists. A broader voice AI scorecard should connect containment to dollars before the pilot starts.

2. What work can it complete in your systems?

An integration logo proves that two systems can connect. Revenue work requires the agent to use catalog, inventory, pricing, promotions, cart, payment, order-management, and customer-history data. It also has to make changes safely when a call goes off the happy path.

Give each vendor scenarios drawn from real calls: a clean purchase; an invalid promotion or unavailable item; and a service call that develops into a sale. Add a failure case, such as a payment timeout after the caller confirms an order.

Require a live sandbox demonstration and inspect the records afterward. Check whether the agent:

  • verifies price, inventory, address, and payment status before confirming the order;

  • prevents a retry from creating a duplicate charge or order;

  • has access only to the systems and actions needed for the task;

  • records what it did and when;

  • handles timeouts, partial failures, and unclear caller responses safely; and

  • gives the human agent the transcript, information already collected, and exact point of failure.

These checks show whether the agent can finish the job cleanly. A fluent conversation or successful API call covers only one part of the transaction.

In Simple’s Omaha Steaks deployment, a narrow sales flow expanded to roughly 100 searchable products and wrote completed orders into the company’s homegrown order system. That supports the claim that the agent can execute the workflow. A separate comparison is needed to show how many additional orders it produced.

3. How does it perform when a caller is ready to act?

A production pilot should test the complete sales motion: recognize what the caller wants, make an appropriate offer, complete the action, and preserve the customer experience.

Agree on the test rules before routing traffic. Decide which callers qualify, how calls will be split, which number determines success, which warning signs can stop the test, how long later outcomes will be tracked, and how much improvement is worth buying. Keep the offer mix, season, and caller population comparable between the AI and current paths.

Show the total number of calls and the absolute lift alongside percentage rates. Continue the test until there is enough traffic to distinguish a repeatable improvement from normal week-to-week variation.

For a sales deployment, track completed-order rate, profit per eligible call, upsell acceptance, abandonment, transfers, repeat contacts, cancellations, refunds, complaints, and payment failures. Read the measures together. A system that increases order count through heavier discounts may reduce profit. Higher containment can also create more work later if repeat calls rise.

Omaha Steaks provides a useful starting point. Its temporary holiday agents attached upgrades on about 22% of eligible calls, while Simple reports that its deployed AI runs at 28% to 30%. The comparison covers one retail workflow, and the source is a vendor case study. The published account does not provide the traffic count or enough detail to show whether the callers, offers, and time periods were directly comparable. Buyers should rerun the comparison on their own traffic under fixed rules.

Abandonment needs the same discipline. Omaha Steaks reports a decline from 16% two years ago to 3% today. The improvement matters, although staffing, routing, demand, or other changes during those two years may explain part of it. Count how many additional callers complete profitable orders, then compare that result with callers kept on the existing path.

4. How does the contract count success?

Pricing model is a weak product category. Per-resolution, usage, capacity, and outcome-based contracts can all support service or sales. Buyers need to understand what each model rewards and whether they can check the invoice against their own data.

Stripe’s guide to outcome-based pricing lays out the essentials: both sides need one definition of success, a shared data source, clear rules for shared credit and edge cases, and a regular way to compare counts. Before comparing rates, define:

  • the exact event that triggers a charge and the system that records it;

  • how long after the call a transfer, reopened issue, repeat contact, cancellation, refund, or partial completion can change the count;

  • who receives credit when a customer starts online, changes channels, or finishes with a person;

  • what happens during fraud, vendor outages, customer-system outages, and test traffic;

  • how long records are kept and how disputes, audits, and credits work; and

  • how minimums, caps, overages, and peak-season volume affect the bill.

Record operational success and billing status separately. A useful call may fall outside the contract’s billable definition, and a charged event may later be reversed. Keeping both records makes disagreements easier to resolve.

Recreate the invoice using the previous 12 months of traffic, including the busiest weeks. Run low, expected, and high cases. This exercise exposes vague billing units, expensive minimums, and bad peak assumptions before the contract reaches production.

5. What can reference customers prove?

Case studies show a vendor under favorable conditions. A reference becomes useful when its operation resembles yours. Look for similar call types, transaction complexity, systems, regulation, and seasonality.

For each performance claim, ask for:

  • the starting point, what the AI was compared with, and how traffic was divided;

  • the raw counts behind the rate, total traffic, and measurement dates;

  • which calls were included or excluded and how long later outcomes were tracked;

  • the company system that supplied the result;

  • changes in staffing, routing, offers, pricing, or marketing during the same period; and

  • performance after launch, including the weakest month.

Speak with the operator who owns the metric. Ask finance or analytics to confirm where the numbers came from. A conversion or recovered-revenue claim is impossible to judge when the vendor cannot show who was counted, what they were compared with, or how credit was assigned.

The reference call should end with a one-sentence pass/fail test for your pilot. For example: “For callers eligible to place a catalog order, the AI must raise profit per call by at least 8% while keeping any increase in refunds and seven-day repeat contacts below one percentage point, compared with similar callers on our current path over six weeks.” If the vendor cannot turn its strongest case study into a test like this, the case study is positioning.

Choose for the work, then prove the value

A support-heavy line may benefit most from faster resolution at lower cost. Measure repeat contacts and customer satisfaction alongside containment so failures cannot disappear into later calls.

A sales-heavy line should test completed transactions, recovered demand, and profit against the existing path. Mixed lines need separate rules and scorecards by call type; a blended average can hide gains in one area and losses in another.

Choose the vendor that can complete the priority task in the system of record and show the added value under counting rules your team can audit. Keep the category label as shorthand. Let the pilot determine what the product is worth. The broader voice AI buyer’s guide covers the security, latency, implementation, and governance checks that sit beside these five commercial tests.

Put your best rep on every call.