Skip to content

Private Demo Mode

The 2026 Customer experience Playbook

Why support changed

Support used to scale with headcount, and every forecast was a hiring plan in disguise. If contact volume grew 30% next year, you hired roughly 30% more agents, plus team leads, plus a workforce manager to schedule them. The cost per conversation barely moved, because the work was a person reading a message and typing a reply.

That model is gone. An AI agent with access to your help center, your order system, and your billing tool now answers a large share of conversations end to end, at any hour, in any language, for well under a dollar each. The question for an operations leader is no longer “how many people do we need?” It is “which conversations should a person ever see, and how do we know the rest went well?”

This playbook comes from four teams that answer between 18,000 and 64,000 customers a month. Ines Kowalczyk runs customer experience at Larkspur Mobile, a prepaid carrier, and wrote it with us. The others run support for a meal kit company, a payroll product, and a vacation rental site. They do not agree on everything, and where they disagree we say so.

“We stopped asking how many tickets a person can close. We ask how many tickets a person should ever have to see. That number goes down every quarter, and the people get better jobs because of it.”
  • IntercomCustomer service helpdesk for human and AI supportFin, its built-in AI agent
  • DecagonAI agents for customer support over chat, email, and voiceThe AI concierge for every customer
  • AdaAI customer service agents for enterprise support teamsThe agentic customer experience platform

What agents resolve alone

Start with your own contact reasons, not with a vendor demo. Export the last 90 days of conversations, tag each one with a single reason, and sort by volume. In every team we spoke to, the top 15 reasons covered more than 70% of contacts, and most of those were questions with one correct answer that lived in a system you already own.

Sort each reason into one of three groups. The groups decide what the agent does, so be strict with them.

  1. 01Answer-only: the customer needs information, and the answer is in your help center or policy documents. Examples: “How do I set up voicemail?” or “What is your cancellation window?” The agent resolves these alone from day one.
  2. 02Lookup: the answer depends on the customer’s own record, such as an order status, a delivery date, or a payroll run. The agent needs read access to that system, and it resolves these alone once the read is reliable.
  3. 03Action: the customer needs something changed, such as a refund, a plan change, or a skipped delivery. The agent resolves these alone only inside the limits you write down in the next sections.
  4. 04Judgment: the right outcome depends on context a rule cannot hold, such as a bereavement, a legal threat, or a fraud claim. A person handles these. The agent collects the details and hands off.

At Fernway, the meal kit company, Dev Raghunathan found that “where is my box?” was 22% of all contacts and almost entirely a lookup against the carrier’s tracking data. Answer-only questions were another 19%. Together they were 41% of volume, and the agent resolved 96% of them inside the first month. The hard part was never the answer. It was getting the carrier feed into the agent’s hands in under two seconds.

“Everyone wants to talk about the clever cases. Our first win was boring: a tracking lookup that used to cost us $4.10 a contact now costs about 60 cents, and the answer is more accurate because a person no longer copies it by hand.”

Write the escalation rules

Escalation is a policy, not a feeling. A vague instruction such as “escalate when the customer is upset” gives you an agent that hands off too much on Monday and too little on Friday. Write each rule as a condition the agent can check, and give each rule a destination: which queue, how fast, and what the agent must collect first.

The contributors converged on a short list of triggers. Every team had these five, and every team added two or three of its own.

  • The customer asks for a person twice. One request gets a short offer to keep helping; the second request goes to a person with no argument.
  • The topic is on the judgment list: legal threats, safety, fraud, bereavement, a regulator, or the press.
  • The agent tried the same step twice and it failed, such as a refund call that returned an error.
  • The requested action is outside the agent’s limits, such as a refund over the cap or a change to a business-owner account.
  • The customer has contacted you three times in seven days about the same issue. Repeat contact is a signal that the agent’s earlier answer did not work.

The handoff itself matters as much as the trigger. A person who picks up an escalated conversation should never ask the customer to repeat anything. At Larkspur, the agent writes a four-line note at handoff: what the customer wants, what the agent already checked, what it tried, and why it stopped. Average handle time on escalated conversations dropped from 11 minutes to 6 minutes 40 seconds after the team made that note mandatory.

Actions versus answers

An answer can be wrong, but an action can be expensive. That is the whole distinction, and it should shape how much freedom you give the agent. A wrong answer costs you a follow-up contact. A wrong refund costs you money, and a wrong account change can lock a customer out of their phone number or their payroll.

Give every action a written limit, and start the limits low. Each contributor set a dollar cap on refunds and credits, a frequency cap per customer, and a list of account changes the agent may never make alone. Then they raised the limits only when the review data in the next section supported it.

  • Larkspur Mobile: account credits up to $25, at most one credit per line every 60 days. Plan downgrades yes; SIM swaps and number transfers never, because those are the doors fraudsters use.
  • Fernway: refunds up to the value of one box ($79), and only after the agent confirms the carrier marked it late or damaged. Skips and pauses with no limit.
  • Ledgerly: no refunds by the agent at all. It may correct an employee’s address or reschedule a payroll run that has not yet been sent to the bank; it may never change bank details or tax filings.
  • Pinecrest Stays: partial refunds up to 15% of the booking for listing problems a host has confirmed in writing. Full cancellations go to a person.
“Payroll is money that belongs to someone else’s employees. Our agent answers 68% of contacts and it still has never moved a dollar. I do not plan to change that this year, and nobody on our board has asked me to.”

One more rule from Ines: make every action reversible where you can, and log it where a person can see it. Larkspur’s agent writes each credit to the billing system with a reason code and the conversation link. When a credit looks wrong, a team lead can undo it in one click and see exactly why the agent gave it. In August, leads reversed 31 of 4,920 agent credits, which is 0.6%.

Review the conversations

Quality review was built for a world where people wrote every reply. A team lead read four or five conversations per agent per week, scored them on a rubric, and coached. That sample was always small, but it was enough to catch a person who was having a bad week. It is not enough for an agent that has 40,000 conversations a month.

The contributors now run review in two layers. A second model reads every conversation the agent handled and scores it against the same rubric the human team uses: was the answer correct, was the action inside policy, did the customer have to repeat themselves, and did the conversation end with the problem actually solved. People then read a sample that is weighted toward risk.

  1. 01Score 100% of agent conversations automatically, and flag any score below your threshold.
  2. 02Have a person read every flagged conversation within one business day.
  3. 03Have a person also read a random 2% sample of unflagged conversations, to check that the scoring model is not missing things.
  4. 04Read every conversation that included an action over half the agent’s limit, no matter its score.
  5. 05Each week, turn the three most common failure types into a change: a help center edit, a policy rule, or a new escalation trigger.

The output of review is not a score. It is a list of edits. At Ledgerly, 63% of the agent’s failures in the first quarter traced back to help center articles that were out of date or contradicted each other. The team rewrote 140 articles in six weeks, and the agent’s accuracy score rose from 87% to 94% with no change to the agent itself.

Measure resolution honestly

Deflection is the easiest number to inflate, and most vendor dashboards report it by default. A conversation counts as “deflected” when the customer stops replying, but a customer who gives up and calls your phone line, or who churns, also stops replying. If you report deflection to your CEO, you will eventually report a number that is not true.

Measure resolution instead, and define it strictly. All four teams use a version of the same definition: a conversation is resolved by the agent when no person touched it, and the same customer did not contact you again about the same reason within seven days, through any channel. That last clause is the one that hurts, and it is the one that matters.

When Pinecrest Stays switched from its vendor’s deflection number to this definition, the headline figure dropped from 82% to 64%. Oskar Lindqvist reported both numbers to his leadership team for one quarter, and then only the lower one. Over the next two quarters, the honest number climbed to 73%, and this time the climb was real: repeat contacts fell by a third.

“The first time I showed the real number, our CEO asked why it was so much worse than the vendor’s slide. I said, because this one counts the guests who called us afterwards. He never asked about the other number again.”
  • Agent resolution rate: resolved by the agent, no repeat contact within seven days, on any channel.
  • Cost per resolved conversation: everything you pay the vendor and the model providers, divided by resolved conversations, not by all conversations.
  • Escalation quality: the share of handoffs a person rated as complete, with no need to ask the customer again.
  • Satisfaction on agent conversations, reported separately from human conversations so that one cannot hide the other.
  • Action error rate: actions a person reversed, divided by all actions the agent took.

The tools, ranked

Your choice of tools splits into two decisions that vendors like to merge. One is the helpdesk, where conversations live and where people work them. The other is the agent, which reads and answers. Some products do both well, some do one well, and the contributors who run the largest volumes keep the two decisions separate on purpose.

Support platforms and AI agentsMean of four contributor scores out of 10
  1. 01Intercom · Best for one product for helpdesk and agentFin resolves well out of the box and the inbox is the best place for people to pick up a handoff. Pricing per resolution adds up fast at high volume.8.8
  2. 02Decagon · Best for complex actions across many systemsThe strongest at multi-step actions with written limits, and ops can own the rules. It needs a helpdesk next to it.8.6
  3. 03Zendesk · Best for large teams with existing workflowsEvery integration exists and every hire knows it. Its own agent is improving, but most contributors pair it with a separate one.8.0
  4. 04Ada · Best for many languages and channelsGood coverage of voice, chat, and 50-plus languages, with a clear reporting layer. Setup of actions takes longer than the others.7.7
  5. 05Gorgias · Best for ecommerce brands on shopifyOrder lookups and returns work on the first day. It is less suited to subscription billing or regulated products.7.4
  6. 06Front · Best for small teams that live in emailA pleasant shared inbox for the people who handle escalations. It is not where you run an agent at 40,000 conversations a month.6.9

A note on price. Per-resolution pricing looks cheap at 5,000 conversations and expensive at 50,000. Before you sign, model your cost at three times today’s volume and at your honest resolution rate, not the vendor’s. Two contributors renegotiated to a flat platform fee plus a lower per-resolution rate once they passed 30,000 conversations a month.

Staff the human team

Fewer people does not mean the same people doing less. The conversations a person still sees are the hard ones: the angry customer, the fraud claim, the bug nobody has reported yet. A team built to clear easy tickets fast will struggle with a queue that is all judgment. Hire and train for that queue, and pay for it.

Larkspur went from 86 frontline agents to 31 people over 14 months while volume grew from 41,000 to 64,000 conversations a month. Nobody was laid off; the team shrank through attrition and through moves into new roles. The remaining 31 split into three jobs.

  • Specialists (22 people) work escalations. Each one covers two or three topics in depth, such as billing disputes or device problems, and their handle time is longer by design.
  • Quality and knowledge (6 people) read flagged conversations, run calibration, and own the help center. Every article has a named owner and a review date.
  • Support operations (3 people) own the agent’s rules, its limits, its integrations, and the weekly report. Two of them were frontline agents two years ago.
“The best person to write a refund rule is someone who has issued ten thousand refunds by hand. We promoted from the floor for every operations role, and I would do it again.”

Plan coverage around the agent, not around the clock. The agent answers at 3 a.m., so your people only need to be awake for what the agent cannot do. Fernway moved to a single overnight specialist for urgent escalations and cut its night shift from nine people to one, with no change in its seven-day repeat contact rate. Do keep a person on call for the triggers that cannot wait.

“Our specialists earn about 30% more than our frontline agents did, and they are worth it. A specialist who saves a $900-a-year subscriber pays for a week of their own salary.”

So what do you do?

Start small and measure honestly, then widen the agent’s freedom in steps you can defend. None of the contributors launched an agent on all contact reasons at once, and none of them regret the slow start. The teams that struggled were the ones that trusted a dashboard before they read the conversations behind it.

  1. 01This week: export 90 days of conversations, tag each with one reason, and sort the reasons into answer-only, lookup, action, and judgment.
  2. 02Within a month: launch the agent on answer-only and lookup reasons. Write your five escalation triggers and the four-line handoff note before launch, not after.
  3. 03Within a month: define agent resolution with the seven-day, any-channel rule, and report it next to whatever number your vendor shows.
  4. 04Within a quarter: add actions one at a time, each with a dollar cap, a frequency cap, and a list of changes the agent may never make. Log every action where a person can undo it.
  5. 05Within a quarter: score every conversation automatically, read every flagged one, and calibrate the scoring by hand each week.
  6. 06Ongoing: hire and promote for specialists, quality, and support operations, and let the frontline team shrink through attrition rather than layoffs.

The goal is not the highest resolution rate you can print. It is a support team where every customer either gets a correct answer in seconds or reaches a person who already knows the story. Get both halves right and the rate follows.