How Many Support Tickets Can AI Deflect? Benchmarks by Industry
Published AI ticket deflection and resolution rates from Intercom, Zendesk, Salesforce and Klarna, what each one measures, and how to measure your own honestly.

Short answer
Published AI ticket deflection rates cluster in two bands. New deployments typically resolve 25–50% of the conversations the AI handles, and mature ones report 60–80%: Intercom says Fin averages 76% across its customers, Salesforce reports 76% on its own help site, and Klarna's assistant handled two-thirds of chats in its first month. Every one of those numbers is vendor-reported, measured differently, and counts abandoned chats as wins. PepoChat has no published statistics; the figures here are all third-party.
If you are deciding whether an AI support agent is worth setting up, the first question is always "how many tickets will it actually take off my plate?" The honest answer is a range, and the range depends more on how you count than on which product you buy. This post is for founders and support leads at small teams who want a realistic target.
Two things up front. First, PepoChat does not publish a deflection statistic. Every figure below comes from a page we opened at a third party, with the date and what was measured next to it, and you should treat vendor figures as marketing until you have reproduced them on your own traffic. Second, "deflection", "containment" and "resolution" are used interchangeably by people who bill you on them, so we define them first.
What does "AI ticket deflection rate" actually mean?
Four terms get mixed together, and they measure different things.
Deflection rate is the share of would-be tickets that never reach a human because the customer found an answer first, from a help centre article, a search result or a bot. Deflection is measured against tickets that would otherwise have been created, which is a counterfactual nobody can observe directly, so it is usually estimated from a drop in ticket volume after launch.
Containment rate is the share of conversations the bot handled that did not escalate to a person. The formula is one minus escalated conversations divided by total conversations. Containment says nothing about whether the customer got what they came for; a visitor who gave up and closed the tab is "contained".
Resolution rate is the share of conversations the AI was involved in that ended with the issue solved. The catch is the word "solved". Intercom's definition counts a resolution when, after Fin's last answer, "the customer either confirms the answer was satisfactory (confirmed resolution), or exits the conversation without requesting further assistance (assumed resolution)". An assumed resolution is a customer who went quiet, and Intercom's separate help centre for Fin puts the silence window at 24 hours.
Verified or CSAT-weighted resolution is the strictest version: a conversation counts only when there is a positive signal from the customer, such as a thumbs-up, a "that helped" reply, or no reopen within a set period. Zendesk formalised this split in May 2026, and its announcement defines a contained resolution as "a meaningful request is handled by the AI agent, without human involvement or follow-up" and a verified resolution as the same "with additional signals confirming the outcome was complete and satisfactory". Zendesk bills only for verified resolutions.
The denominator matters as much as the numerator. Intercom's resolution rate is "the number of conversations Fin resolved, divided by the number of conversations Fin was involved" in, and it reports a separate involvement rate for the share of all conversations Fin touched. A 76% resolution rate on 60% involvement is 46% of total volume. When you read any benchmark, ask: resolved by whose confirmation, and divided by what?
| Term | Numerator | Denominator | What it hides |
|---|---|---|---|
| Deflection | Tickets that never got created | Tickets that would have been created | The counterfactual is estimated, not measured |
| Containment | Conversations with no human escalation | All bot conversations | Customers who gave up |
| Resolution (assumed) | Confirmed plus silent-after-answer conversations | Conversations the AI was involved in | Silence counted as success; involvement rate |
| Verified resolution | Conversations with a positive customer signal | Conversations the AI was involved in | Customers who were satisfied but never clicked |
What AI ticket deflection rates do vendors report?
Here are the figures we could actually find on the source's own pages, with dates. All of them are self-reported by the vendor or the company deploying the bot.
Intercom Fin. In a March 2026 blog post, Intercom wrote that "more than 12,000 teams use Fin" and that its "average resolution rate across customers has increased every month and now stands at 76%". That is an average across customers using the assumed-resolution definition above. In a separate March 2026 post about its own support team, Intercom said Fin "now resolves over 81%" of its support volume, having started at "over 25%" during an early beta. On Intercom's community forum, a support staff member told a customer that "a good resolution rate starts at around 30–50%", and users in the same thread reported 25%, 45% and 55–60%.
Zendesk. Zendesk's June 2026 guide to automated resolution rate gives the formula as issues fully resolved by AI divided by total issues handled by AI, and lists five conditions for a true resolution, including "require no follow-up support for the same issue". Its published Lush case study reports a 60% first-contact resolution rate for the retailer's AI agent, saving "5 minutes per ticket" and "360 agent hours each month"; the page carries no date.
Salesforce. A December 2024 customer story about Agentforce on Salesforce's own help site reports "76% of customer inquiries resolved without a human" across "more than 1.7 million conversations", while "escalating only 5% of inquiries to a human Support Engineer". Note the gap: 76% resolved and 5% escalated leaves roughly a fifth of conversations that were neither, which is the abandoned-chat problem in one line. Salesforce's undated Agentforce metrics page also lists a 58% chat containment rate at CVS Health, 62% of support cases resolved autonomously at Xero, 89% of "routine messaging inquiries" at Canada Goose and a "95% case deflection rate" at Live Nation.
Klarna. Klarna's February 2024 press release said its OpenAI-powered assistant had "2.3 million conversations, two-thirds of Klarna's customer service chats" in its first month, doing "the equivalent work of 700 full-time agents", with customer satisfaction "on par with human agents", a "25% drop in repeat inquiries" and resolution time down from 11 minutes to under 2. Fifteen months later, in May 2025, CX Dive reported that Klarna was recruiting human agents again, with its CEO saying cost had been "too predominant" in the original decision and that "there will always be a human if you want". Two-thirds of chats handled is a share-of-volume figure, not a resolution rate.
Gartner. In March 2025 Gartner predicted, as reported by CX Today, that "agentic AI will autonomously resolve 80 percent of common customer service issues without human intervention by 2029", with a 30% reduction in operating costs. It is a forecast about common issues, and it is quoted in vendor decks as if it were a current benchmark.

Chatbot resolution rate benchmarks by industry and use case
There is no neutral, audited dataset of AI resolution rates by industry. What exists is a scatter of case studies, each measured on the vendor's terms. The table below groups the published examples above by the kind of work involved, and gives the range they imply. Treat the ranges as "what companies have publicly claimed", not "what you should expect".
| Industry / use case | Published examples (source, date) | Implied range | Caveats |
|---|---|---|---|
| Software and SaaS product support | Salesforce help site 76% resolved (Salesforce, Dec 2024); Intercom's own support 81% (Intercom, Mar 2026); Xero 62% of cases (Salesforce metrics page, undated) | 60–80% | All three are the vendor or a flagship customer with a large, well-maintained help centre |
| Ecommerce and retail | Lush 60% first-contact resolution (Zendesk, undated); Canada Goose 89% of routine messaging inquiries (Salesforce, undated) | 60–90% of routine questions | "Routine" is doing a lot of work; order status and returns are easy, complaints are not |
| Fintech and payments | Klarna two-thirds of chats handled (Klarna, Feb 2024) | ~65% share of volume | Share of chats, not resolution; Klarna later rehired humans for quality |
| Healthcare and pharmacy | CVS Health 58% chat containment (Salesforce metrics page, undated) | 50–60% containment | Containment, not resolution; regulated topics are routed to people by design |
| Events and ticketing | Live Nation 95% case deflection (Salesforce metrics page, undated) | Up to 95% | Deflection of a narrow, high-volume question set around events |
| Any industry, first 30–90 days | Intercom staff guidance 30–50% (community, 2025); Intercom beta start "over 25%" (Mar 2026) | 25–50% | The realistic starting band for a small team with a thin knowledge base |
| Cross-customer vendor average | Intercom Fin 76% (Mar 2026) | 76% | Assumed resolutions included; averaged across 12,000+ teams of every size |
Three patterns hold across the table. Narrow, repetitive question sets (order status, event logistics) produce the highest numbers. Regulated or emotionally loaded domains produce the lowest, on purpose, because the design routes those conversations to people. And every company in the top band had months or years of content work behind the number; Intercom's own team went from 25% to 81% over roughly three years and created a dedicated knowledge-manager role to do it.
Why do vendor-reported deflection numbers run high?
None of the companies above is lying. The numbers are high because of how the counting works, and the same traps will inflate your own dashboard if you let them.
Silence counts as success. Under an assumed-resolution rule, a customer who asked "how do I cancel?", got a wrong answer, and left to email you from a different tab is a resolved conversation. Intercom's help centre is explicit that a resolution is deducted only if the customer "later returns to the same conversation seeking further assistance". A new conversation, a phone call or a chargeback does not undo the count. Salesforce's 76% resolved next to 5% escalated shows the size of the grey zone: the remaining conversations went nowhere, and the vendor-friendly reading is to leave them out of the denominator.
The denominator is conversations the bot was involved in, not all conversations. If the bot only fires on the pricing page, or only in English, or only after a pre-chat form that half your visitors abandon, the denominator shrinks and the rate rises. Ask for involvement rate alongside resolution rate, and multiply them.
Per-resolution billing creates an incentive to count generously. When the vendor charges per resolution, every definitional edge case is money. One customer on Intercom's forum described being charged for resolutions in cases where a human agent stepped in before the customer could ask for one, because under the rules the customer had "already received an answer and hasn't asked for additional help". Zendesk's move to bill only on verified resolutions is a direct response to that pressure, and a healthy sign, but it is also a reminder that the older numbers were computed on the looser rule.
Greetings and small talk. Some systems exclude a conversation where the bot only said hello (Intercom does; Zendesk's "unassisted conversation" category does too). Others count every session, so a bot with an aggressive greeting on every page can show a large or small denominator depending on the rule.
Reopens and channel shift. A ticket that comes back three days later on email is a failed resolution that the chat metric never sees. The only way to catch it is to join chat transcripts to your ticket system by customer, which almost nobody does.
Case-study survivorship. The companies on a vendor's metrics page are the ones that agreed to be there; the customer who churned at 22% does not get a logo tile.

How to measure your own AI deflection rate honestly
You need a definition written down before launch, one baseline, and a weekly habit.
1. Baseline before you switch anything on
Count four weeks of inbound conversations by channel (chat, email, form, phone) and by topic. Ten to fifteen topics is plenty: "where is my order", "how do I cancel", "password reset", "pricing question" and so on. Write down the weekly totals. Without this, you cannot compute deflection later, because you will not know what volume would have been.
2. Run the top-ten-questions test
Take the ten most common topics from the baseline and write two real customer phrasings for each, copied from actual tickets, spelling mistakes included. Ask the AI support agent all twenty before it goes live. Score each answer on three points: correct, complete, and either finished the task or offered a person. Anything under 15 of 20 means the knowledge base has gaps, and no launch metric will fix that. Repeat the test monthly; it is the cheapest leading indicator you have. The training guide covers how to fill the gaps it finds.
3. Define "resolved" and stick to it
Pick one rule and write it in the same document as the baseline. A defensible small-team rule is: a conversation is resolved when the visitor gave a thumbs-up, or replied with thanks, or asked nothing further within 24 hours and did not open another conversation or email on the same topic within seven days. Everything else is unresolved or escalated. That is stricter than most vendor rules, which is the point.
4. Read a sample every week
Open twenty random conversations a week and label them by hand: resolved, abandoned, escalated, wrong answer. After a month you will have a hand-labelled rate to compare against whatever your dashboard says, and the gap between the two is your correction factor. In PepoChat the team inbox shows every conversation with its unresolved, escalated or resolved status, the full transcript, and the visitor's thumbs-up or thumbs-down vote, and the analytics view groups conversations by auto-labelled topic in weekly buckets, so the sampling is a filter rather than an export.
5. Compute three numbers, not one
Each month, report containment (conversations that did not escalate ÷ all bot conversations), verified resolution (positive-signal conversations ÷ all bot conversations), and deflection (baseline weekly tickets minus current weekly human tickets, adjusted for any change in traffic). If containment is 70% and verified resolution is 35%, you have a bot that stops people rather than helps them, and the next section is where to look.
What moves the deflection number?
Vendor variance is smaller than setup variance. The same product can sit at 25% or 70% on two sites, and the difference is almost always one of the following.
Content coverage. The AI support agent answers from your knowledge base, so a question that is not covered in the knowledge base cannot be resolved, only escalated. Most launches under 40% are content problems: the refund policy is a PDF nobody uploaded, or the help centre describes last year's interface. Run the top-ten test, find the misses, add the source.
Answer honesty. A bot that guesses will score well on containment and badly on verified resolution, because the customer leaves with a wrong answer and comes back angry on email. Grounded answers with a clear "I don't know, let me get a person" lose a few containment points and gain real resolutions; how to stop your chatbot hallucinating goes through the settings involved.
Handoff quality. Escalation is not failure; a bad escalation is. If the handoff drops the transcript, forces the customer to repeat themselves, or lands in an inbox nobody watches on weekends, the customer's second message goes to email and your deflection number quietly falls. Design the human handoff so that the AI escalates when it finds nothing relevant or when the visitor asks for a person, and so the operator sees the whole conversation.
Actions. The biggest step change in the case studies above comes from letting the agent do things rather than describe them. "Where is my order" answered from a policy page is a partial resolution; answered with the live tracking status from the store, it is a complete one. The same goes for checking a subscription, filing a ticket in your helpdesk, or booking a call. Connecting an order lookup is usually the single highest-leverage change for an ecommerce site, and connecting a chatbot to Shopify order status shows what it involves.
Scope and question design. Suggested questions in the widget steer visitors toward things the bot is good at, and a greeting that sets expectations ("I can help with orders, billing and setup; for anything else I'll get a person") reduces abandoned chats. Restricting the bot to the pages where its knowledge applies raises the rate honestly, as long as you report involvement rate next to it.
Reply budget on free tiers. If you are on a capped plan, the cap is part of your measurement. On PepoChat's free plan, for example, the 500 AI replies a month are the ceiling; after that, new conversations go straight to the team inbox with a message rather than failing, which is fine for customers but will show up as a drop in containment at the end of a busy month. The pricing page has the current limits.
What to do next
Start with the baseline and the top-ten-questions test, because they cost nothing and tell you more about your likely rate than any benchmark in this post. Then decide what "resolved" means for you and write it down. If you are still working out what kind of tool you need, what is an AI support agent separates agents from scripted bots and copilots, and the use cases page shows the setups (order status, subscriptions, booking) that produce the highest verified rates. Every feature on the free plan is available at app.pepochat.com/sign-up, so you can run the twenty-question test on your own content before anyone on your team commits to a number.
Frequently asked questions
- What is a good AI ticket deflection rate?
- Published figures put new deployments at 25–50% of the conversations the AI handles and mature ones at 60–80%. Intercom staff describe 30–50% as a good starting point. Those numbers are vendor-reported and mostly count silent customers as resolved, so a verified resolution rate of 40–50% within three months is a strong result for a small team.
- What is the difference between deflection, containment and resolution rate?
- Deflection is the share of would-be tickets that never reached a human, estimated from a drop in volume. Containment is the share of bot conversations that did not escalate, whether or not the customer was helped. Resolution is the share of AI-handled conversations where the issue was solved, and verified resolution requires a positive customer signal such as a thumbs-up.
- What resolution rate does Intercom Fin report?
- In March 2026 Intercom said Fin's average resolution rate across more than 12,000 customer teams stood at 76%, and that Fin resolves over 81% of Intercom's own support volume. Intercom's definition counts both confirmed resolutions and assumed ones, where the customer simply stops replying after an answer, so the figure is not directly comparable with stricter measures.
- Why are vendor-reported deflection numbers higher than what teams see?
- Most vendor definitions count a conversation as resolved when the customer goes quiet, use only the conversations the bot was involved in as the denominator, and never see a reopen that arrives by email. Per-resolution billing also rewards generous counting. Case studies are drawn from successful customers, so published averages sit above what a typical first deployment achieves.
- How do I measure my own chatbot resolution rate?
- Record four weeks of ticket volume by topic before launch, write down one definition of resolved, then hand-label twenty random conversations a week as resolved, abandoned, escalated or wrong. Report containment, verified resolution and deflection separately each month, and run a monthly test of the ten most common questions in two real customer phrasings each.
- Does PepoChat publish a deflection or resolution rate?
- No. PepoChat has no published deflection statistics, and every number in this article comes from a third-party source with its date and definition. What PepoChat does provide is the data to measure your own rate: conversation statuses, full transcripts, visitor thumbs-up and thumbs-down votes, and weekly analytics by auto-labelled topic, all included on the free plan.
Try this on your own site in ten minutes
PepoChat includes every feature on the free plan — 500 AI replies and 10 knowledge sources a month, no credit card.
Keep reading
Chatbase vs PepoChat: Which AI Support Agent Fits Your Team in 2026?
An honest side-by-side of Chatbase and PepoChat on pricing, free tier, sources, handoff, booking, actions, channels and voice, plus a pick by team type.
Intercom Fin Pricing Explained (and What a Flat-Plan Alternative Costs)
What Intercom Fin's $0.99 per resolution adds up to once you count outcomes and seats, with a worked 2,000-conversation example and when a flat plan wins.