The honest answer is "it depends on your repetition rate" — but here's how to actually calculate your number before you commit to a build.
The math that matters more than any industry benchmark
Vendors love to quote "70% deflection" as a universal number. It isn't — it's specific to companies with a high repeat-question rate, and applying it to your business without checking your own data will set the wrong expectations internally. The number you actually need: pull three months of tickets, categorize them, and calculate what percentage falls into your top 10-15 recurring topics. That percentage is roughly your deflection ceiling — the maximum a well-built bot could realistically handle without human involvement.
We've seen this range from 35% (a company with genuinely varied, complex support needs — think custom enterprise software) to 78% (a company with high-volume, repetitive questions — think e-commerce order status and sizing). Both are legitimate outcomes; the mistake is assuming your business will hit the higher end without checking.
What "deflection" actually measures, and what it doesn't
Deflection rate — the percentage of conversations the bot resolves without human involvement — is the headline metric, but it's incomplete on its own. We track four numbers together: deflection rate, average resolution time, escalation accuracy (did the bot correctly recognize when it needed to hand off, versus giving a wrong answer confidently), and customer satisfaction on bot-only interactions versus human-assisted ones. A bot with 80% deflection and poor escalation accuracy is worse than one with 55% deflection and near-perfect escalation — the first one is quietly giving wrong answers to the 20% it shouldn't have handled.
A concrete example with real numbers
A mid-size healthcare services client came to us fielding roughly 1,800 support tickets a month, with 3 full-time support staff. Ticket analysis showed 62% fell into 12 recurring categories: appointment scheduling questions, insurance coverage basics, portal login issues, and a handful of others. Post-launch, the bot resolved 58% of total volume independently (slightly below the 62% ceiling, because some "recurring" questions still had enough variation to need human judgment), reduced average first-response time from 3.1 hours to under 2 minutes for bot-handled queries, and freed enough staff capacity that the client postponed a planned support hire for roughly 8 months.
Where the estimate goes wrong
The most common overestimate comes from counting "questions with a documented answer" as the ceiling, rather than "questions customers actually phrase simply enough for a bot to recognize." A policy might be documented clearly, but if customers ask about it in twelve different indirect ways, retrieval accuracy drops and so does real-world deflection. We test against actual historical phrasing, not idealized questions, before committing to a number.
How Ndakum approaches it
Before we build anything, our AI Chatbot engagements start with this exact analysis on your ticket data — you get a real deflection estimate specific to your business, not an industry average, before deciding whether to move forward.
Curious whether this fits your business?
A short conversation will tell us both. No pressure, no obligation.
Book a consultation