AI and Customer Service: What Are Companies Actually Achieving?
A research-led look at the evidence behind the claims
Artificial intelligence has moved rapidly from being an interesting experiment in customer service to becoming part of the operating model of some of the world's largest companies.
Customer-service teams are using AI to answer customers directly, help agents find information, draft responses, summarise conversations and reduce the amount of routine work handled by people.
And the numbers being reported can be impressive.
Companies are reporting large reductions in response times, improvements in first-time resolution, millions of customer conversations handled by AI and significant cost savings.
But there is an important question behind all of these numbers:
What do they actually tell us about customer service?
To explore this, we looked at publicly available evidence from companies that have implemented AI in customer-service operations.
Rather than simply collecting impressive-sounding statistics, we have tried to look at what each measure actually means, where the evidence comes from and what it does — and does not — demonstrate.
How we researched this
We looked primarily for publicly available information from the companies themselves, including regulatory filings, investor presentations and published customer case studies from technology providers.
We have prioritised original sources where possible.
That distinction matters.
A result published in a company's regulatory filing is different from an independently audited study. A technology provider's customer case study is different again. And a company's own measurement of customer satisfaction should not automatically be treated as independent evidence that customers prefer AI.
Throughout this article, we therefore identify the source of the claims.
The figures below should be read as reported results, rather than as a controlled scientific comparison of AI and human customer service.
1. Klarna: AI handling a large share of customer conversations
Klarna is probably one of the most widely discussed examples of AI being used in customer service.
In February 2024, the company launched an AI assistant developed in partnership with OpenAI.
OpenAI initially reported that, during the assistant's first month, it handled 2.3 million conversations — approximately two-thirds of Klarna's customer-service chats at the time. Klarna said the assistant was doing work equivalent to around 700 full-time agents, resolving customer issues in less than two minutes compared with 11 minutes previously, and reducing repeat enquiries by 25%. It also said customer satisfaction was on par with human agents.
The figures continued to develop as Klarna expanded the system.
In its third-quarter 2025 results, Klarna reported that its AI assistant was doing the equivalent work of 853 full-time agents. The company's investor presentation said the assistant was handling around 28 million annualised conversations, solving 81% of customer-service chats, and delivering approximately $58 million in annualised cost savings as of 30 September 2025. Klarna also said customer satisfaction was on par with human agents.
That was a snapshot of the programme at the end of September 2025.
Klarna's subsequent 2025 Annual Report, covering the full year to 31 December 2025, reported that the AI assistant had handled 80% of customer-service chats during the year, with no drop in consumer satisfaction levels. The company's detailed filing also reported work equivalent to more than 850 full-time agents and approximately $59 million in cost savings during 2025.
What does this tell us?
This is unusually strong evidence of scale.
The AI isn't being used simply as a small experimental chatbot. Klarna's own reported figures indicate that it became responsible for a substantial proportion of customer-service interactions.
The figures also show why the date of a statistic matters.
The Q3 results provide a snapshot of the programme as it stood at September 2025. The subsequent annual report covers the entire 2025 calendar year. The percentages should therefore not be treated as contradictory measurements of exactly the same period.
There is also evidence of an operational benefit. Klarna reports substantial cost savings, shorter resolution times and fewer repeat enquiries.
What should we be careful about?
The phrase "equivalent to 853 full-time agents" is particularly important to understand.
Klarna describes this as an estimate based on reductions in chat and telephone conversations handled by full-time agents following the launch of its AI assistant.
That is a calculation of workload capacity — not necessarily a statement that 853 individual employees were made redundant.
There is a similar qualification around customer satisfaction.
Klarna says its AI-handled chats rank on par with human agents in consumer satisfaction, based on its service-chat data and consumer satisfaction surveys. That is useful evidence, but it remains company-reported evidence, rather than an independent controlled comparison.
The Klarna example therefore gives us a fairly clear picture of what AI can achieve operationally at scale.
It gives us less certainty about the broader question of whether customers generally prefer AI to human service.
And that distinction — between operational efficiency and customer experience — is one we see repeatedly throughout the other case studies.
2. Vodafone: measuring first-time resolution
Vodafone provides an interesting contrast because one of its headline measures is not primarily about cost.
The company has been developing SuperTOBi, a generative-AI version of its digital assistant, alongside SuperAgent, an AI tool designed to support customer-service employees.
Vodafone reported that initial testing of SuperTOBi produced an approximately 50% improvement in first-time resolution for critical customer journeys such as complex billing enquiries. In its FY25 H1 results presentation, Vodafone also reported a roughly 50% improvement in first-time resolution based on interactions with 7 million customers.
Why is this interesting?
First-time resolution is much closer to the customer experience than a simple measure of AI adoption.
A customer generally doesn't care how sophisticated the underlying technology is.
They care whether their problem is solved.
If a customer can get their problem resolved during the first interaction, that potentially represents a meaningful improvement in service.
But there is a catch
"First-time resolution" is a useful metric, but it needs context.
A higher first-time-resolution rate doesn't automatically tell us whether the underlying answer was better, whether the customer was happier or whether the issue stayed resolved.
And Vodafone's figures are company-reported rather than the result of an independent controlled study.
Still, this is a good example of an AI programme being measured against an outcome that customers actually experience.
3. Simplyhealth: AI helping people rather than replacing them
Not all of the interesting examples involve AI talking directly to customers.
Simplyhealth has used generative AI to help customer-service employees respond to emails.
According to Salesforce, Simplyhealth initially trained its system to create knowledge-based email replies to frequently asked questions. The AI-generated response is reviewed and edited by an employee before being sent to the customer.
The process that previously took around 12 minutes was reduced to approximately one minute. Salesforce reported that Simplyhealth was handling more than 600 such emails each week, saving more than 90 hours of staff time, while the company reported productivity improvements of up to 90%.
More recent Salesforce reporting says Simplyhealth has since expanded its use of AI, including autonomous handling of routine enquiries, and now reports around 120 hours of weekly time savings.
Why does this matter?
It demonstrates that "AI customer service" doesn't necessarily mean replacing the customer-service agent.
There is another possibility:
AI can remove some of the work around the conversation so that the human has more time for the conversation itself.
For relatively straightforward enquiries, drafting an answer can be repetitive.
If AI can retrieve the relevant information and produce a useful first draft, the employee can spend more time checking the answer, understanding the customer's circumstances and dealing with the parts of the interaction that require judgement.
What should we be careful about?
These figures come from Salesforce's customer case study and Simplyhealth's own reported results.
They are therefore useful evidence of what the company says it achieved, but they are not independent experimental results.
And "up to 90%" is not the same thing as saying that every customer-service interaction became 90% more productive.
The more specific and useful finding is that Simplyhealth reported substantial time savings from using AI for particular categories of customer communication.
4. Sicredi: a small pilot with measurable results
One of the more useful examples we found comes from Brazil.
Sicredi worked with IBM to develop a generative-AI assistant designed to help customer-support representatives find answers in the organisation's support documentation.
The company spent three weeks co-creating the assistant and then tested it for 20 days.
During the pilot, support representatives served 6,500 members asking questions about Sicredi's consortium product.
Compared with the previous month's customer-support data for that product, Sicredi reported:
a 10–12% improvement in queries resolved without involving a product specialist;
a 1% improvement in Net Promoter Score;
an 8% reduction in support-call abandonment caused by waiting times; and
a reduction in average time to resolve customer queries.
Why this example matters
It is a relatively small experiment, but it illustrates something important about evaluating AI.
The system was not judged simply on how many questions it could answer.
The company looked at several measures:
resolution without escalation, customer advocacy, abandonment and resolution time.
That gives us a more rounded picture of what the technology was doing.
It also shows why pilot projects can be valuable.
Rather than making a broad claim that "AI improves customer service", the company can compare a defined group, over a defined period, against previous performance.
What should we be careful about?
This was a 20-day pilot, and the results were published by IBM.
It therefore shouldn't be treated as proof that the same improvements will occur across every customer-service operation.
But it is a useful example of how AI projects can be evaluated more meaningfully.
5. Lenovo: AI as an agent assistant
Lenovo provides another example of the human-plus-AI model.
Its Premier Support operation uses Microsoft Dynamics 365 Contact Center and Customer Service with Copilot.
When a customer engages with support, AI helps the service representative identify possible solutions using historical service interactions. The system can also create a summary after the interaction.
Microsoft's published customer story says Lenovo achieved:
15% higher agent productivity
20% lower average handling time
record-high customer satisfaction.
What does this tell us?
This is another example where AI isn't necessarily replacing the person dealing with the customer.
Instead, it is reducing the amount of time the employee spends searching for information and documenting the interaction.
A customer may never know that AI was involved.
The visible result is simply that the employee can potentially deal with the problem more quickly.
And again, the measurement matters
"Agent productivity" is an operational metric.
It isn't automatically the same thing as better customer service.
However, Lenovo also reports higher customer-satisfaction ratings, which at least begins to connect the efficiency improvement with a customer outcome.
The evidence is still based on a customer story published by Microsoft, so it should be understood as a reported business result rather than an independently verified experiment.
6. Nationwide: perhaps the most revealing use of AI is behind the scenes
Nationwide offers another useful example because its AI implementation has been positioned as a copilot rather than an autopilot.
Microsoft reports that Nationwide is using GPT-4 through Azure OpenAI to help with customer correspondence.
According to Microsoft, AI-assisted letters reduced response times from approximately 45 minutes to around 10–15 minutes.
The customer doesn't necessarily interact with an AI system at all.
Instead, the employee uses AI to help prepare the response.
This changes the question
When we talk about AI and customer service, it is easy to focus on chatbots.
But there is a much larger opportunity behind the scenes.
Customer-service organisations spend enormous amounts of time:
searching knowledge bases;
reading previous interactions;
summarising calls;
drafting emails;
categorising enquiries;
finding relevant policies;
updating case notes.
These activities are largely invisible to the customer.
If AI can reduce that workload, the customer may experience the benefit as a faster or more informed response — without ever knowing AI was involved.
This may be one of the most important areas of AI adoption in customer service.
7. NatWest: when the headline number needs careful reading
NatWest provides a particularly interesting example of why we need to read AI claims carefully.
The bank told a UK Parliament inquiry that its Cora digital assistant handled more than 11 million customer interactions in 2024.
NatWest also reported that its generative-AI upgrade, Cora+, delivered up to a 150% increase in customer satisfaction, while a pilot had halved the number of query cases requiring colleagues to intervene.
At first glance, a 150% increase in customer satisfaction sounds extraordinary.
But this is precisely where research-style reporting needs to slow down.
NatWest's evidence does not mean that customers became "150% happier" in some universal measure of satisfaction.
The statement relates to the bank's own measurement of Cora+ and describes an increase in customer satisfaction associated with the system.
Without knowing the precise baseline, sample, methodology and calculation behind the percentage, the number cannot sensibly be compared with another company's customer-satisfaction figure.
The lesson
A percentage is not automatically a comparable metric.
"50% improvement in first-time resolution", "20% reduction in handling time" and "150% increase in customer satisfaction" sound like three numbers that could sit neatly beside each other.
They can't.
They measure different things, using different methodologies, in different organisations.
That is one of the biggest problems with trying to create a simple league table of AI customer service.
What do these case studies have in common?
Despite the differences between the companies, several patterns appear repeatedly.
1. AI is producing measurable operational improvements
Across the examples, companies report:
shorter response times;
shorter handling times;
more enquiries resolved without escalation;
fewer repeat enquiries;
higher agent productivity;
fewer abandoned interactions;
significant volumes of conversations handled automatically;
lower service costs.
This is probably the clearest area of evidence so far.
Companies can measure operational performance relatively easily.
Time is measurable.
Volume is measurable.
Cost is measurable.
The number of interactions handled by a system is measurable.
2. Customer outcomes are harder to measure
Customer service is ultimately about the customer.
But customer outcomes are more complicated.
A shorter interaction isn't necessarily a better interaction.
A conversation resolved in two minutes may be excellent if the customer's problem is genuinely solved.
It may be terrible if the customer has to contact the company again tomorrow.
This is why metrics such as repeat contact, first-time resolution, customer satisfaction, abandonment and escalation are potentially more interesting than raw AI adoption numbers.
Klarna's reported reduction in repeat enquiries is therefore more informative than simply knowing how many conversations its AI handles.
Likewise, Vodafone's first-time-resolution measurement addresses a customer-service outcome rather than simply reporting chatbot usage.
3. "AI handled X% of conversations" doesn't tell the whole story
This is perhaps the easiest metric to misunderstand.
Suppose an AI system handles 80% of conversations.
That sounds like an enormous achievement.
But there are several questions underneath it:
What kinds of conversations are included?
How difficult are they?
How many are actually resolved?
How many customers contact the company again?
How many are eventually transferred to a human?
What happens to the remaining 20%?
And perhaps most importantly:
What does the customer experience?
Conversation volume tells us about the scale of automation.
It doesn't, on its own, tell us about service quality.
4. "Equivalent to X agents" needs particularly careful interpretation
The Klarna example illustrates this well.
Klarna says its AI does work equivalent to more than 850 full-time agents.
That is a useful way of communicating the scale of the workload being automated.
But it shouldn't automatically be interpreted as more than 850 jobs being eliminated.
An FTE-equivalent calculation measures workload capacity.
A company can use that capacity in several ways:
reducing staffing requirements;
handling a growing customer base without proportional hiring;
reducing outsourced service costs;
moving employees to more complex work;
extending service hours;
or some combination of these.
The number tells us something important about capacity.
It doesn't tell us by itself what happened to the people who previously performed the work.
5. The most interesting implementations may be the least visible
The examples from Simplyhealth, Lenovo and Nationwide all point towards the same idea.
AI doesn't have to sit between the customer and the employee.
It can sit beside the employee.
That means:
Customer → Human → AI assistance
rather than:
Customer → AI → Human if necessary
The first model may be particularly important for complex customer service.
AI can search, summarise, draft and recommend.
The human can interpret, empathise, make judgements and take responsibility for the final response.
That is a very different proposition from simply trying to automate the entire conversation.
So, is AI actually improving customer service?
The evidence we found suggests that companies are achieving real and measurable improvements in customer-service operations.
There is public evidence of faster responses, shorter handling times, greater automation, fewer repeat enquiries, more first-time resolution and significant reported cost savings.
But the evidence is much less straightforward when the question becomes:
Does AI provide a better customer experience than a human?
There isn't a single answer in the data we examined.
Some companies report customer satisfaction that is comparable with human service.
Others report improvements in satisfaction.
Some are measuring operational outcomes rather than satisfaction at all.
And many of the published figures come from the companies themselves or from the technology providers supplying the AI.
That doesn't make the results meaningless.
It simply means we should understand what kind of evidence we are looking at.
The customer-service AI scorecard
Perhaps the most useful way to evaluate future claims is not to ask:
"How much AI is this company using?"
Instead, ask five questions:
1. What has actually been automated?
Is AI answering customers directly, or helping employees?
2. What is being measured?
Is the headline number about cost, speed, volume, productivity, resolution or customer satisfaction?
3. Is the result company-reported?
If so, has the methodology been independently verified?
4. What happened to repeat contact?
A fast answer is not necessarily a successful answer.
5. What happened to the customer?
Ultimately, the most important question is whether the customer got what they needed, with less effort and less frustration.
The bigger question for customer service
The most interesting conclusion from this research may be that AI customer service is not really one thing.
There are at least three different developments happening at once.
AI as the agent:
The technology communicates directly with the customer and attempts to resolve the enquiry.
AI as the assistant:
The technology helps a human agent find information, draft responses and complete administrative work.
AI as the infrastructure:
The technology works behind the scenes to route, classify, summarise and analyse customer interactions.
Companies are already reporting measurable benefits from all three.
But they are not interchangeable.
And as more businesses publish increasingly impressive AI statistics, the ability to distinguish between automation, efficiency and genuine customer-service improvement will become increasingly important.
The next stage of AI in customer service may therefore be less about asking whether companies are using AI.
They clearly are.
The more useful question is:
Are they using it to make customer service more efficient — or to make customer service better?
Those are not necessarily the same thing.
And the evidence, so far, suggests we should keep measuring both.
Sources and methodology
This article is based on publicly available information published by the companies and technology providers discussed.
Primary sources used include Klarna's SEC filings and investor presentations, Vodafone's investor materials, the UK Parliament's published evidence from NatWest, and customer case studies published by IBM, Salesforce and Microsoft.
Where a result is company-reported or appears in a technology-provider case study, it is described as such. We have not treated those figures as independently verified research.
That distinction is important because AI customer-service metrics are still developing, and apparently similar percentages can represent very different things.
For this reason, the figures in this article are intended to illustrate what companies are reporting — not to create a league table of AI customer-service performance.
Prepared by ChatGPT and prompted on 25th September 2026

No comments:
Post a Comment