Skip to content
Chatbotscape
Verified
Customer Effort Score· Support metrics
Customer Effort Score (CES) is a single-question survey metric that measures how much work a customer had to do to get their issue handled, asked immediately after a support interaction. It exists in two incompatible versions: the original 2010 five-point question about how much effort the customer expended, where a low score is good, and the later seven-point agreement statement about whether the company made things easy, where a high score is good. Because the two run in opposite directions and are scored differently, a CES figure means nothing without the wording and scale attached.
By Chatbotscape Editorial· Methodology· Published 19 August 2026· Updated 19 August 2026

Customer Effort Score — The Metric That Prices What the Customer Had to Do (2026)

Quick answer: CES asks a customer, right after you have dealt with them, how hard that was. It is the third of the three standard feedback metrics, alongside CSAT and NPS, and it is the one that fits a chatbot best, because effort is the variable a bot actually moves. It is also the one with the shakiest evidence and the messiest arithmetic. Two different questions have carried the name since 2010, they point in opposite directions, and vendors report them by at least three different methods. If you adopt it, write down your wording, your scale and your scoring rule, and treat any external benchmark as uninterpretable unless it publishes all three.

Two questions, one name, opposite directions

The metric was introduced in 2010 in a Harvard Business Review article by Matthew Dixon, Karen Freeman and Nicholas Toman under the headline "Stop Trying to Delight Your Customers." The argument was that companies overspend on exceeding expectations and underspend on removing obstacles, and that the thing which actually predicts whether a customer stays is how much work you made them do.

The original instrument was a single five-point item: How much effort did you personally have to put forth to handle your request?, with 1 meaning very low effort and 5 meaning very high effort. You averaged the responses. A lower score was better.

The scale was later rewritten. The company that funded the original research was purchased by the Gartner Group, and the item became a seven-point agreement statement along the lines of The company made it easy for me to handle my issue, with respondents agreeing or disagreeing. Gartner's own scoring guidance is a top-three-box calculation: count everyone answering 5, 6 or 7 and divide by all respondents. A higher score is now better. The revised wording does not contain the word "effort" at all.

Both versions are called CES. That is the single most important practical fact on this page, and it has three consequences:

Original (2010)Revised
Question typeHow much effort did you expendAgreement that the company made it easy
Points57
DirectionLower is betterHigher is better
Usual scoringMean of all responsesPercentage answering 5, 6 or 7
Contains the word "effort"YesNo

First, a raw number is meaningless on its own. A CES of 5 is the worst possible result on the original scale and a middling-to-good one on the revised scale. Second, benchmarks published before and after the change are not comparable, and in the benchmark material we have looked at the instrument is usually not stated. Third, even within the revised version, some organizations report the top-three-box percentage and others report the mean, and those are different numbers from identical data.

So before you can use CES at all you have to answer three questions about your own implementation: what exactly does the question say, how many points does the scale have, and are you reporting a mean or a box percentage. Nobody can answer them for you, and no dashboard default should be trusted to have answered them the way you assume.

What it measures that CSAT and NPS do not

The three metrics are often described as interchangeable ways of asking "was that good." They are not. They differ in what they are about and in when they can honestly be asked.

MetricWhat the customer is ratingScopeHonest to ask when
CSATHow they feel about the interaction that just happenedOne interactionImmediately after any contact
CES (this page)How much work that interaction cost themOne interactionImmediately after a service contact
NPSTheir intention toward the company as a wholeThe relationshipPeriodically, not after every ticket

The gap between CSAT and CES is narrower than it looks and it matters for automation. A customer can be satisfied and still have worked hard. They rephrased their question three times, got sent to a knowledge-base article that did not fit, went back, finally reached a person, and the person sorted it out in ninety seconds. Asked how satisfied they are with the resolution, many people say they are satisfied, because the outcome was good and the human was pleasant. CSAT charges nothing for the four minutes of grinding that preceded it. CES is the instrument built to charge for exactly that.

The limitation is the mirror image. The 2010 authors were explicit that their research "focused exclusively on contact-center interactions," and the validation data came from people who had contacted a company. CES has nothing to say about customers who never get in touch, and that group is usually much the larger one: in the independent study discussed below, Sauro reports that only 21 percent of respondents had contacted the firm at all across a two-year window. It is a service-interaction metric and it should not be lifted onto the whole relationship. That is what NPS is nominally for, with its own well-worn problems.

The evidence, read rather than repeated

CES arrived with strong claims. In the original article the authors reported a validation study of 75,000 customers who had contacted B2B and B2C companies by phone or online, said that 84 percent of those customers had not had their expectations exceeded, and set out figures intended to show that satisfaction is a poor predictor of loyalty: 20 percent of the customers they classed as satisfied said they intended to leave, and 28 percent of the dissatisfied ones said they intended to stay. On CES itself they reported that among low-effort customers, 94 percent intended to repurchase and 88 percent intended to increase spending, while 81 percent of the high-effort customers intended to spread negative word of mouth.

Those numbers are quoted throughout the CES literature. What is quoted far less is that they were published without the supporting detail needed to check the comparison, a point made directly by Jeff Sauro of MeasuringU, whose 2019 review of the metric is where the scale history above also comes from: the article did not report the corresponding figures for satisfaction and NPS, so the claim that CES beat them cannot be evaluated from the article itself.

There is independent evidence, and it is more equivocal than either the enthusiasts or the debunkers usually let on. De Haan, Verhoef and Wiesel published a peer-reviewed comparison of customer satisfaction, NPS and CES in the International Journal of Research in Marketing in June 2015, using multi-level probit models controlling for respondent self-selection. We read the published abstract on 18 August 2026. It states that top-two-box customer satisfaction "performs best for predicting customer retention," that focusing on the extremes of a scale beats using the full scale, that the best metric differs by industry and by unit of analysis, and that combining metrics improves prediction further.

The dataset is worth stating precisely, because it bounds how far the result travels. The abstract describes customers of 93 firms across 18 industries; Sauro's summary specifies 6,649 Dutch respondents and 93 Dutch firms. It is a single-market study, and a US or Brazilian reader should hold that in mind before treating it as settled everywhere.

Note also what the abstract does and does not say. It says satisfaction won. It does not, in the text we read, say that CES failed outright. Sauro's summary of the paper goes further: he reports correlations with two-year retention of .184 for top-two-box satisfaction, .170 for NPS and -.073 for CES, and describes CES as a poor predictor. We did not read the full paper, only the abstract, so we attribute those correlations to Sauro rather than to the study. We would also flag a tension inside his own account. He states that the version tested was "the CES (original five-point version)," and on that scale a high score means high effort, so a negative correlation with retention is the direction the theory predicts rather than a refutation of it. What survives regardless of how the sign is read is the magnitude. At roughly .07 against .18 and .17, CES was the weakest of the three in that dataset, and that is the claim we would carry forward.

Two further details from the same source are worth having. Only 21 percent of the respondents in that study had contacted the firm at all across a two-year window, which is the applicability limit from the previous section showing up as data. And published CES benchmarks are thin. One widely circulated figure for the revised scale, an average of about 5.5, comes from a survey vendor rather than from independent research, and Sauro notes both that the vendor "appears to be using the new scale" and that 5.5 is close to the historical average across many seven-point multipoint scales.

None of this makes CES useless. It makes it a diagnostic rather than a scoreboard. Our reading is that CES earns its place when you are trying to find friction inside a process you control, and does not earn its place as a headline number reported upward. Our chatbot KPIs guide is about that second problem.

The bot problem: the survey only reaches the people who finished

Here is the part that is specific to automation, and it is our reasoning rather than a published finding.

A post-interaction survey can only be answered by someone who was still there at the end of the interaction. For a human support conversation that is usually most people, because conversations with agents tend to end rather than evaporate. For a chatbot it is not. A meaningful share of bot sessions end because the person gave up, closed the tab and went to find a phone number, and our abandonment rate entry is about counting them.

Those people are not a random sample of your users. They are, close to by definition, the highest-effort cases you produced. If they never see the survey, your bot's CES is computed over the survivors, and it will read better than the experience it is supposed to describe. The bias is not noise. It has a direction, and the direction flatters you.

Our CSAT entry already makes the general version of this point, that a 90 percent score from 8 percent of users is not the same as a 90 percent score from half of them. What we want to add here is that for effort specifically the non-responders are not merely fewer, they are systematically the ones whose answer would have been worst. Three practical consequences follow:

  • Publish the response rate next to every CES figure, always. A CES with no denominator is a number about the people you did not annoy.
  • Read CES against abandonment, not on its own. If effort looks fine and abandonment is climbing, believe the abandonment.
  • Do not compare bot CES to agent CES as if they were like for like. The human channel has a milder version of the same bias, so the comparison is tilted before you start.

What our own deflection pages left out

We should apply this to ourselves, because we have published the diagnosis without the instrument.

Our deflection rate entry states plainly that the metric "doesn't distinguish 'bot answered well' from 'user gave up'." Our deflection versus containment entry makes the same case at length: a bot that frustrates users into abandoning looks identical to one that resolved the question, because both sessions ended without escalation. That analysis is right and we stand behind it.

The corrective both pages prescribed, until this release, was CSAT. Deflection rate recommended "CSAT as the quality floor"; deflection versus containment recommended adding a post-chat CSAT survey for a quarter and recalibrating from it. We think that is the right instinct pointed at a slightly wrong instrument. The failure mode those pages describe is a work failure, not a feeling failure. Someone who ground through four rounds of rephrasing before getting an answer has been failed in the specific way deflection hides, and on our reading satisfaction is the metric least likely to register it, because the outcome was fine in the end. CES is the metric shaped like that failure.

We have checked this rather than assumed it. Before this publishing run, the word "effort" appeared nowhere in either deflection entry, nor in our CSAT or NPS entries, and the exact phrase "customer effort score" matched no file anywhere on this site. It now matches eight, and the composition is the point: this entry, today's companion guide on customer service automation, and six pages that were amended in the same release to link here. In other words the gap was real and we closed it in the act of describing it. Across our fifteen published platform reviews, which were not amended, "deflection" appears in ten and "CSAT" in four, while "customer effort" appears in none. We have been consistently better at asking vendors about the volume metric than about either quality metric, and we have never once asked about this one.

Neither deflection page is wrong and we are not retracting them. The honest statement is that CSAT is a floor worth holding and CES is the sharper instrument for the particular gap those pages identify, and that until this release we had no entry to point them at. Both have been amended in the same release to point here, rather than left to be corrected quietly later.

Where it breaks

Reporting a number without the instrument. A CES of 5.2 is uninterpretable. Wording, points, direction and scoring rule travel with the number or the number does not travel.

Surveying everyone after everything. CES is a service-interaction metric. Firing it after a marketing message or a browse session produces answers to a question the customer did not experience.

Treating it as a loyalty forecast. The strongest independent comparison we found put satisfaction ahead of it for predicting retention. Use it to locate friction in a flow you can change.

Optimizing the survey instead of the process. Effort scores improve when you stop asking the people who struggled. Watch the response rate and abandonment or you will measure your own survey design.

Using it as the escalation trigger. By the time a bad effort rating arrives the conversation is over. Effort should be caught live, through repeat-rephrasing and escalation signals, which is what our escalation playbook is for. CES tells you afterwards whether the catching worked.

FAQ

What is customer effort score?

It is a single-question survey metric, asked straight after a support interaction, that measures how much work the customer had to do to get their issue handled. It was introduced in a 2010 Harvard Business Review article on the argument that reducing effort predicts loyalty better than exceeding expectations does.

What is the customer effort score formula?

There is no single one, which is the practical problem with the metric. On the original five-point version you take the mean of all responses and a lower result is better. On the revised seven-point agreement version, Gartner's guidance is a top-three-box calculation: the count of respondents answering 5, 6 or 7 divided by all respondents, where a higher result is better. Some organizations report the mean of the seven-point scale instead. All three are called CES.

Is a high CES good or bad?

It depends entirely on which version you are running, and in the material we have read this is the misreading that recurs most. On the original question, which asks how much effort the customer expended, a high score is bad. On the revised statement, which asks whether the company made things easy, a high score is good. Never accept a CES figure without the question wording attached.

What is a good customer effort score?

We publish no benchmark, because we have no dataset that would support one and because the benchmarks we have looked at usually do not state which instrument produced them. One widely circulated figure for the revised scale, an average around 5.5, comes from a survey vendor rather than independent research and is close to the historical average across seven-point scales generally. The comparison worth making is your own score over time, on an unchanged instrument, read next to your response rate.

CES or CSAT for a chatbot?

Both, if you can, and CES is the one more people are missing. CSAT is a good general floor and is easier to explain. CES is better matched to the specific way bots fail, which is by making people work rather than by making them unhappy. If you can only run one survey question, run the one you will actually act on.

Does CES predict whether customers stay?

Less well than satisfaction does, on the strongest independent evidence we found. A 2015 peer-reviewed comparison across 93 firms and 18 industries concluded that top-two-box customer satisfaction performed best for predicting retention. Secondary summaries of that paper report CES as the weakest of the three metrics tested. We read the abstract rather than the full paper and say so on this page.

Why does my chatbot's effort score look better than the experience feels?

Most likely because the survey only reaches people who stayed to the end. Users who gave up mid-conversation are both the highest-effort cases and the least likely to answer, so the score is computed over survivors. Check the survey response rate and the abandonment rate before believing a good CES.

Which platforms report CES natively?

We have not tested CES reporting comparatively across platforms and publish no ranking for it. The thing to check is whether the survey question wording is configurable, since a platform that only offers a fixed satisfaction question cannot run CES at all whatever else its analytics do. Our chatbot metrics guide covers assembling the measurement set, and our CSAT calculator handles the arithmetic for the satisfaction half.

Sources

  • Sauro, Jeff, PhD. 10 Things to Know about the Customer Effort Score, MeasuringU, published 6 November 2019, read 18 August 2026 — the source of this entry's scale history and of the reported correlations. Specifically: the 2010 Harvard Business Review origin under the headline "Stop Trying to Delight Your Customers"; the validation study of 75,000 customers contacting B2B and B2C companies by phone or online, the 84 percent whose expectations were not exceeded, and the 20 percent satisfied-but-leaving and 28 percent dissatisfied-but-staying figures; the low-effort repurchase (94 percent) and increased-spend (88 percent) intentions and the high-effort negative word-of-mouth figure (81 percent); the original five-point item "How much effort did you personally have to put forth to handle your request?" scored as a mean with lower better; the acquisition of the funding firm by Gartner and the change to a seven-point agreement scale scored as a top-three-box percentage with the word "effort" removed from the item; the observation that others report the mean and still others vary the number of points; the reported two-year retention correlations of .184 (top-two-box satisfaction), .170 (NPS) and -.073 (CES); the 21 percent of respondents who had contacted the firm within the two-year window; the vendor-published average of 5.5 on the revised scale and the note that this is near the general average for seven-point scales; and the authors' own statement that their research "focused exclusively on contact-center interactions." measuringu.com
  • de Haan, Evert; Verhoef, Peter C.; and Wiesel, Thorsten. The predictive ability of different customer feedback metrics for retention, International Journal of Research in Marketing, volume 32, issue 2, June 2015, pages 195-206, doi 10.1016/j.ijresmar.2015.02.004 — published abstract read 18 August 2026 via the University of Groningen research portal; the full text was not read. The abstract is the source for the study comparing customer satisfaction, the Net Promoter Score and the Customer Effort Score, the data covering customers of 93 firms across 18 industries, the multi-level probit models controlling for respondent self-selection bias, and the stated overall findings that top-two-box customer satisfaction "performs best for predicting customer retention," that focusing on scale extremes is preferable to using the full scale, that the best metric differs by industry and unit of analysis, and that combining metrics improves prediction. The abstract does not state that CES had no significant effect; that stronger characterization appears in secondary summaries and is attributed on this page to Sauro rather than to the paper. The respondent count (6,649), the Dutch nationality of the sample and firms, and the statement that the version tested was the original five-point instrument come from Sauro rather than from the abstract. research.rug.nl
  • Dixon, Matthew; Freeman, Karen; and Toman, Nicholas. Stop Trying to Delight Your Customers, Harvard Business Review, July-August 2010 — cited on this page for its title, authorship and date only. We did not read the article directly. Every figure this page attributes to it reaches us through Sauro's review, which quotes and summarizes it, and is presented on that basis. hbr.org
  • Chatbotscape. Chatbot deflection rate and deflection vs containment — the two entries quoted in this page's self-assessment section. Both were re-read on 18 August 2026 to confirm the quoted wording, and to confirm that neither used the word "effort" before the amendments described in the bullet below, which introduced it into both. The CSAT-as-corrective recommendation attributed to them was present in both prior to those amendments and remains, now alongside a pointer to this entry.
  • Site-wide string counts, re-run 18 August 2026 after this page's final edit and after the backlinks described below were added. Search string grep -ril "customer effort score" --include="*.md" . matched zero files before this publishing run and matches eight after it: this entry; its same-day companion /academy/customer-service-automation-guide; and six existing pages amended in the same release to link here, namely glossary/chatbot-csat, glossary/chatbot-deflection-rate, glossary/net-promoter-score, glossary/customer-service-chatbot, academy/chatbot-metrics-guide and academy/chatbot-escalation-playbook. A seventh existing page, glossary/chatbot-deflection-vs-containment, links to this entry without using the full phrase. Across the fifteen published platform reviews, which were not amended in this release (sample-reviews/{aisensy,blip,botpenguin,botpress,chatbase,chatfuel,intercom,landbot,manychat,sendpulse,tars,tidio,typebot,voiceflow,wati}-review.md, excluding _methodology-* and the .v1.backup file), case-insensitive matches were "deflection" in 10, "CSAT" in 4 and "customer effort" in 0. Published as a transparency statement about our own coverage, not as a claim about the platforms.
  • Ahrefs Keywords Explorer, US overview and volume-by-country, queried 18 August 2026 — the search-demand, difficulty, parent-topic and country-split figures in this entry's keyword note, including the checks behind deferring 'first contact resolution' and declining 'escalation matrix' and 'average handle time'.
  • Chatbotscape platform reviews — Intercom, Tidio and SendPulse are platforms we have evaluated hands-on and are the three carried in this entry's related-platform metadata. We have run no comparative test of survey-question configurability or CES reporting across them, and this entry therefore carries no ranking.
  • Chatbotscape evaluation methodology. /methodology (continuously updated).