Multilingual Support Quality: The Metrics That Reflect Perception

Every support dashboard has a language dimension, and most of them measure the wrong thing. They count how many languages the team can staff, how many agents speak a second language, how many translation seats you bought. None of those are what a customer experiences. Intercom published survey data that makes the gap unusually concrete, and the most useful number in it is not a satisfaction figure at all.
The survey covered 170 non-native English speaking SaaS customers and 135 support team leads. Among the findings: 88% of support teams say they offer customer support in more than one language, but only 28% of end users say they actually see support offered in their own language. Those two numbers sitting next to each other are the whole argument. A large majority of teams believe they have solved languages. The overwhelming majority of customers disagree.

Treat survey numbers as survey numbers
Before going further, one caveat that applies to everything below. These figures come from a vendor survey of its own customer base, with a sample size the post does publish but a sampling method it does not. Read them as directional evidence about how a self-selected group of SaaS buyers and support leads perceive languages, not as population statistics. Any team that quotes 29% or 70% as a general fact about customers is misusing the source.
What survives that caveat is the shape of the argument, and the shape is robust: buyers who feel able to communicate in their own language report higher satisfaction and higher loyalty, and they report more tolerance for product problems and longer patience when replies are slow. That is consistent with the ordinary mechanics of support work. Translation quality is not a nice-to-have layer on top of a conversation. It changes what the conversation is.
Why language rarely shows up in your quality metrics
The reason language is missing from most quality reviews is structural rather than accidental. Reviews are usually sampled from a ticket queue, and the queue is dominated by your largest language. A team that is 70% English by volume can run a weekly quality review, find nothing wrong, and be shipping bad Japanese replies the entire time.
The fix is unglamorous: sample quality reviews per language, at a fixed volume, on a schedule. Not weighted by volume — fixed volume per language, because the point is to detect the defect in the small language, which is exactly where the defect is cheapest to fix and most expensive to ignore.
Three measurements survive this treatment.
Perceived language correctness. Not whether the reply is grammatical, but whether the customer accepted the reply as being addressed to them. The cheapest proxy is whether the customer repeats the question in the same language. If a Portuguese customer switches to English to restate a question that was answered in Portuguese, that is a language failure recorded in the transcript as a language preference.
Termbase adherence. Sample replies and check whether the terms that your team has formally agreed on were actually used. This is mechanically checkable, and it is the metric most teams can start measuring this week without buying anything. The reasoning behind why term adherence predicts perceived quality is set out in the termbase piece.
Untranslated reply rate. The share of replies sent in a language the customer did not write in, with no machine translation applied at all. In a team that just hired its first bilingual agent, this number is frequently the most actionable one on the list, because it points at a routing problem rather than a tooling problem.
Where the ceiling really is
One finding in the survey deserves separate treatment because it is a constraint no software removes. Asked what makes comprehensive multilingual support hard, the answer was hiring: 85% of support managers said it is difficult to find representatives who speak more than one language. The post also notes that once you do find one, retention is expensive and holding a consistent quality bar across both dimensions is harder still.
That is a real ceiling and it argues against plans that treat language coverage as something to be switched on. It also argues for measuring the gap between declared and perceived coverage, because that gap is where tooling can actually help: a team that cannot hire twelve bilingual agents can still make its existing agents competent in two more languages, but only if it has a termbase, a defined quality bar per language, and a review cadence. Those are exactly the three things the quality control approach on this site is built around.
What to change on Monday
Pick your three smallest languages by volume, not your three largest. Set a fixed weekly quality-review quota for each. Track perceived correctness, term adherence, and untranslated reply rate. Expect the untranslated rate to be the worst of the three on the first run, and expect that to be a routing finding rather than a hiring finding.
If you want the operational side of language routing — detection thresholds, queue ordering, escalation — it is covered in the piece on routing a multilingual chat queue. And if the conversation-level quality is fine but the messages feel wrong anyway, the failure is usually register rather than vocabulary; that is the subject of the tone problem.





