Insights

Real-Time Translation and the Tone Problem: Accurate Words, Wrong Conversation

The same message rendered in two languages with different levels of formality

Machine translation has been good enough at words for some time. What it still does badly is register — and register is what decides whether a reply reads as a professional business response or as something slightly off that the customer notices without being able to say why.

This is the gap worth understanding, because a team that benchmarks translation accuracy will conclude that their setup is working while customers quietly form a different impression.

Accurate words in the wrong register still read as an off reply

What register actually is

Register is the layer that says how formal, how warm, how direct, and how committed a message is. It is carried by word choice, but also by sentence length, punctuation, contractions, and what is left unsaid.

A request for a refund and a request for a refund with an implicit threat are lexically almost identical. In one, the customer is being patient. In the other, they have already decided to escalate. Machine translation tends to produce the polite version of both, which is the wrong one half the time.

The same applies in the other direction. A reply that is technically accurate but written at a slightly wrong level of formality can read as distant in a market that expects warmth, or as inappropriate in a market that expects distance.

Three cases that survive literal translation badly

Politeness that reads as evasion. "We will consider it" is a statement about the future in English. In several languages it is a formula that means no. Translating it literally into a language where it also means no produces a reply that promises consideration the seller never intended. This is the category with the highest commercial cost, because it creates an expectation that the next message has to retract.

Legal and contractual weight. Words like "will", "may", "approximately", and "subject to" carry obligations. A translation that softens one of them does not introduce an error in the meaning — it introduces a commitment, or removes one. In technical and B2B conversations this is the risk that matters, and it is invisible to anyone reading only the translation.

Urgency. An angry customer who writes briefly is often read as calm, because brevity is neutral. In translation, brief and calm look the same. Teams that escalate on tone lose the escalation, because the machine-rendered version is too tidy to trigger anything.

What to do about it

No amount of translation quality fixes register, because register is not a vocabulary problem. Three practices help:

  • Keep the original visible to the person replying. This is the cheapest intervention and the most effective. An agent who can see the customer's own words registers the frustration that the translation smoothed away.
  • Treat certain phrases as human-only. A short list of your own high-risk phrasings — your standard refusals, your commitment language, your apology wording — written once and reviewed, used verbatim instead of translated. Ten phrases cover most of the exposure.
  • Read before sending on anything that commits you. A second pass by a person, restricted to messages that make a promise or a concession, catches the category of error that costs the most and occurs least often.

What to measure

Accuracy against a reference is hard to compute on live chat and not very informative. Two proxies are more useful:

  • The share of replies that required an edit after translation. A rising number means the model or the glossary is drifting for your traffic.
  • The share of conversations that escalate after a slow start. If translation is smoothing tone out of the early messages, escalation is happening later than it should, and this number is where it shows.

Both are cheap to collect and neither requires a reference set.

The practical summary

Translation quality on words is close to solved and will keep improving on its own. Register will not improve on its own, because improving it requires knowing what your business considers appropriate in each market, and that is knowledge only you have.

Encoding that knowledge as a short list of phrasings used verbatim is worth more than any model upgrade. For how that list fits into the wider pipeline, see AI translation quality control. For where replies get composed, see real-time multilingual support and a queue that reads the language.


Source note: This article draws on Translate AI's article "How to Translate a Conversation in Real Time (and Not Sound Like a Machine)" (https://www.translate-ai.app/articles/translate-conversation-in-real-time), which discusses real-time translation quality. The register analysis, the three high-risk categories, and the recommended practices are our own operational view. No accuracy or performance figures from the source are used as verified data here.

Share:fXintgwa
Telegram客服TG频道双向客服WhatsApp返回顶部