Insights

AI Translation Quality Control: What to Check When a Human Cannot Read Everything

A review queue separating routine messages from commitment-bearing ones

Once a team runs customer conversations through machine translation, a question appears that has no good default answer: how do you check quality when nobody on the team can read every language you serve?

The instinct is to sample randomly and review. That produces a number nobody can act on. The better approach is to review by risk — to accept that most messages do not need checking, identify the small set that do, and spend the human attention there.

Quality control that samples randomly spends attention where nothing is at stake

Why random sampling fails as a control

Random sampling measures the average. Average translation quality on customer chat is high and not interesting, because most messages are greetings, logistics, and confirmations. The messages that damage a relationship are rare by construction — they are rare because they are the exception.

A sample of fifty messages will contain one or two consequential messages, and there is no way to know in advance which ones those were. So the review either misses them, or you escalate to reviewing everything, which is the same as not having a control.

The fix is to stop sampling at random and start classifying by consequence.

Four categories, and only one of them needs a human every time

Category 1 — routine. Greetings, opening hours, order status, shipping estimates, confirmations. Let these pass. Reading them costs more than the errors are worth.

Category 2 — commitment-bearing. Anything that states a delivery date, a price, a refund, a specification, or a commitment to do something. These get a human read before sending. This category is small — often under a tenth of traffic — and it contains nearly all the commercial risk.

Category 3 — emotional. Messages where the customer is frustrated, and messages where they are being told no. These get flagged for awareness rather than review. The agent is not checking the translation for accuracy; they are being told that this conversation has already gone badly, which changes how carefully they answer it.

Category 4 — unknown. Anything the system is not confident about. This is the category that grows silently, and it needs a weekly look at its size. If it is not shrinking, the confidence threshold is wrong.

Build the commitment-bearing list from your own history

The most useful artefact here is not a glossary. It is a list of the phrasings your business actually commits to, written once, in the original languages, and used verbatim.

In practice this is short: your standard shipping estimate, your refund conditions, your warranty statement, your three most common refusals, your apology wording. Ten to twenty phrases. Once they exist, the agent composes from them rather than translating the phrase that came to mind, and the largest remaining category of commercial error largely disappears.

The second artefact is a monthly review of the unknown category, asking one question: what did the system not understand, and is that a gap in vocabulary or a gap in the model? Vocabulary gaps are fixed by adding terms. Model gaps are not fixed by adding terms, and treating them as vocabulary gaps is how a glossary grows to four hundred entries and the error rate does not move.

What the numbers mean

Once the categories exist, three numbers are worth reporting:

  • Commitment-bearing volume as a share of traffic. A rising share usually means a change in customer mix, not a change in quality.
  • Edit rate within commitment-bearing messages. This is your real quality number, and it should be stable or falling. A jump means something changed — a new model, a new market, a new phrasing.
  • Unknown-category size. The one that tells you whether to keep tuning the threshold or stop.

None of these require a reference corpus, a linguist, or a scoring rubric. All three are counts you can produce from data you already generate.

Where this sits

Quality control is the last stage of a pipeline that begins with detection and ends with a human read. For the stage before it, see real-time multilingual support. For the specific errors that survive good word-level accuracy, see the tone problem in real-time translation. For the routing that decides who sees the message in the first place, see a queue that reads the language and shared inbox ownership.


Source note: This article draws on The Translation Team's article "AI Translation Quality: Speed Without Losing Your Message" (https://www.thetranslationteam.com/blog/ai-translation-quality), which discusses AI translation quality. The four-category control scheme, the verbatim commitment list, and the reporting set are our own operational design. No quality figures or vendor claims from the source are used as verified data here.

Share:fXintgwa
Telegram客服TG频道双向客服WhatsApp返回顶部