September 18, 2026

How accurate is AI translation for legal contracts?


We ran the same legal indemnification clause through ten AI models, translated separately into Spanish and French. In Spanish, every single model converged on the legally precise term: "indemnizar," to indemnify. In French, the same clause split three ways. Most models got it right. Two used weaker, non-binding phrasing. One model, Gemini, translated the clause as "will guarantee" instead of "will indemnify," a materially different legal commitment. Nothing about the French output looked wrong. The grammar was clean, the sentence read naturally, and the mistake was invisible unless you already knew what to look for. That's the real risk in translating a contract clause with AI: not garbled text, but a fluent sentence that quietly promises something different from what you signed up for.

How accurate is AI translation for legal contracts?

Sometimes, completely. The indemnification clause above is the clearest example: ten models, zero disagreement, in Spanish. Few results in cross-model testing come out this clean. But accuracy on a legal clause depends on the specific clause in the specific language pair, and that changes case by case. We've seen the same pattern show up elsewhere. A limitation-of-liability clause translated into Spanish scored 9.4 to 9.5 across four of five models, and then one model, Mistral, scored a flat 0. The four strong scores hid a deeper split. ChatGPT and Claude used the phrase "que surjan de." Qwen and DeepSeek used "derivados de." That is a scope-altering difference. A quality score that only measures fluency will not catch it.

A separate test on an "at-will employment" clause translated into German found 0% consensus on the core term. DeepSeek and ChatGPT produced functional, usable phrasing. Claude and Qwen each invoked an unrelated German legal concept that doesn't map onto the source meaning. Mistral scored 0 again. Four models, four different outcomes, on a single term in a single sentence.

Not every result works against AI. On a binding-arbitration clause translated into Japanese, the cheapest model in the comparison, roughly 60 times cheaper than the most expensive one, scored highest on the consensus metric. The pricier model, Claude Opus, actually used the more legally correct professional register, even though it didn't win the top score. The highest-scoring model in a comparison and the most legally appropriate model aren't always the same model, which is why seeing every model's output matters more than trusting whichever one scores highest.

Why do AI models disagree on legal terminology?

Legal language is precise in a way that doesn't always have a clean equivalent across legal systems. "Indemnify" carries a specific, binding meaning in contract law. "Guarantee" sounds similar and would fit the same sentence without looking wrong, but it is a different commitment. A model trained to produce fluent, natural-sounding text has no built-in reason to prefer the narrower legal meaning over the more common everyday word, especially when both fit the sentence grammatically. We ran a broader comparison across four domains: legal, technical, marketing, and academic text. The gap between models was smallest on straightforward technical writing. It was widest on legal and academic samples. That is exactly where terminology precision matters most, and where it is hardest to catch a wrong choice without a side-by-side comparison.

A garbled sentence gets caught in a five-second read. A fluent sentence that quietly swaps "indemnify" for "guarantee," or "compensate" for the legally precise "exempt from liability," looks correct because it is correct English (or French, or German). The mistake only surfaces when someone tries to enforce the clause.

This isn't a hypothetical risk. A clause can look correct and still shift financial exposure or jurisdiction. This is a recurring pattern in cross-border contract disputes. MachineTranslation.com has covered a real case where a liability clause changed meaning in translation. The resulting ambiguity became part of an expensive legal dispute. Nobody noticed at signing. That's the pattern worth planning around.

How do you check if a legal translation is accurate?

We run every translation through multiple AI models at once. A panel called "Your translation, wrapped" shows four things: how many models worked on it, what percentage agreed, how many terms they disagreed on, and how fast they reached a final answer. For legal content, that agreement percentage means more than it does for a casual message. It tells you something specific: whether the models are converging on the same legal meaning, not just similar-sounding wording.

What are examples of legal contract clauses translated by AI?

One clause needs its own explanation, since it is the test behind this article's headline finding.

ClauseTarget languageWhat the test showed
Indemnification clauseSpanishAll 10 models converged on the legally precise term, "indemnizar." The same clause tested in French split three ways, with one model materially changing the commitment to "will guarantee."

The governing law clause is the widest single test we've run: the same clause, translated across 18 language pairs, so you can see exactly where a standard piece of contract boilerplate holds up and where it doesn't.

Two Portuguese rows aren't linked to a general language-pair page above, since MachineTranslation.com splits Portuguese into Brazil and Portugal variants and the source data didn't specify which.

Do you need a certified human translator for a legal contract?

For anything you plan to sign, file, or rely on in a dispute, yes. This matches the actual standard the legal translation industry follows: ISO 20771. It is the international standard for legal translation services, and it explicitly excludes the use of output from machine translation, even with post-editing, from its scope. It's a scope boundary the standard draws on purpose, because a certified legal translator is accountable for the result in a way a model comparison isn't.

In practice: use AI for the fast, informed first pass, and to find out exactly where models disagree, the way the indemnification and at-will-employment tests above did. Then route the final version through Human Verification before it becomes the version anyone signs. A professional linguist reviews the AI output and returns it with a 100% accuracy guarantee, which is the level of accountability a contract actually needs.

FAQ

1. How accurate is AI translation for legal contracts?

It varies by language pair and clause type. We found all 10 models converged on the correct legal term for an indemnification clause translated into Spanish. For the same clause into French, models split three ways, and one materially changed the legal commitment from "indemnify" to "guarantee." The same contract can be safe in one language and risky in another.

2. Why do AI models disagree on legal terminology?

Legal terms often have a precise meaning in one legal system that doesn't map cleanly onto another, so models have to choose between a literal rendering and a functionally equivalent one. We've seen this produce a 0% consensus rate on a single term, with different models each choosing a different, non-overlapping translation, none of which was flagged as wrong by the models themselves.

3. How do you check if a legal translation is accurate?

We run every translation through multiple AI models at once. A panel called "Your translation, wrapped" shows how many models worked on it, what percentage agreed, how many terms they disagreed on, and how fast they reached a final answer. For legal content specifically, we treat anything under 90% agreement, or several disputed terms concentrated on one clause, as a signal to route that clause to a human reviewer rather than ship it as-is.

4. Do I need a certified human translator for a legal contract?

For anything you intend to sign, file with a court or regulator, or rely on in a dispute, yes. ISO 20771, the international standard for legal translation, explicitly excludes machine translation output, even with post-editing, from its scope. Use AI to get a fast, informed first draft and to catch where models disagree, then route the final version through a certified human legal translator before it's binding.

5. What happens if a contract is translated incorrectly?

Courts generally enforce the wording of the contract as translated, not what either party intended. A mistranslated liability, indemnification, or governing-law clause can shift financial exposure, change which court has jurisdiction, or weaken a party's position in a dispute, often without anyone noticing until the disagreement is already underway.

Photo of Rachelle Garcia

By Rachelle Garcia

Connect on LinkedIn

Rachelle leads product and AI at Tomedes, where she runs the experiments that turn internal data into better translation experiences. She writes about what actually happens when you build AI products such as MachineTranslation.com — the numbers, the surprises, and the parts that don't go to plan.

Share: