August 21, 2026
We asked 10 AI models to translate a routine client email into Spanish. Every model handled the grammar correctly. Every model also used the informal "tú" for a message that called for the formal "usted." Adding five more models to the test didn't change that.
That gap, between what these models get right by structure and what they get wrong by judgment, is the clearest thing our testing has found on English to Spanish. Grammar converges. Register doesn't. And knowing which is which changes what you should trust a single model to catch on its own, in this direction or in the reverse, Spanish to English.
Short answer: Across a 10-model, two-round test, every model got structural tasks right: gender agreement, subjunctive mood, legal and medical terminology. Every model also missed the same thing: none used formal "usted" for a business message, in either round. Consensus fixes disagreement between models. It can't catch a mistake every model in the pool shares.
We measured average model agreement across MachineTranslation.com's ten highest-volume language pairs over a three-month window. English to Spanish came in at 83.8%, the second-lowest of the ten, despite Spanish being one of the most heavily-resourced languages any model will see. English to Hindi ranked highest, at 91.9%. Japanese to English ranked lowest, at 83.6%, just below Spanish.


According to internal MachineTranslation.com data. Agreement here does not track resourcing. Hindi and Tagalog, both less-resourced than Spanish, both rank higher. Full methodology in our model agreement study.
Resourcing predicts quality. It doesn't predict agreement. Agreement tracks how much ambiguity sits in a sentence, and Spanish has more of it than the resourcing numbers suggest, mostly around one specific kind of decision: register.
To see where that ambiguity does and doesn't show up, we ran the same 10-model, two-round test on English to Spanish that we ran on English to French, across formality, gender agreement, subjunctive mood, legal terminology, medical terminology, idiom, and ambiguity. Two structural categories came back as a clean sweep.
| Category | Spanish | French |
|---|---|---|
| Gender agreement | 10 of 10 identical | Split into 3 strategies |
| Subjunctive mood | 10 of 10 correct | Avoidance grew 1 of 5 → 3 of 10 |
Every model in both rounds handled Spanish gender agreement identically, and every model used the subjunctive correctly. French told a different story on both counts. Feminizing "manager" split three ways. Subjunctive avoidance, models switching to an infinitive construction rather than committing to the mood, got worse as we added more models, not better.

Both are grammatically correct. The difference is process versus outcome, a genuine distinction in Spanish, not an error on DeepSeek's part. This is the kind of disagreement consensus is built for: five valid readings of one sentence, resolved by which one the pool actually converges on.
Formality did not behave like gender or subjunctive. For a client-facing business message, every model in our test, all 10, both rounds, defaulted to the informal register and never produced the formal "usted."

According to internal MachineTranslation.com data. Expanding the model pool from 5 to 10 engines did not change the outcome. Every model, in both rounds, made the same register choice.
Gender agreement and subjunctive mood are structural: there's a grammatically correct answer, and more models means more chances to converge on it. Formality is a judgment call about audience and context, the kind of thing a sentence often doesn't state outright. If every model was trained toward the same default assumption about that judgment call, adding more models just adds more of the same assumption. We've started calling this a silent failure: a translation that reads as correct while quietly making a decision on the reader's behalf.
A model that gets gender agreement and subjunctive mood right on every sentence looks reliable. That reliability doesn't extend to formality, and formality is the one a Spanish-speaking business reader will notice first. Addressing a client with "tú" instead of "usted" in a first contact, a contract cover note, or a support reply reads as familiar in a context that called for formal. No model in our test corrected for it, on its own.
01. Legal and medical terminology in our test held up consistently across models. Structural risk here is low.
02. Formality does not correct itself with scale. If a translation will reach a client, a regulator, or anyone outside your organization, register needs a human check or an explicit instruction, not just a better model.
03. Ambiguity in casual or idiomatic phrasing behaves like formality, not like grammar: it needs to be surfaced, not averaged away.
SMART runs every model in the pool on a sentence and returns whichever version the majority converge on. That is exactly the mechanism that resolves the subjunctive example above: four models land on "se haga," one lands on "esté hecho," and SMART surfaces the majority reading.
It's a different mechanism for the formality miss, and worth being direct about the limit. If every model in the pool shares the same blind spot, there is no majority to diverge from, consensus has nothing to correct against. That's precisely why our AI Translation Agent exists as a separate layer: it asks directly about register, audience, and tone when a sentence doesn't state them, rather than assuming an answer the way a single pass through any one model, or all ten, would.
No single model handles English to Spanish perfectly on its own. In our 10-model test, every model got grammar-level tasks like gender agreement and subjunctive mood right, but every model also missed the same formality cue. We built SMART to run all of them at once and surface the majority answer, and our AI Translation Agent to ask about tone and audience directly when that context is missing.
Not reliably. In our test, all 10 models across two rounds defaulted to informal register for a client-facing business message and never used the formal "usted." This did not improve as we expanded the model pool from 5 to 10 engines.
We measured 83.8% average model agreement on this pair across a three-month window, near the bottom of our ten highest-volume language pairs, even though Spanish is one of the most heavily-resourced languages we translate.
Not on its own. Our test found the informal-register miss was identical across a 5-model round and a 10-model round. Consensus resolves disagreement between models with different answers. It doesn't catch a mistake every model in the pool happens to share

By Rachelle Garcia
Connect on LinkedInRachelle leads product and AI at Tomedes, where she runs the experiments that turn internal data into better translation experiences. She writes about what actually happens when you build AI products such as MachineTranslation.com — the numbers, the surprises, and the parts that don't go to plan.