July 17, 2026
Every model I tested, all 10, in both Spanish and French, in both rounds, translated a message to a client using the informal register. Not one used usted or vous. This is the one result in the entire test that stayed constant regardless of how many models joined the pool.
I ran this same kind of test on Japanese and Vietnamese first, and formality was a real finding there too, but it was a finding that resolved: adding more models fixed a casual-register miss in Japanese. This time, in two languages that mark formality with a completely different mechanism (a distinct second-person pronoun, not an honorific verb system), the miss held steady across every model and every round.
That's the throughline of this article, and it's also the real test of AI translation accuracy across languages: Spanish and French don't share Japanese or Vietnamese's failure points at all. They introduce two the first test couldn't touch: grammatical gender, which doesn't exist as a translation problem in either Japanese or Vietnamese, and subjunctive mood, and several models avoid entirely instead of getting it wrong.
The AI models used in these tests are the latest variants exclusive to paying users on MachineTranslation.com. Testing GPT-5.5 on its own means a separate OpenAI subscription. Testing Claude Opus 4-7 means a separate Anthropic subscription. Same for Gemini 3.5-Flash, GLM-5-Turbo, and MiniMax M2.7, five separate accounts just to compare premium output for yourself. A paying user on MachineTranslation.com gets all five in the same comparison view as the baseline models, with no extra subscriptions required.
A note on why this matters beyond curiosity: automated metrics like BLEU and COMET are built to score fluency and n-gram overlap. They're not built to catch a model quietly picking the wrong grammatical gender strategy, or routing around a subjunctive clause instead of using it correctly. That's exactly the category of error this test is designed to surface.
Japanese and Vietnamese don't have this problem. Spanish and French mark gender on nouns, adjectives, and sometimes verbs, which means a translation isn't just right or wrong on meaning, it also has to agree internally across the whole sentence.

In Spanish, this wasn't a contest. Every one of the 10 models, in both rounds, wrote "La gerente revisó el contrato y lo firmó ella misma", word for word identical. "Gerente" doesn't inflect for gender, so there was nothing for the models to disagree about once they'd correctly identified the antecedent as feminine, which all of them did.
French split into three distinct, individually correct strategies:

Round 2, French: three valid lexical choices for "manager," all with correct gender agreement internally.

"Directrice" explicitly feminizes the noun. "Responsable" is gender-neutral in form and relies entirely on the article and pronoun to carry gender. "La manager", used only by DeepSeek V4-Pro, and only in Round 2, borrows the English word and marks it grammatically feminine with the French article. All three are correct French. None of them are the same sentence.
Grammatical gender in French is a place where models disagree on strategy while each stays internally consistent, not a place where they get it wrong. A single-model tool would give you one of these three answers with no indication that two other equally valid options existed, which matters if your organization has a house style preference between "directrice" and "responsable" and doesn't know it's an active choice being made on its behalf.

Spanish showed no ambiguity here. All 10 models, both rounds, used the subjunctive correctly, varying only in verb choice (entregue, envíe, presente), never in mood.
French told a different story, and it's the most interesting single result in this entire test.

In Round 1, Gemini 3.1-Pro-Preview was the only baseline model to avoid the subjunctive-triggering "que le fournisseur soumette" in favor of an infinitive construction, "recommandons au fournisseur de soumettre." Everyone else used the grammatically correct subjunctive.
In Round 2, that avoidance strategy spread. Claude Opus 4-7 and Gemini 3.5-Flash, two of the five premium variants, joined Gemini 3.1 in the infinitive-construction camp. An infinitive construction is perfectly correct French, so no model here made a grammar error. What changed is a behavioral pattern worth naming honestly: newer models didn't get better at the subjunctive. Some of them learned to write around it.
None of the models produced an incorrect subjunctive form. The finding is subtler than an error: expanding the pool increased the share of models avoiding the subjunctive construction, from 1 of 5 baseline models to 3 of the 10 in Round 2. If your use case specifically needs subjunctive-mood phrasing, for a style guide or an educational context, this is a real, measurable drift, even though every individual output is grammatically defensible.
These five categories mirror the first article's test set directly, which makes the comparison between languages, not just between models, possible for the first time.

The client-context sentence, "Hey, quick question, are you free to hop on a call?", got the informal treatment from every model, in both languages, in both rounds. Anyone using an AI translator English to Spanish for client-facing messages would hit the same miss, regardless of which model they picked. In article one, a comparable formality miss in Japanese was corrected once the pool expanded. Here, it wasn't.

This is the most operationally relevant miss in the whole article for a business audience. If you're translating client communications into Spanish or French, the model's default won't protect you here. It has to be checked manually, every time, regardless of which models are in the pool.
Every model, both rounds, both languages, translated "indemnify" correctly at the concept level. But the exact wording diverged sharply between the two languages. Spanish used "indemnizará" in all 10 outputs, a direct cognate of the English legal term, no variation. French split three ways: "indemnisera" (the majority), "devra indemniser," and "s'engage à indemniser."

One output in Round 2 is worth flagging specifically. Gemini 3.5-Flash, one of the premium variants, wrote "garantira le client" instead of any form of "indemniser." "Will guarantee" and "will indemnify" are not interchangeable legal commitments. This is the French equivalent of article one's 免責 versus 補償 problem in Japanese: the concept survives, but the specific legal term a contract actually needs doesn't always make it through, even in a model added specifically because it's newer. It's also the clearest answer this test gives on the best AI model for French translation in a legal context: none of them alone, since the miss came from a premium model, not the baseline.
Spanish's clean convergence on "indemnizará" supports a real pattern: languages that share Latin-derived legal vocabulary with English produce more consistent AI translation of legal terms than languages that don't. But French, which also descends from Latin, still fragmented, and one premium model introduced a real precision miss. Shared etymology reduces disagreement. It doesn't eliminate the need to check the specific term a contract requires.
This is the clean positive result of the article, and worth stating as directly as the misses: in both Spanish and French, GPT-4.1-nano was the lone model using an imprecise term ("daño hepático" / "troubles hépatiques", both closer to general "liver damage" than the correct "impaired liver function"), and it stayed a minority of one in both rounds, correctly outvoted by the consensus in both rounds.
Unlike the Vietnamese medical split in article one, which never resolved across either round, Spanish and French medical terminology converged correctly and stayed converged. This is what the consensus argument looks like when it works exactly as intended: one weaker model's imprecision gets outvoted instead of amplified.
"Let's circle back once the dust settles" produced a real translation failure in Spanish: Qwen 3.7-Max, in both rounds, and DeepSeek V4-Pro, in Round 2, translated "dust settles" literally as "se asiente/asienta el polvo", a calque that doesn't function as an idiom in Spanish the way it does in English.

French didn't produce a comparable error. Models split between "la poussière retombe" (a legitimate, if literal-sounding, French expression) and more generic phrasing like "les choses se calment." Both are valid communication choices. Calling this a right-versus-wrong split would overstate a finding that isn't really there.
Same test as article one: a phrase that's genuinely ambiguous in English, designed to expose disagreement instead of testing grammar.


In both Spanish and French, expanding the model pool did not uniformly preserve the ambiguity better than the baseline did. In French, GPT-4.1-nano was the lone baseline model to flatten "That's sick" into generic "C'est génial." In Round 2, two premium variants, Claude Opus 4-7 and GPT-5.5, landed on that same flattened answer, and DeepSeek V4-Pro switched away from the slang it had used correctly in Round 1.
This is exactly the case for running translations through SMART instead of a single model. GPT-5.5 and Claude Opus 4-7 are individually capable, newer than most of the baseline, and both landed on the flattened answer here with nothing in their own output to signal that most of the field disagreed with them. Seeing all ten outputs side by side doesn't make the newest model always right. It makes the disagreement visible when the newest model isn't.
GPT-4.1-nano has now flattened deliberate ambiguity in all four languages tested across both articles: Japanese, Vietnamese, Spanish, and French. Four languages is enough of a pattern to name directly, not four unrelated incidents.
Article one's throughline was that a single model's fluency isn't evidence of correctness. This one adds two failure modes that don't exist in Japanese or Vietnamese at all: models can resolve grammatical gender through genuinely different, equally valid strategies, and they can avoid a grammatical mood entirely instead of getting it wrong, and that avoidance grew more common as the pool of models expanded, not less.
The formality result is the one that should change how a business actually uses this data. It didn't improve with a larger pool, in either language, which means it's not a problem consensus solves by itself; it's a problem that needs a manual check regardless of which models are involved. The ambiguity result is the one that should change how "premium" gets interpreted: newer and more capable isn't the same as more careful, and the only way to catch the difference is to see the disagreement, not to trust the newest name in the list.
You can run any of these same sentences through MachineTranslation.com and see the full model-by-model breakdown yourself. And for the legal and medical categories specifically, where a wrong term carries real consequences, a human verification is built for exactly this kind of check.
Spanish showed more convergence across models on grammatical gender, subjunctive mood, and legal terminology, largely because Spanish legal vocabulary shares direct Latin-derived cognates with English. French produced more genuine strategy splits, particularly on gender-neutral job titles and subjunctive-triggering constructions. Neither language is uniformly easier; the failure points differ.
Yes, but they resolve it differently by language. In Spanish, all 10 models agreed on gender agreement across both rounds. In French, models split three ways between a gender-neutral noun, an explicitly feminized noun, and an English loanword, all grammatically valid, none identical. French models resolved gender agreement correctly but inconsistently, each choosing a different valid strategy.
No. Spanish subjunctive usage was already correct and consistent across all 10 models. In French, some models avoided the subjunctive-triggering construction entirely by switching to an infinitive phrasing, and this avoidance became more common, not less, once premium-tier model variants were added in Round 2.
None of the 10 models tested used the formal register (usted or vous) for a client-context message, in either language, in either round. This was the one finding that didn't improve when the model pool expanded, and it needs to be checked manually; no model in this test defaulted to formal address on its own.
Spanish legal vocabulary descends more directly from Latin, so terms like "indemnizará" have a single dominant cognate form. French legal phrasing has more competing near-synonyms, which is why models split three ways on the same clause, and why one premium-tier variant substituted a materially different legal meaning.

By Rachelle Garcia
Connect on LinkedInRachelle leads product and AI at Tomedes, where she runs the experiments that turn internal data into better translation experiences. She writes about what actually happens when you build AI products such as MachineTranslation.com — the numbers, the surprises, and the parts that don't go to plan.