September 8, 2026
Run a client-facing business message through AI translation and the grammar will very likely be correct. The tone might not be. In one internal test, a message meant for a new client was translated into Spanish and French across ten different AI models, in two separate rounds. Every model, in both languages, in both rounds, defaulted to the informal register, "tú" instead of "usted," "tu" instead of "vous," for a message that called for the formal one. Widening the comparison from five models to ten didn't fix it, because every model in the pool made the same call. That's not a translation error in the traditional sense. The words were right. The relationship was wrong.
Formality isn't stated in most English source sentences, because English doesn't force the choice the way Spanish, French, German, and Japanese do. "Please review the attached proposal" carries no grammatical marker for formal or casual address. A model has to infer it, and in testing, every model inferred the same way: casual by default. Structural grammar wasn't the problem. The same test found every model correctly handled gender agreement, subjunctive mood, and legal or medical terminology in the identical message. It was specifically the human judgment call, the one an experienced assistant would make without thinking twice, that every model got wrong, together.
Once you get past the formality question, individual business phrases score well, and it's worth seeing where. "Please find attached," translated into Spanish, produced near-identical results from two leading models, ChatGPT scored 9.5 and Claude scored 9.4, differing by a single dropped comma. A separate model diverged with an unusual word order, and the platform's own agreed pick correctly set that version aside. On a French version of the same idea, "Veuillez trouver ci-joint notre rapport annuel" ("Please find attached our annual report"), five models reached full, 100% agreement with zero disputed terms in under seven seconds. Simple, self-contained business phrases are close to a solved problem.
Job titles are where that confidence should drop. Testing the French sentence "Elle est directrice adjointe du département commercial" across five engines produced four different English readings: "deputy director of commercial department," "assistant director," "deputy director of the sales department," and one model that returned no usable output at all. If that sentence were sitting in an email signature or an org chart you were localizing, the model you happened to pick would have quietly changed someone's job title.
You don't need to run your own ten-model test to catch this. Every translation on MachineTranslation.com opens a panel called "Your translation, wrapped" that shows how many AI models worked on that specific translation, what percentage agreed, how many terms they disagreed on, and how fast they reached a final answer. In one live example, six models worked on a translation, reached 92% agreement, disagreed on four terms, and reached a final answer in 1.6 seconds, with a short breakdown showing which model ran most concise, most thorough, most formal, and most natural on that specific text.

A 92% agreement rate with a handful of disputed terms is a reasonable candidate to ship as-is for routine business content. Lower agreement, or several disputed terms concentrated in one sentence, is the signal to route that specific line to a human reviewer instead of the whole document.
That last point is the practical takeaway for a marketing or ops team using this day to day: treat the agreement percentage as a triage signal, not a pass/fail grade. High agreement on a routine confirmation email is fine to send. A formality question, a job title, or a legal-adjacent phrase like "without prejudice" sitting inside an otherwise high-agreement email is exactly the kind of single line worth a second look before it goes out under your company's name.
Three of these are worth calling out on their own, because each one produced a genuinely different result when we tested it live.
| Phrase | Target language | What the test showed |
|---|---|---|
| "Please find attached" | Spanish | ChatGPT and Claude scored within a single dropped comma of each other (9.5 vs. 9.4). One model diverged with an unusual word order, correctly excluded from the agreed pick. |
| "Please find attached our annual report" | French | The cleanest possible result: full agreement across every model tested, zero disputed terms, under seven seconds. |
| "Deputy director, commercial department" | French | Four different readings across five models: "deputy director," "assistant director," "deputy director of the sales department," and one model that returned nothing usable. |
The remaining phrases below cover the everyday range of business correspondence: requests, confirmations, follow-ups, and a handful of common travel questions for when the email turns into an in-person meeting.
ON-THE-GO TRAVEL PHRASESPhrase | Language pair |
|---|---|
| "How much does the trip cost?" | Spanish → German |
| "What time is checkout?" | Spanish → French |
| "What time is check-in?" | French → Portuguese |
| "Where can I buy tickets?" | French → Portuguese |
For routine, high-agreement correspondence, an AI translation is a fine send-as-is. For a first message to a new client, anything contract-adjacent, or a message that sets the tone for a relationship you want to keep, run it past a native speaker before it goes out, specifically for register. That's the one category of mistake a model comparison alone won't reliably catch, because the whole pool can share the same blind spot. We route anything at that level of stakes through Human Verification: a professional translator reviews the AI output and returns it with a 100% accuracy guarantee, on top of whatever the models already agreed on.
For self-contained phrases like "please find attached" or "to whom it may concern," yes, we found models agree closely and the output is reliable. The risk shows up in full client-facing messages, where models can get every word right and still default to the wrong level of formality, which reads as a tone problem rather than a translation error.
We ran a client-facing message across ten AI models, in two rounds, and every model defaulted to the informal "tú" in Spanish and "tu" in French, never producing the formal "usted" or "vous" the message called for, even after we expanded the pool from 5 to 10 engines. Formality is a judgment call the source sentence didn't state outright, and if every model in a comparison shares the same blind spot, comparing more models doesn't fix it.
We run every translation through multiple AI models at once, and a panel called "Your translation, wrapped" shows how many models worked on it, what percentage agreed, how many terms they disagreed on, and how fast they reached a final answer, so disagreement is visible instead of hidden inside one confident-looking result.
Use it for the first draft and to catch obvious wording issues fast. For a first message to a new client, a contract-adjacent email, or anything that sets the tone for an ongoing relationship, we'd recommend having a native speaker confirm the register before you send it.
"Tú" is the informal "you," used with friends, family, or peers you already know casually. "Usted" is the formal register expected in professional correspondence, especially with a new client or a superior. Business emails in Spanish default to "usted" unless the recipient has explicitly invited a more casual tone.

By Rachelle Garcia
Connect on LinkedInRachelle leads product and AI at Tomedes, where she runs the experiments that turn internal data into better translation experiences. She writes about what actually happens when you build AI products such as MachineTranslation.com — the numbers, the surprises, and the parts that don't go to plan.