September 25, 2026

For everyday communication in a clinic or a hospital, yes, AI translation is safe, and a genuine improvement over having no language support at all. For anything tied to diagnosis, medication, consent, or a patient's immediate safety, we don't believe a single AI output is enough on its own. History already shows what one unchecked mistranslation can cost.
In 1980, a young man named Willie Ramirez was brought into a Florida hospital. His family, speaking Spanish, said he was "intoxicado." They meant poisoned. The staff heard "intoxicated."
He was treated for a drug overdose instead of the brain hemorrhage he actually had. By the time anyone caught the error, he was quadriplegic. The hospital settled for $71 million.
One word, misunderstood once, changed a life and cost a fortune. We've written about this case before. It's the clearest example we know of what medical translation has to get right.
A wrong word in a marketing email costs you a customer. A wrong word in a treatment room costs someone their health. That's what's actually at stake.
Comparing several AI models against each other catches more errors than trusting any one of them alone. We've built our whole platform around that belief. But comparison is a safety net, not a guarantee, and medical translation is exactly the place where the gap between the two matters most.
A medical term almost always has a close, everyday-sounding substitute that isn't the same thing clinically. The consequence of getting it wrong shows up immediately, in a treatment decision, not later in a dispute somewhere.
"Intoxicado" and "intoxicated" sound like they should mean the same thing. They don't. That gap is where the risk lives.
We've seen a version of this in our own testing. The same clinical term, translated the same way, came out correct in two languages and incorrect in a third.
Agreement between models tells you they landed on the same answer. It doesn't tell you the answer is right. In most content, that's a minor risk. In medical content, it's the whole risk.
This is why we don't ask people to simply trust the technology. We ask them to look at what it agrees on, what it doesn't, and to decide from there how much confidence the moment actually calls for.
For content that needed nuance, legal, medical, marketing, those tools just weren't enough. There was a gap between what users needed and what machine translation was delivering.
Every translation on our platform runs through multiple AI models at once. A summary called "Your translation, wrapped" shows how many models weighed in, how closely they agreed, and where they didn't.
We built it this way on purpose. A tool that hides disagreement behind one confident-sounding answer is more dangerous in a hospital setting than a tool that shows you exactly where to look twice.
That transparency doesn't replace judgment. It gives the person using the translation, a nurse, an administrator, a patient's family member, the information they need to decide whether this is a moment to trust the result or a moment to ask someone qualified to double-check it.
Below are 20 real medical phrases we've tested and published, from routine intake questions to the kind of urgent, high-stakes phrases where a mistake carries the most weight.
| Phrase | Target language |
|---|---|
| "Where does it hurt?" | Spanish |
| "Where does it hurt?" | Arabic |
| "Take one tablet twice daily" | Spanish |
| "Take one tablet twice daily" | Arabic |
| "Take one tablet twice daily" | French |
| Consent for a medical procedure | Spanish |
| Consent for a medical procedure | French |
| "You'll feel a slight pinch" | Spanish |
Where the law and the standards of the profession are concerned, yes. The U.S. National CLAS Standards, the federal framework for language access in healthcare, are direct about this: Standard 7 calls for language assistance provided by people whose competence has been established through training and certification, and it specifically discourages the use of untrained individuals in that role.
That standard exists because of cases like Willie Ramirez's.
We built Human Verification for exactly this reason. Let AI do what it does well: draft quickly, compare itself against other models, and surface disagreement before anyone acts on it.
Then let a certified professional carry the accountability that a comparison of models, however careful, cannot carry on its own. That's the whole point of building this technology responsibly.
For everyday phrases and routine communication, it's a fast, reliable first step. For anything tied to diagnosis, medication, consent, or a patient's immediate safety, we don't believe it's safe to rely on AI alone. Comparing multiple AI models catches more errors than trusting a single one, but a hospital's safety comes from certified human review at the point where a mistake would actually hurt someone.
Because a medical term can have a close, common-sounding substitute that isn't clinically the same thing, and the consequence of that substitution shows up immediately, in a treatment decision, not later in a dispute. We've seen this in our own testing: the same clinical term translated correctly in two languages and incorrectly in a third, using the same method each time.
Every translation runs through multiple AI models at once, and a summary shows how many models weighed in, how much they agreed, and where they didn't. That transparency is the point: we'd rather show you a disagreement than hide it behind one confident-looking answer.
Where the law and the standards of the profession are concerned, yes. The U.S. National CLAS Standards call for trained, competent language assistance and specifically discourage untrained individuals from filling that role. AI can support that work by drafting and cross-checking, but it doesn't replace the accountability a certified interpreter carries.
Sometimes nothing, because the error is caught. Sometimes a great deal, because it isn't. A single mistranslated word once led a hospital to treat a patient for the wrong condition entirely, delaying the correct diagnosis until the damage was permanent. That's the standard medical translation has to be held to: not "usually fine," but "checked."

By Ofer Tirosh
Connect on LinkedInOfer founded Tomedes in 2007 and now leads the company's push to combine two decades of human translation expertise with AI. He writes about where the language industry is actually heading and the shifts nobody's ready for, the bets that paid off, and the ones that didn't.