September 8, 2026
Quick context if you're new here: I'm the founder of Tomedes, the translation company behind MachineTranslation.com. We've been doing human translation since 2007. When AI started changing the industry, we didn't sit that out, we built this AI translation platform. That background is why I care about getting this update right.
You've always been able to see the individual models behind the consensus pick, not just the winning result. That part isn't new. What changed yesterday is what those models are labeled.
We used to show you just "ChatGPT." Now we show you "gpt-4o-mini." We used to show you "Claude." Now we show you "claude-haiku-4-5." That distinction matters more than it sounds like it should, because "ChatGPT" isn't one model, it's a brand name sitting on top of a family of very different models, and which one actually ran is the difference between a translation you can trust and one you're guessing about.
Nothing about how MachineTranslation.com picks the best translation changed. What changed is the precision of what we label it with.
Here's what that looks like on a real translation, run by someone who isn't even logged in:

Four of the eight models agree with the consensus pick word for word. The other four don't, and that split was visible before. What's new is that you can now see the split came from Mistral Small and AmazonNova's lite model, specifically, not some undifferentiated "Mistral" or "Amazon" black box. If a translation ever looks off, you know exactly which variant to blame or double-check, instead of a brand name that could mean five different things.
MachineTranslation.com still runs your text through up to 22 models behind the scenes to reach that consensus pick, using the consensus system we call SMART, the same one we've written about before. A default set of those models, drawn from providers including ChatGPT, Gemini, Claude, Qwen, Mistral AI, Mercury, AmazonNova, and DeepSeek, has always displayed automatically alongside the consensus result. What's new is that each one is now labeled with the exact model variant behind it, like gpt-4o-mini or claude-haiku-4-5, instead of just the provider's brand name. That default set is a sample, not the full 22, and which variants appear in it will shift over time as we swap in newer versions. If you want to see more than the default set, there's a plus button on the result panel. Add any additional model from the pool and its translation appears right there next to the others, labeled the same way, no separate account or API key required.

This part isn't new, but it's worth being straightforward about it now that the labels are precise enough to actually show it.
The consensus pick is built from a standard set of models and basic LLMs, including variants like gpt-4o-mini. You still get the full comparison view, the plus button, and document translation.
The consensus pick is built from the newest, most capable model variants we support, shown side by side with the individual results, the same way.
This is the part the old "ChatGPT" label actively hid, and it's the real reason the variant matters more than the brand. We've tested this directly. Running the Spanish idiom "llevarse el gato al agua" through five different ChatGPT versions at once, gpt-4o-mini and gpt-4.1-nano, the smaller, faster variants, both gave the same literal, wrong translation. The larger variants in the same family, gpt-4.1-mini, gpt-5.4-mini, and gpt-5.4, each landed on a different, correct idiomatic translation. Same company. Same brand name. Four different answers depending on which variant actually ran.
We ran the same kind of test on a German idiom across eight GPT and Claude sub-versions, and the pattern repeated: gpt-4o-mini and Claude Haiku both translated it literally, which is wrong, while the other six variants got it right, including a smaller model that outperformed a larger, older one. Model size and model age don't move in a straight line with translation quality, and "the ChatGPT model" was never specific enough to tell you which side of that line you were getting.
That's the honest version of the free-versus-paid trade-off. Both plans compare multiple models against each other, always have. The free plan's default pool leans on faster, lighter variants like gpt-4o-mini, the same variant that mistranslated the idiom above. Paid plans draw from the newer, larger variants that got it right. I run this test, and variations of it, regularly, not because I distrust the platform, but because the data it produces is the clearest argument I know for why running one model at a time is not enough. We've made this trade-off on purpose, and I've explained the reasoning behind it before: cheaper models carry real quality risks on complex or idiomatic text, and giving every user our best model regardless of plan isn't sustainable. It also obscures what professional-grade translation actually costs to deliver. Naming the exact variant, instead of hiding it behind a single provider name, means that trade-off is something you can actually see, not something we're asking you to take on faith.
A brand name isn't a specification. "ChatGPT" could mean gpt-4o-mini or GPT-5.4, and those aren't close to the same translator. One is a fast, cheap model built for casual use. The other handles idiom and legal register at a level the first one can't touch. Calling both of them "ChatGPT" and leaving it at that was never dishonest, exactly, but it wasn't specific enough to be useful either.
I've said before that confidence in a translation should come from convergence, not assumption, and that when independent models reach the same output, that convergence means something, and when they don't, the divergence means something too. That argument only holds up if you actually know which models are converging. "ChatGPT and Claude agree" tells you less than "gpt-4o-mini and claude-haiku-4-5 agree," because the second one tells you these are both fast, lightweight variants agreeing on simple text, not flagship models agreeing on something hard. We were never hiding which models ran. But a provider name doesn't tell you anything you can act on, and naming the actual variant is what makes the comparison mean something.
There's a reason model agreement is useful at all, and it's not something we invented at Tomedes. Sampling several independent model outputs and checking where they agree is a well-established way to catch an individual model's mistakes, a method researchers call self-consistency: correct answers tend to show up consistently across independent samples, while errors tend to scatter. That logic only works if you know what's actually being sampled. A consensus built from eight variants of the same underlying model isn't the same signal as one built from eight genuinely different models, and now you can tell the difference.
Yes. The default set is a starting point, not a ceiling. Anyone can use the plus button to pull in additional models from the same pool the platform already draws on, whether that's to sanity-check a specific phrase, compare a model you already know and trust, or just see more opinions before sending something to a client. Paid plans can pull from the full advanced pool, including the premium models we've tested and compared directly against each other.

If you're on the free plan, nothing about your workflow changes. You'll just see it more precisely now: not "Gemini," but which Gemini. If you're on a paid plan, the same precision applies to a stronger model pool, plus the option to compare any of the premium variants directly. Either way, the next time MachineTranslation.com hands you a translation, you know exactly which models are standing behind it, not just which companies made them.
MachineTranslation.com compares up to 22 AI models, including ChatGPT, Gemini, Claude, Qwen, Mistral AI, Mercury, AmazonNova, and DeepSeek. The consensus pick and a default set of individual model results have always displayed together; each is now labeled with the exact model variant, like gpt-4o-mini or claude-haiku-4-5, and you can add more with the plus button.
Free and unregistered users get a consensus pick built from a standard model pool. Paid users get one built from the newest, most advanced variants we support, shown the same way. Both plans work the same way; the model pool behind it is what changes.
A default set of models runs automatically on every translation. You can use the plus button to add more models yourself, on top of whatever the platform already selected as the consensus result.
No. The consensus pick still runs and is still shown first. The individual models are additional information, not a replacement, so you can see why that result won.
Models disagree most on idioms, register, and culturally specific phrasing, where more than one answer is defensible. When several land on the same wording, that agreement is a signal the translation is safe. When they split, it's worth a second look.
You can see this for yourself the next time you translate something at MachineTranslation.com, free plan or paid. If you want the fuller picture of how often these models actually agree across languages, that's a piece Rachelle (our Head of AI) wrote on the data behind it.

By Ofer Tirosh
Connect on LinkedInOfer founded Tomedes in 2007 and now leads the company's push to combine two decades of human translation expertise with AI. He writes about where the language industry is actually heading and the shifts nobody's ready for, the bets that paid off, and the ones that didn't.