October 3, 2026 · 15 min read
For Professionals: AI Translation Accuracy Backed by COMET and MQM
Research backed primer for professionals on AI translation. Why COMET and MQM matter, key failure modes, and a human in the loop checklist to deploy safely.

AI translation now produces fluent, often accurate output for many high-resource language pairs and every day, low-risk text. For clinical, legal, or other high-stakes content, though, a hybrid human-in-the-loop workflow remains the safe standard. The right way to judge any system is through semantic-fidelity metrics like COMET and structured human review (MQM), not fluency alone.
TL;DR:
- Large language models typically outperform dedicated neural machine translation systems on meaning-focused metrics like COMET, especially after direct translation training.
- Automatic metrics such as BLEU can be misleading because they reward word overlap without assessing accuracy of meaning, making human evaluation essential for high-stakes content.
- Clinical, legal, and speech input translations require human review, as errors in these domains can lead to serious misunderstandings despite high fluency scores.
- Fine-tuning models directly on translation pairs and implementing human-in-the-loop review significantly improve accuracy and reduce terminology drift in practical workflows.
- Before adopting an AI translation system, conduct domain-specific testing with actual content, verify metric transparency, and ensure support for glossary and human review integration.
Table of Contents
- How translation accuracy is measured: BLEU, COMET, and MQM
- What recent benchmarks reveal about LLMs versus traditional systems
- Domain and language gaps: clinical, legal, and speech-based accuracy
- Common failure modes: why fluent doesn’t mean correct
- Improving accuracy: fine-tuning, glossaries, and human-in-the-loop workflows
- A practical checklist before you adopt an AI translation tool
- Where AI translation is heading, and what that means for how you plan
- How Live Caption AI handles domain accuracy for live events
- FAQ
- Sources
How translation accuracy is measured: BLEU, COMET, and MQM
For years, BLEU was the default yardstick for machine translation. It counts overlapping words and phrases between a machine output and a reference translation. The problem: BLEU rewards matching text, not matching meaning. A sentence can score poorly on BLEU while being a perfectly accurate translation, and it can score well while missing the point entirely.
COMET works differently. It uses neural embeddings, and increasingly LLM-based judges, to compare the meaning of a translation against the source and reference, not just the word choice. This makes it far better at catching the kind of error that actually confuses a reader, like a flipped meaning or a dropped qualifier, even when the sentence still reads smoothly.
MQM (Multidimensional Quality Metrics) goes a step further. Instead of a single score, trained annotators tag specific spans of text with specific error categories: mistranslation, omission, terminology, fluency, and more, each weighted by severity. It is slower and more expensive than automatic scoring, but it is the closest thing the field has to ground truth. A large MQM-based re-evaluation of WMT systems found that human translations were still preferred over machine output once errors were scored this way, and that accuracy and mistranslation errors, not fluency, drove most of the gap.
What this means in practice:
- For casual or internal content, BLEU-style scores or a quick fluency check are usually enough.
- For customer-facing or brand content, COMET-style semantic scoring catches meaning errors fluency checks miss.
- For clinical, legal, or safety-critical content, nothing substitutes for MQM-style human review by a domain-qualified bilingual reviewer.
Knowing which metric a vendor quotes tells you a lot. A vendor that only cites BLEU may be hiding a system that reads well but drifts in meaning.
What recent benchmarks reveal about LLMs versus traditional systems
The clearest recent evidence comes from the WMT25 General Machine Translation Shared Task, the annual benchmark the field uses to compare translation systems head to head. It confirmed a pattern researchers have been tracking for a couple of years: large language models tend to beat dedicated neural machine translation (NMT) systems on semantic-fidelity metrics like COMET, while NMT systems sometimes hold their own, or win, on lexical metrics like BLEU.
One submission to WMT25 reported a notably higher COMET score for its LLM-based system compared to a strong NMT baseline on a comparable benchmark, a gap large enough to matter in real deployments, according to DLUT and GTCOM’s WMT25 system report. That same report found that fine-tuning an LLM with direct translation supervision, training it straight on source-to-target pairs, consistently outperformed a refinement strategy where the model first drafts and then edits its own output. The gain held across different model sizes and language pairs, which suggests it is a reasonably robust finding rather than a quirk of one setup.
This BLEU-COMET dichotomy matters because it shows two systems can disagree sharply depending on which yardstick you use. A system optimized for word-level overlap can look weaker on COMET even when it is, by some measures, more literal. A system that scores high on COMET can occasionally paraphrase more loosely than a strict terminology-locked workflow wants.
The WMT25 task organizers added an important caveat: automatic scores can be biased by the test data itself, and human evaluation sometimes reverses the rankings that automatic metrics produce. Speech-domain inputs, transcripts with disfluencies, false starts, and overlapping speech, proved especially hard for every system tested, and performance on low-resource language pairs lagged well behind high-resource pairs like English-Spanish or English-French. The organizers’ core recommendation was blunt: don’t trust a single automatic score to settle a claim about accuracy. Run human evaluation, especially for anything domain-specific or safety-relevant.
A separate WMT24 metric-reliability analysis reinforces this. It found that embedding-based automatic metrics sometimes agreed better with professional judgment than crowd-sourced human ratings did, but professional MQM evaluation still surfaced errors that neither automatic scores nor casual human raters caught. The takeaway isn’t that metrics are useless. It’s that no single number tells the whole story, and the number that matters depends on what you’re translating.

Domain and language gaps: clinical, legal, and speech-based accuracy
Accuracy is not one number. It changes with the domain, the language pair, and whether the source is written text or messy speech.
The sharpest evidence comes from medicine, where the cost of an error is highest. A comparative analysis of machine-translated patient discharge instructions tested GPT models and Google Translate on real clinical documents. But sentence-level accuracy hides a problem.
In other words, a system can get almost every sentence right and still produce a document that tells a patient the wrong dosage, because one bad sentence out of twenty is enough to cause harm. The study found error rates varied meaningfully by language, with some languages, including Russian, showing more clinically impactful mistakes than Spanish or Chinese.
Why does domain and language matter this much?
- Resource availability drives quality: high-resource pairs like English-Spanish have far more training data than lower-resource pairs, and accuracy tracks that gap closely.
- Terminology density matters: medical, legal, and financial text packs specialized vocabulary into short spans, so a single mistranslated term can change the entire meaning of a sentence.
- Speech-derived input introduces disfluencies, interruptions, and ambiguous phrasing that written-text training data doesn’t prepare a model for, which is part of why WMT25 flagged speech-domain robustness as a persistent weak spot.
A practical rule follows from this: AI-only translation is reasonable for low-stakes, high-resource-language content where an error causes minor confusion at worst. Human review is not optional for clinical instructions, legal notices, financial disclosures, or any content in a lower-resource language pair, regardless of how fluent the output reads.
Common failure modes: why fluent doesn’t mean correct
The most dangerous AI translation errors are the ones that look fine. Researchers and professional translators call this illusionary fluency: an output so smooth and grammatically correct that a reader with no knowledge of the source language has no way to detect that the meaning has shifted. The American Translators Association has flagged this as the primary risk of relying on AI translation without a review step, because the very thing that makes a translation easy to trust, fluency, is exactly what masks a subtle but serious error.
Common patterns to watch for:
- Terminology drift: a specialized term gets translated as a generic synonym that loses its precise meaning.
- Dropped negation or modifiers: words like “not,” “only,” or “except” get lost in longer sentences, reversing the intended meaning.
- Wrong units or measurements: dosages, dates, and currency amounts get mistranslated or left in the source format.
- Source copying: the system leaves a phrase untranslated or repeats source-language text when it is uncertain.
- Gender and agreement errors: pronouns or adjectives disagree with their subject in languages with grammatical gender, which can distort meaning or sound disrespectful.
The consequences scale with the stakes. In marketing, a dropped modifier is embarrassing. In medicine, it can mean a patient misreads a dosage instruction. In legal or financial documents, a single mistranslated clause can change an obligation entirely.
Pro Tip: Have a bilingual reviewer read translated output aloud alongside the source, sentence by sentence, rather than skimming for overall fluency. Errors hide in details a fast read skips over.
Improving accuracy: fine-tuning, glossaries, and human-in-the-loop workflows
Three levers reliably move accuracy in the right direction: better fine-tuning, stricter terminology control, and a human review step. None of them work in isolation.
On the technical side, the DLUT/GTCOM WMT25 report found that fine-tuning a model directly on source-to-target translation pairs, rather than training it to draft and then refine its own output, produced more consistent accuracy gains across model sizes and language pairs. That’s a useful detail for anyone evaluating a vendor’s technical claims: ask whether their system is trained for direct translation or for a draft-then-refine pipeline, since the former has stronger benchmark support.
On the workflow side, the evidence for human-in-the-loop review is substantial. A 2025 Nature study of human-in-the-loop strategies in medical translation found that combining machine translation with human review cut total task time by more than half compared with human-only translation, while also scoring higher on quality. That’s a strong signal that the hybrid model isn’t a compromise. It can outperform either extreme.
A practical deployment checklist:
- Build a domain glossary of terms that must translate consistently, and lock those terms so the system can’t substitute synonyms.
- Require a domain-qualified human post-editor for any content above your defined risk threshold.
- Run periodic MQM-style sampling on live output to catch drift before it becomes a pattern.
- Feed corrected samples back into fine-tuning cycles rather than treating review as a one-way filter.
For live spoken content specifically, a workable pipeline looks like: live audio capture, on-device prefiltering to clean up disfluencies, model translation constrained by a locked glossary, and a domain-aware pass before anything gets published or archived. We built Live Caption AI around this logic for real-time settings: our Medical, Worship, and Finance vocabulary models, plus the ability to add your own Custom Terms, exist specifically to reduce the terminology drift that generic speech-to-text tools introduce, with translation available into one of 29 languages per session. For teams managing the post-editing side of a workflow like this, general audio post-production collaboration practices offer useful structure for coordinating reviewers across a live or recorded pipeline, and our own guide to auto-translating live captions walks through pairing automated output with a human check for events.
A practical checklist before you adopt an AI translation tool
Before trusting any system with real content, run a pilot test rather than taking a vendor’s accuracy claim at face value.
- Test on your own representative data, not a vendor’s demo sentences. Use actual documents or transcripts from your domain.
- Ask which metrics they report. COMET or MQM-based evaluation is a stronger signal than BLEU alone.
- Request language-pair-specific evidence. A system strong in English-Spanish may perform far worse in a lower-resource pair you actually need.
- Confirm glossary or custom-terminology support. Without it, terminology drift is close to guaranteed over time.
- Check whether human-in-the-loop review is built into the workflow or bolted on as an afterthought.
- Ask about data retention and privacy, especially for clinical, legal, or financial content, since audio and text retention policies vary widely across vendors.
Set different acceptance thresholds for different risk levels. High-risk content, like clinical instructions or legal notices, should get full human review until you have enough historical data to trust a narrower sampling rate.
Pro Tip: Ask a vendor directly whether their published accuracy numbers come from BLEU, COMET, or MQM evaluation. If they can’t answer clearly, treat that as a red flag about how rigorously the system has actually been tested.
Where AI translation is heading, and what that means for how you plan
Progress in AI translation is incremental, not sudden. Fine-tuned LLMs keep closing the gap on semantic-fidelity benchmarks, and evaluation methods are maturing alongside them, but every benchmark gain still comes with caveats about speech robustness and low-resource languages that won’t disappear overnight.
The practical lesson is to invest in your own evaluation process rather than chasing a vendor’s parity claim. A system that beats another on a published benchmark may still fail on your specific domain, your specific terminology, or your specific language pair. Build your own small test set, score it with COMET or MQM-style review, and revisit that test periodically as models change.
One more consideration deserves attention: as high-resource languages keep improving faster than low-resource ones, the gap in translation quality between them can widen rather than close. Anyone deploying AI translation at scale should watch that gap deliberately, not assume it will fix itself.
— Ryan
How Live Caption AI handles domain accuracy for live events
Live, spoken content is where translation accuracy gets hardest to control, and it’s exactly where we focus. Live Caption AI turns any phone into a caption receiver through a QR code, so there’s no hardware to install and no stenographer to book.

Our vocabulary models are built to catch terminology that generic speech-to-text tools routinely miss, and you can add your own Custom Terms on top of them for anything specific to your organization. Translation covers multiple languages per session, and session audio is never stored.
This fits a range of settings where live accuracy matters:
- Clinical and professional meetings, where a mistranslated term carries real consequences.
- Worship services, where congregations often include multiple language communities in the same room.
- Conferences and events, where attendees expect captions without renting interpretation equipment.
Our Professional plan is built for exactly these settings, and full plan details, including the Free tier and Business pricing, are on our pricing page.
This article is general information, not a substitute for advice from a qualified doctor. Consult a qualified healthcare professional about your own circumstances before acting on anything here.
FAQ
Which AI is most accurate for translation?
No single system wins across every language pair and domain: accuracy depends heavily on the specific pair and content type being translated. Recent benchmark work shows large language models often score higher on semantic-fidelity metrics like COMET, while traditional neural machine translation systems can still compete on lexical metrics like BLEU, according to WMT25 shared-task findings.
What are the downsides of AI translation?
The main risks are illusionary fluency, where an error-filled translation still reads smoothly, along with terminology drift, dropped negations, and weaker performance on lower-resource languages and speech-based input. These patterns are documented across WMT25 and in clinical translation research, which found clinically impactful errors even when sentence-level accuracy looked strong.
How accurate is ChatGPT in translating?
In a clinical comparative analysis, a GPT model reached roughly 97% sentence-level accuracy for English-to-Spanish and 95% for English-to-Chinese discharge instructions, but instruction-set level inaccuracies still appeared in a meaningful share of cases depending on the target language, according to this comparative analysis. That gap between sentence-level and document-level accuracy is why human review still matters for clinical use.
Is AI better at translation than Google Translate?
The same clinical study tested Google Translate alongside GPT models and found both reached high sentence-level accuracy for Spanish and Chinese, with similar instruction-set level limitations across languages, per this research. Neither tool eliminated the need for human review on clinically significant content.
How is translation quality actually tested?
Quality is measured with a mix of automatic metrics, like BLEU for word overlap and COMET for semantic similarity, and structured human evaluation methods like MQM, which has trained reviewers tag specific errors by type and severity. Human evaluation remains the most reliable check, since MQM-based studies have found it can reveal accuracy gaps that automatic scores miss.
Sources
- Nature article on human-in-the-loop strategies in medical translation
- Evaluation of the accuracy and safety of machine translation of patient-specific discharge instructions: a comparative analysis
- Findings of the WMT25 General Machine Translation Shared Task
- MQM-based analysis of MT vs human translations