Skip to content
AI Atlas
PaperActive

SalamandraTA at WMT 2026 Terminology Shared Task: Hard Examples Are Better Teachers

arxiv.org/abs/2609.09999

quality89

Updated 5 h ago · first seen 11 Sept 2026

paper_01M294G6DAK3F0TY94BVJ0GTKQ

Published
11 Sept 2026
T1 · 5 h ago
arXiv
2609.09999
T1 · 5 h ago
Category
cs.CL
T1 · 5 h ago

Abstract

Terminology-aware translation asks for more than a correct translation: the output must use the exact terms a glossary prescribes. The standard recipe, fine-tuning on glossary-annotated translation pairs, hides an inefficiency: for most examples the glossary prescribes exactly what the model would have produced anyway, so they teach nothing about following a glossary. We therefore keep only the examples where the model's own translation contradicts the glossary. In a controlled study at fixed data volume, this selection alone raises term accuracy from 78.7% to 89.9%. The filtered data, built by a two-way synthetic pipeline on open models, is part of the instruction-tuning mixture of our public release SalamandraTA-7b-instruct v3.0, which, used exactly as released and wrapped in a document-level inference pipeline, forms the BSC submission to the WMT26 Terminology Shared Task Track 1. At the official WMT26 evaluation, our system achieves 94.2% term success at 74.6 chrF++, with only two of the twenty-two submissions outperforming it on both metrics. On last year's benchmark, it also surpasses our GRPO-based system, despite being trained solely with ordinary supervised fine-tuning.

Authors 2

Xixian Liao, Maite Melero

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

Arxiv announce type
replace

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

arXiv id
2609.09999

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

Categories
cs.CL

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

Primary category
cs.CL

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

Published
11 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 5 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

5 h ago

Conflicts

None