A new study shows that translating benchmark questions into another language can hide the fact that a model has already seen them.
Researchers deliberately fed four open-weight instruction-tuned language models Arabic translations of two well-known English evaluation sets, MMLU and XQuAD, at varying levels of exposure, then tested the same models on the original English versions. Two standard post-hoc contamination detectors, TS-Guessing and Min-K%++, mostly failed to flag anything unusual - Min-K%++ stayed at or below chance, and TS-Guessing showed only a weak, model-specific signal on MMLU. Despite those clean readings, the models' English MMLU scores climbed as their Arabic exposure increased. The researchers also built their own diagnostic, Translation-Aware Contamination Detection, which checks whether a model's answers stay consistent across languages and across reordered answer choices, and found consistency well above what random chance would produce.
That gap is the real finding: a model can quietly absorb test content through a translated version of a benchmark and post inflated scores, while the tools built specifically to catch that behavior report nothing wrong. It is a narrow case - four models, two benchmarks, one language pair - but it points at a blind spot in how labs currently check their training data for leakage.
Most contamination debates assume leaked eval data looks the same no matter who finds it. This result says otherwise: the leak does not need to match the test's language to corrupt the score.