AI/ tokenization · llm-pricing · multilingual-ai · research

French Text Costs Up to 58% More Tokens Than English

A study of seven major LLM tokenizers finds French pays up to 58% more tokens than English, and a prototype tokenizer narrows the gap only slightly.

French costs up to 58% more tokens than English to process through major AI models, according to a new study.

Researchers tested seven widely used 2026 tokenizers, including OpenAI's o200k, Llama 3, Qwen3, DeepSeek V3/V4, Gemma 3, Mistral's Tekken, and Anthropic's Claude generation-5 tokenizer checked via Anthropic's counting API. Using a 124-language translation set and the Universal Declaration of Human Rights, they found French needs 31% to 58% more tokens than English, while Simplified Chinese ranges from 5% fewer to 40% more and beats French on six of the seven tokenizers. France's regional and overseas languages fare worse still, running 1.6 to 3.3 times the English token count. The team then built a French-tuned prototype tokenizer, Baracoda FR v1.2, which trims 11.5% off French token counts and 3.7% off English ones compared with Tekken, at the same vocabulary size.

Token counts are not trivia. They are the unit AI companies bill by and the limit that defines a context window, so a language-based token premium becomes a real cost and capacity penalty. That gap gets worse in agentic setups that keep re-sending conversation history, and tiered pricing plus fixed context windows amplify it further for French speakers and, more sharply, for France's regional languages.

Even the researchers' own fix has limits: Baracoda performs worse on other languages, and it does not beat the existing French-specific CroissantLLM tokenizer at a comparable vocabulary size, proof that patching a tokenizer for one language rarely comes free.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →