Gemma 4 pays more attention to how a claim is worded than to which document arrived first.
In a paper titled "Words Speak Louder Than Order: A Behavioral Evaluation of Gemma 4," posted to arXiv (2609.30716) on September 28, 2026, researchers fed Google's Gemma 4-e4b model pairs of conflicting documents and tracked which one it trusted. The study ran a targeted suite of 13 test items through 784 forward passes across ten counterbalanced conditions, built specifically to separate the effect of a document's framing (official guideline versus fresh update) from its reading position (first versus second). Framing won. When the two cues pointed in different directions, the model's answer tracked the source's tone far more than whether it was read first or second.
That matters for anyone building retrieval systems that hand a model two sources and hope it picks the right one. The model does have a built-in preference for whatever it reads first, but the same paper found that bias swings by a factor of 5 or more depending purely on wording, and it spikes hardest when the two documents are near-identical templates. Vary the phrasing between sources and the positional bias fades, which suggests uniform source formatting is a bigger risk than document order itself.
Translation for anyone stacking documents into a RAG pipeline: skip the officious tone if you want the model weighing evidence instead of letterhead.