AI email assistants do less work simply because you asked nicely instead of bluntly.
A new study tested a retrieval-augmented generation pipeline and two tool-using email agents by holding the underlying request, available evidence, and expected outcome fixed while varying how the request was phrased. Researchers built validated variants across five communication-style dimensions and four rule-based English dialect conditions. Indirect requests hurt performance on all three systems, and formal phrasing hurt both agentic systems. The causes differed: verbose wording mainly confused the lexical retriever by burying the relevant email, while indirect and dialect variants still caused failures even when the right email was found, mostly by making agents skip required actions rather than take wrong ones.
This matters because almost every email-agent benchmark tests one canonical phrasing per task, so this failure mode has been invisible. A demo that nails a clean, direct prompt can still quietly drop tasks when a real user hedges, writes formally, or just talks differently. That gap is the difference between a tool that looks reliable in a sales deck and one that is reliable in an actual inbox.
It is a useful reminder that a correct-sounding reply is not the same as a completed task, and that most agent benchmarks are still grading the easy version of the test.