AI/ ai · proactive-assistants · machine-learning · on-device-ai

Small Graph Model Beats LLMs at Knowing When to Interrupt You

A tiny graph model predicts when proactive AI assistants should speak, beating nine rival architectures on accuracy while running far faster.

A team of AI researchers built a small graph-based model that decides when a proactive assistant should interrupt you, and it beats large language models at the job while using a fraction of the compute.

The paper, posted to arXiv, describes a temporal-graph-learning (TGL) controller that treats a user's activity, such as emails, messages, and calendar events, as a graph, where entities like people and projects recur over time and link separate interactions together. That structure lets the model predict, in a single forward pass, both whether a moment is worth interrupting for and which entities should inform the resulting suggestion. The researchers tested it against nine other trigger architectures, including two LLMs that make the call in a single forward pass, using AUC (area under the ROC curve), a standard measure of how well a model separates true interrupt-worthy moments from ones best left alone. TGL posted the highest AUC of all nine, processing each event in 11.13 milliseconds on a GPU server, which the authors clock at 4 to 7 times faster than the two single-forward LLM triggers it outranked.

That combination, more accurate and much cheaper, undercuts the assumption that proactive AI needs a language model running in the background just to decide whether to speak. A shared TGL controller also lifted F1 scores by an average of 16.7 points across 14 different downstream suggestion-generating models, meaning the gains hold regardless of which LLM actually writes the final suggestion.

It even runs on a laptop at under 14 milliseconds with a 220 MiB footprint, which is the real pitch here: proactive assistants that don't need a data center just to know when to shut up.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →