A new benchmark scores XR interfaces on how fast they get you to the right feature, not on how cleverly they guess what you want.
The paper, titled 'A Benchmarking Framework for Context-aware XR Interfaces,' introduces ContextXR, which maps an XR app as a graph of 'facets' - grouped sets of related capabilities tied to a single user intent. On top of that structure, the researchers built MineXR++, a dataset that adds facet-level labels to existing XR interface data, and defined three test tasks: context factor analysis, initial facet suggestion, and next facet suggestion. Each suggestion method is scored with a simulated-interaction metric called navigation and search cost - essentially how much clicking and hunting it takes a simulated user to reach the functionality they need. The team ran that metric against three approaches: simple global-popularity ranking, relational retrieval, and LLM-based suggestion.
XR interface research has mostly relied on small, one-off user studies that are hard to compare across labs or methods. A shared graph representation plus a reproducible cost metric means different adaptation approaches - old-school popularity lists, retrieval systems, LLM suggesters - can finally be scored on the same yardstick instead of separate demo videos.
Worth remembering: this measures speed to a feature, not whether the feature was the right call. A method that's fast but wrong still posts a good score if a user can override it quickly - so treat the leaderboard as a stopwatch, not a taste test.