Researchers have built a benchmark to find out if AI models can actually answer plain-English questions about building systems - and the early scores are rough.
The benchmark, called Build2SPARQL, pairs natural-language questions with SPARQL queries (the query language for knowledge graphs) drawn from 201 real building knowledge graphs built on the Brick and ASHRAE 223P ontologies, which model things like sensors, equipment, and zones in a structured, machine-readable format. The queries themselves come from graph-traversal code, not from a language model, so the correct answers are not dependent on any model's quirks. Only the questions are AI-generated, phrased five different ways and spanning six query patterns, from simple chains to aggregations. That process produced 6,136 verified queries and 30,680 questions, with human raters confirming the questions matched the queries nearly 99% of the time.
This matters because building automation is quietly becoming a knowledge-graph problem: more HVAC systems, meters, and sensors are getting mapped into these structured graphs so software can reason about them. The bottleneck has been a lack of training and test data for translating operator questions into queries, the same gap that slowed text-to-SQL systems for years before benchmarks caught up.
The numbers show how far there is to go. Zero-shot accuracy on three open-weight models ranged from 0.2% to 20%. Giving the models three retrieved examples pushed that to 56-65% - a real jump, but still a coin flip's distance from reliable enough to trust with a live building system.