AI/ knowledge-graphs · llm · sparql · semantic-web

AI System Reads Graph Data Before Writing Queries

A new neurosymbolic pipeline called QRAKEN improves natural-language-to-SPARQL accuracy by grounding queries in actual graph data instead of schema assumptions.

Researchers have built a system that stops AI models from confidently querying knowledge graphs wrong.

The problem: large language models translating plain-English questions into SPARQL, the query language for RDF knowledge graphs, often produce queries that are technically valid but misread what data actually exists in the graph. A new pipeline called QRAKEN fixes this by distilling the graph's real content first. An offline step builds TTQL, a compact summary of which multi-hop patterns, value frequencies, and example literals actually appear in the data. The LLM then uses that summary to write queries, with automated checks rejecting any query pattern the data doesn't support. On the CK25 benchmark, QRAKEN scored 30-32% better than the best comparable competitor using the same underlying models.

This matters because it reframes a classic LLM failure mode: models that know a schema's shape but not its substance. Most Text-to-SPARQL tools lean on schema definitions like SHACL to constrain output, but schemas describe what's allowed, not what's populated. QRAKEN's ablations show the empirical TTQL patterns, not the schema checks, drive almost all the improvement, beating a SHACL baseline by 64%. That's a meaningful distinction for anyone building natural-language interfaces over real-world databases, where the gap between documented structure and actual content is often where things break.

The approach also worked with two local, open-weight 35-billion-parameter models at no added cost, matching the top closed-model competitor. The researchers are upfront that this is tested on one modest-sized benchmark, and that dumping the entire TTQL summary into the prompt may not scale to sprawling, loosely structured graphs.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →