AI/ text-to-sql · benchmarks · llm-evaluation · databases

Study Finds AI Text-to-SQL Models Barely Need Full Schema

A new benchmark shows retrieval quality matters more than schema detail for AI database queries, and exposes a memorization loophole in testing.

A new benchmark argues that the real bottleneck in AI database queries isn't how you format the schema. It's whether you find the right tables at all.

Researchers built BudgetSchemaBench, a diagnostic for text-to-SQL systems (the AI tools that turn plain-English questions into database queries). It forces models to work with a fixed budget of schema information, meaning the list of tables and columns describing a database, ranging from 2.5% to 50% of an 80-database catalog, then checks whether the resulting SQL actually runs and returns the right answer. Rather than relying on humans or other AI models to label which tables matter, the team extracted those labels directly from the correct SQL queries themselves. They also compared three ways of formatting schema text and tracked two search methods, lexical keyword matching and dense semantic retrieval, used to find relevant tables before the AI writes its query.

The standout result: a weaker, keyword-based search gained 18 percentage points in accuracy when given 20 times more schema budget, while a stronger semantic search gained just 3 points, because it already found the right tables with almost nothing to go on. Stranger still, when researchers deleted the correct tables from the prompt entirely, the AI still named the exact missing table in 94.6% of its correct answers, meaning it had memorized the schema from training data rather than reading it fresh.

That's a quiet problem for anyone grading these systems. If a model can guess a table name from memory, you're not always testing what you think you're testing. The paper's real takeaway is blunter than most schema-formatting research: the fancy serialization choices barely moved the needle, within 2 percentage points of each other, while retrieval quality moved it by double digits.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →