A new benchmark uses chess puzzles to test whether tweaking a prompt, not retraining a model, can make AI systems noticeably smarter.
Researchers built the benchmark from 1,118 Lichess puzzles to study automatic prompt optimization, the practice of rewriting instructions fed to a frozen model rather than retraining it. They ran six optimization algorithms against eight target models, grading results with exact-match scoring and chess-engine evaluation of alternative moves. The paper names the strongest model tested as "Gemini 3.5 Flash," used as the meta-model that drafts and critiques candidate prompts - that name doesn't match any model Google has publicly released, so we're flagging it as stated in the unreviewed paper rather than a confirmed, verifiable fact. Even that top performer solved only about 55 percent of the puzzles, and the researchers say the entire study cost around $800 to run.
Most LLM benchmarks go stale fast: models memorize test sets or scores saturate within months. This one is designed to regenerate itself with fresh puzzles and adjustable difficulty, which matters more than any single score, since it gives researchers a cheap, hard-to-game way to compare prompt-optimization methods as models keep improving.
Chess has been a convenient stand-in for machine reasoning since Deep Blue; here it's demoted to an $800 scorecard for prompt engineering, not a measure of raw intelligence.