Researchers have built a benchmark that hunts for the exact questions a language model gets right about half the time, rather than the ones it always nails or always flubs.
The approach, called Dynamic Boundary Evaluation, starts from a complaint about standard AI benchmarks: give every model the same fixed set of questions and you mostly learn which models ace everything and which fail everything. DBE instead hunts for each model's boundary - items where its chance of passing is close to a coin flip - using a search method called Skill-Guided Boundary Search that needs only API-level access to the model. The researchers built a difficulty-calibrated item bank covering safety, capability, and truthfulness, validated against nine reference LLMs. They tested it across four categories: refusing harmful requests, avoiding over-refusal of safe ones, following constrained instructions, and resisting sycophancy over multi-turn conversations.
Fixed benchmarks hit ceilings and floors fast, especially as models improve at handling familiar test sets, so the industry keeps leaning on scores that no longer separate real capability gaps. A boundary-seeking method that adapts per model, and grows its item bank when a model falls outside existing coverage, offers a way to keep comparing models on a shared scale even as the field moves past today's hardest benchmarks.
It is still an evaluation of evaluations, not a cure for the deeper problem: benchmarks only measure what someone thought to test for.