AI/ ai · patents · llm-evaluation · legal-tech

Study Tests Whether AI Judges Can Vet Patent Drafts

A benchmark called Vibe Patenting shows AI judges help cheaper models draft better patents, though scores don't always match a human attorney's verdict.

Researchers built a benchmark to test whether an AI can reliably judge its own patent applications, and the results are a qualified yes.

A new paper introduces "Vibe Patenting," a testbed where an AI agent drafts a patent application and a separate large language model acts as judge, scoring the draft and returning structured feedback for revision. Across multiple inventions and several drafting-agent setups, that judge-guided revision loop consistently improved judge-assessed quality, while agents left to revise on their own quickly plateaued. Notably, enough rounds of judge feedback let a cheaper, low-reasoning model approach the performance of a pricier, high-reasoning one. The researchers also checked their AI judge against a professional patent attorney's independent evaluation and found real but inconsistent agreement, with some scoring metrics tracking the attorney's judgment far better than others.

Patent drafting is exactly the kind of high-stakes, rules-heavy writing that AI vendors pitch hardest and lawyers trust least. If a judge model can coach a cheap drafting model toward attorney-grade output, that is a genuine cost lever for firms. But when the judge and the human reviewer disagree on what counts as "good," you have automated confidence, not necessarily quality.

The paper's own data undercuts easy optimism: agreement between the AI judge and the attorney varied sharply by metric, the same calibration problem that has dogged AI grading in coding and writing benchmarks for years.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →