AI/ ai · ai agents · banking · llm evaluation

Study Questions Whether AI Deck Scores Reflect Real Gains

A new study finds an AI harness for banking decks crushes a bare prompt from the same model but roughly ties a stronger model working alone.

A new study of AI-built banking pitch decks finds the fancy scaffolding mostly just catches a smaller model up to a bigger one, and even that claim is hard to pin down.

Researchers tested an agentic harness wrapped around a 27B language model, combining financial calculations, narrative templates and validation checks, built to produce the credit-decision and client-financing decks that corporate and investment banking teams rely on. Graded by five LLM judges in shared-session, text-only reviews with template markers stripped out, the full harness scored 20.4 to 33.6 points higher, out of 95, than the same 27B model writing directly from a short prompt, and it beat the plain-prompt version on all seventeen development deliverables tested. But measured against Opus generating decks directly from a short prompt, with no harness and no scaffolding, the system's edge shrank to a range of -4.7 to +0.8 points, essentially a toss-up. Repeated grading of unchanged decks also produced different scores run to run, blurring the line between genuine improvement and judge noise.

The two comparisons tell different stories, and that's the real finding. Against its own unaided baseline, the harness looks like a decisive win. Against a stronger off-the-shelf model asked to do the job cold, it's close to a wash. Judges also agreed more on which decks were decent across development rounds than on which final deck actually ranked best, a reminder that LLM-as-judge pipelines can be as fickle as the systems they grade.

Any bank tempted to copy this approach should read the fine print: the scaffolding bought real ground against a small model's cold start, but the comparison against Opus is the one worth rereading before anyone calls it a win.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →