AI/ ai · benchmarks · multimodal-ai · finance

A New Benchmark Pits AI Against Earnings Call Analysts

MM-FinEval pairs transcripts, slides, and audio from 2,045 earnings calls to test whether multimodal AI can forecast finances like a human analyst.

A new benchmark called MM-FinEval checks whether AI models can read a quarterly earnings call the way a human analyst does, transcript, slides, and tone of voice included.

Researchers built the dataset from 2,045 S&P 500 earnings calls recorded between 2019 and 2022, pairing word-for-word transcripts with the presentation slides shown on each call and the full audio recording. Every call is scored against 12 financial forecasting tasks. The team then tested 19 existing AI models, split into three types: models that read images and text, models that listen to audio and read text, and "any-to-any" models built to handle all three inputs at once.

The standout finding is that small any-to-any models using all three modalities beat larger proprietary models limited to just two. That suggests tone of voice and slide visuals carry real financial signal beyond the transcript, information a text-only model simply never sees.

It is a useful corrective to the current habit of judging AI's financial reasoning on text-only question sets, which miss the hedging and visual cues that human analysts rely on every earnings season.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →