AI tools that summarize earnings calls are decent at showing their work, but they still get facts wrong.
Researchers built a benchmark called ECTs-100, drawing transcripts from the 100 largest constituents of the S&P 500. They designed a numeric scoring method to check whether an AI's claims about a transcript are backed by real citations from that transcript, without needing human experts to grade every answer by hand. The benchmark also tests what the researchers call conscious incompetence: whether a model admits when a transcript does not contain enough information to answer a question, instead of inventing an answer anyway. Across the models tested, the paper reports strong performance on grounding claims in citations, but weaker performance on whether those claims are actually correct.
Investors and analysts already lean on AI to turn hours of executive hedging into quick notes. This benchmark shows those tools are good at pointing to the right passage but still get the underlying facts wrong, and they do not always recognize when they lack enough information to answer at all. That is the kind of mistake that would get a human analyst fired.
A citation next to a claim does not make the claim true. It just makes it checkable.