AI/ ai · generative-ai · cognitive-science · ai-research

Paper Argues AI Should Be Judged by Process, Not Output

A new arXiv paper says AI mimics the outputs of human thinking without replicating the process behind them, and proposes tests to tell the difference.

A new arXiv paper argues that intelligence is defined by the process of thinking, not by the output it produces, and that most generative AI fails that standard.

The paper's authors say generative AI models are trained on traces, the textual and visual residue of human cognitive work, and generate new samples from that distribution. The result can look like reasoning, problem-solving, or creativity, but the paper argues the underlying activity that produces those outputs in humans is missing or opaque in the machine version. Drawing on a long-standing distinction in cognitive science between weak and strong equivalence, the authors define seven process features meant to test whether a system's cognition genuinely matches a human's, not just its output. They also propose 'process audits' as a way to make that strong equivalence testable, plus design principles meant to keep AI tools from replacing the thinking they claim to assist.

The more pointed claim is a symmetric risk: AI tools that do a person's generative work for them, rather than with them, may leave that person's own capacities unbuilt. That reframes a familiar complaint, that AI writing and coding tools make people lazier, as a testable design problem rather than a vibe.

It is a framework, not a product test, and there is no indication yet of what a process audit would look like applied to a chatbot or a coding assistant. But the vocabulary is useful. Most AI benchmarks still score whether an answer looks right, not whether anything resembling thought produced it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →