AI/ ai · deep-research · benchmarks · agents

LongCat's AI Research Agent Ranks Second, Rivals Unnamed

LongCat's multi-agent research system scored strong marks on three benchmarks and placed second of four in an undisclosed in-house test.

LongCat has published a technical report for an AI system that plans, researches, and writes full reports on its own.

The system, called LongCat-DeepResearch, splits the job into stages: planning agents scan outside sources and draft a plan the team calls a ResearchSpec, then separate research agents investigate and write their assigned sections in parallel, gathering their own evidence as they go. A global review step then flags specific sections for revision instead of rewriting the whole report from scratch. On three public report-quality benchmarks, DeepResearchBench, DeepResearchBench II, and ResearchRubrics, each scored on a scale that tops out near 100, it posted 55.25, 51.35, and 79.83. On an in-house benchmark it scored 76.04, good for second place among four systems tested.

That "second of four" ranking is the detail worth pausing on: LongCat's report never names the other three systems or their scores, so there is no way to tell whether it lost a close race or got beaten badly. The workflow itself, splitting planning from drafting, then editing only what needs it, is a sensible fix for the usual problem of AI-written reports collapsing into mush on revision. A self-reported rank with no visible competitors is not a comparison, it is a claim.

File this under wait-and-see until LongCat-DeepResearch runs against named competitors, in the open, on a benchmark where every score is on the table.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →