AI/ ai · recommender systems · kuaishou · machine learning

Kuaishou's New Recommender AI Skips the Verbose Reasoning Step

Kuaishou's OneLatent compresses AI reasoning into hidden tokens, boosting recommendation accuracy and cutting inference cost dramatically.

Kuaishou has found a way to make AI-powered recommendations smarter without making them slower.

Researchers at the short-video platform built OneLatent, a recommendation system that swaps out long, written-out AI reasoning for a handful of compressed "latent tokens" that capture the same reasoning invisibly. The system first trains on varied reasoning examples generated by a teacher model, then gradually compresses those reasoning chains into the latent tokens through a three-stage alignment process, and finally fine-tunes them for the recommendation task. Tested on an industrial-scale Kuaishou dataset and a public benchmark, OneLatent beat both the reasoning and non-reasoning versions of the same base model on SID@64, a metric that measures how often the correct item lands among a model's top 64 ranked predictions, improving scores by 17.44% and 9.33% respectively. It also delivered over 17 times the inference throughput of the explicit-reasoning approach.

AI reasoning models that think out loud are more accurate but too slow and costly to run at the scale of serving recommendations to hundreds of millions of users. Compressing that reasoning into a fixed set of tokens instead of generating full text sidesteps that tradeoff. Kuaishou already moved OneLatent into production and ran an A/B test in its local-services ads business, reporting a 9.6% revenue lift over prior systems including OneRec and OneReason.

It's a reminder that in recommendation engines, the real contest isn't who reasons best, but who reasons fastest per dollar.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →