AI/ speculative-decoding · speech-recognition · android · on-device-ai

New Decoding Trick Speeds Up Speech AI on Phones

AS2D lets an audio AI's draft and verify steps run in parallel, boosting on-device transcription throughput 42-76% across four Android phones.

A new decoding trick promises faster on-device speech transcription by letting a phone's AI stop waiting for itself.

Speculative decoding already speeds up AI text generation by having a small "drafter" model guess ahead while a larger "target" model checks its work in batches, but the drafter has always had to wait for the target's confirmed output before guessing again. AS2D removes that lockstep for audio models: the drafter keeps generating guesses straight from the incoming audio, and the target verifies and corrects whatever's ready without dictating the drafter's pace. Tested on four Android phones through the MNN inference engine, across two target models, seven datasets, and 12.2 hours of speech, the technique lifted pooled transcription throughput 42-76% over running the target alone, with only 5.7% of test windows landing slower than the baseline versus 58-63% for older speculative methods. A separate benchmark using a larger 7-billion-parameter target and a single-shot execution mode reported up to 78% higher throughput, a bigger number, but from a different measurement setup, not a like-for-like comparison with the pooled 42-76% figure.

The appeal here isn't a smarter model, it's less wasted motion: on phones, every millisecond spent waiting on a chip translates directly into battery drain and lag, which is why voice assistants and live captioning still feel sluggish next to their cloud counterparts. This kind of decoupling could matter more than another leaderboard-topping model release, since it runs on existing hardware and slots into an inference stack developers already use.

It's still a research prototype tested on a narrow set of phones and tasks, but if the approach holds up in production, it's the kind of unglamorous systems engineering that decides whether on-device AI features actually ship, not another benchmark chart.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →