[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-tte-flash-2b-outperforms-explicit-reasoning-without-the-cost":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},7391,"tte-flash-2b-outperforms-explicit-reasoning-without-the-cost","TTE-Flash-2B Outperforms Explicit Reasoning Without the Cost","A new AI model skips writing out its reasoning yet still beats the version that spells it out, on a standard multimodal benchmark.","A multimodal embedding model that reasons silently just beat its more talkative sibling on a major benchmark.\n\nUniversal Multimodal Embedding systems that use Chain-of-Thought reasoning tend to perform well: a generative model writes out its reasoning about a query, then the final representation is pulled from a token that has read both the query and that reasoning text. The catch is that generating all that reasoning is slow and expensive at inference time. A new paper called TTE-Flash swaps the written-out reasoning for latent \"think tokens\" - internal representations trained to imply the same reasoning without spelling it out. The think tokens are trained with the same loss used to generate explicit reasoning, while the embedding tokens that follow them are trained separately with a contrastive loss. The resulting model, TTE-Flash-2B, outperforms its explicit-CoT counterpart on the MMEB-v2 benchmark, at a fixed inference cost no matter how much reasoning the query would otherwise require.\n\nThat matters because CoT reasoning has become one of the more reliable ways to improve embedding quality, but it scales badly - more reasoning tokens means more compute per query. TTE-Flash keeps the accuracy gain while decoupling it from that cost, and the researchers say the latent think tokens can still be decoded back into readable text or mapped to image regions, so the reasoning isn't a total black box. Zero-shot tests across 15 video datasets also show performance improving as the number of think tokens increases, suggesting the approach scales predictably.\n\nIt's one arXiv paper, not a shipped product, so treat \"outperforms\" as a benchmark claim until someone outside the authors' lab reproduces it.","[\"ai\",\"embeddings\",\"chain-of-thought\",\"multimodal\"]","2026-09-23T04:00:00.000Z","2026-09-23T11:15:37.879Z","2026-09-23T11:15:43.663Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek claims TTE-Flash-2B is 'matching' embedding quality while the body says it 'beats' its explicit-CoT counterpart on MMEB-v2 — reconcile the dek and body so they describe the same result (outperforms, not merely matches).","resolved","ai",[30,32,33,34],"embeddings","chain-of-thought","multimodal",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.16638",0,{"sections":41},[42,46,51,56,61,66,71,76,81,86,91,96,101,106],{"name":43,"slug":30,"count":44,"latest_published_at":45},"AI",4430,"2026-09-24T15:24:49.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":50},"Security","security",724,"2026-09-24T04:00:00.000Z",{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",383,"2026-09-24T14:15:05.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",232,"2026-09-24T14:00:29.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Hardware","hardware",177,"2026-09-24T15:24:48.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Science","science",136,"2026-09-24T09:00:00.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Consumer Tech","consumer-tech",120,"2026-09-24T14:24:47.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Software","software",85,"2026-09-23T20:00:00.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Dev Tools","dev-tools",79,"2026-09-22T22:21:13.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Startups","startups",67,"2026-09-24T14:31:00.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Gaming","gaming",45,"2026-09-22T15:35:06.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"General","general",43,"2026-09-21T23:48:56.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"Reviews","reviews",29,"2026-09-24T13:00:00.000Z",{"name":107,"slug":108,"count":109,"latest_published_at":110},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]