[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-models-can-now-coach-themselves-while-solving-problems":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},6966,"ai-models-can-now-coach-themselves-while-solving-problems","AI Models Can Now Coach Themselves While Solving Problems","A new test-time training method has one AI model play both student and teacher, generating practice problems to fix its own reasoning gaps.","A new technique called TTSR has a single AI model quiz, grade, and retrain itself on the fly, without any human-labeled data.\n\nResearchers behind TTSR (Test-Time Self-Reflection) split one pretrained model into two roles that take turns: a Student that attempts test questions, and a Teacher that reviews the Student's failed attempts and writes new practice questions calibrated to what the Student is actually getting wrong. The system keeps a running \"weakness memory\" across iterations and distills it into a short strategy note that gets prepended to the Student's next attempts, then fades out as those weaknesses get resolved. The approach targets two known problems with test-time training: models running out of trustworthy practice problems to learn from, and wasting compute on repeated random retries instead of diagnosing what actually went wrong. On difficult math reasoning benchmarks, it produced consistent gains, held up across different base models, and carried over to reasoning tasks outside math.\n\nMost test-time training methods lean on a model grading its own guesses, which gets unreliable fast on hard questions, since noisy self-labels produce noisy rewards. TTSR sidesteps that by having the model generate targeted follow-up questions instead of just more attempts at the same one, turning failure analysis into new training material. That is a meaningful shift for anyone trying to squeeze more reasoning capability out of a model without extra fine-tuning data or compute-heavy reinforcement learning runs.\n\nIt is still one model quizzing itself with no outside grader, so there is plenty of room for it to confidently reinforce the wrong lesson; the benchmark results are the real test of whether that risk stays manageable.","[\"ai\",\"llms\",\"machine-learning\",\"reasoning\"]","2026-09-18T04:00:00.000Z","2026-09-19T00:50:39.119Z","2026-09-19T00:50:51.029Z","published",null,[],"ai",[24,26,27,28],"llms","machine-learning","reasoning",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2603.03297",0,{"sections":35},[36,39,43,48,53,57,61,66,70,75,80,85,90,95],{"name":37,"slug":24,"count":38,"latest_published_at":18},"AI",4082,{"name":40,"slug":41,"count":42,"latest_published_at":18},"Security","security",661,{"name":44,"slug":45,"count":46,"latest_published_at":47},"Policy","policy",339,"2026-09-17T12:00:00.000Z",{"name":49,"slug":50,"count":51,"latest_published_at":52},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":54,"slug":55,"count":56,"latest_published_at":18},"Hardware","hardware",155,{"name":58,"slug":59,"count":60,"latest_published_at":18},"Science","science",125,{"name":62,"slug":63,"count":64,"latest_published_at":65},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":18},"Dev Tools","dev-tools",78,{"name":71,"slug":72,"count":73,"latest_published_at":74},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]