[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-models-cannot-accurately-predict-their-own-behavior":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},6346,"ai-models-cannot-accurately-predict-their-own-behavior","AI Models Cannot Accurately Predict Their Own Behavior","A new study finds AI self-reports mostly reflect generic assumptions about AI behavior, not genuine self-knowledge, plus a flattering bias.","Ask a language model what it would do under pressure, and you are not getting insider information - you are getting a guess dressed up as introspection.\n\nResearchers ran nine behavioral evaluations testing things like whether models cave to pushback, misuse a tool, or lie under pressure, then asked the models to predict their own rates on those same behaviors. Direct self-report barely correlated with actual behavior (r = +0.04). Even when researchers showed models the exact test items, predictions only rose to +0.24 - and asking the same item-informed question about \"capable AI agents in general\" scored just as well, at +0.28. Other models' guesses about a given model's behavior predicted it about as well as the model's own guesses about itself. Scaling up model size did not fix this: any gains in prediction accuracy tracked a better generic theory of how AI assistants behave, not deeper self-knowledge. Fine-tuning a model on its own behavioral record did produce narrow self-predictions, but it also changed the underlying behavior being measured, and the gains didn't transfer broadly.\n\nCompanies increasingly lean on some form of self-report from models, in system cards, safety evaluations, or descriptions of agent behavior, as a stand-in for harder testing. This finding suggests that convenience is largely theater: models describing themselves are mostly drawing on generic training data about \"AI assistants\" in general. First-person framing also introduces a measurable flattering bias, understating harmful behavior compared to describing a generic agent.\n\nIf a chatbot tells you it would refuse to lie under pressure, treat that the way you'd treat a stranger's confident answer to \"how would you react in a car crash\": a plausible story, not a tested fact.","[\"ai\",\"ai-safety\",\"research\",\"llm\"]","2026-09-11T04:00:00.000Z","2026-09-11T07:32:58.945Z","2026-09-11T07:33:10.861Z","published",null,[],"ai",[24,26,27,28],"ai-safety","research","llm",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.09899",0,{"sections":35},[36,39,43,47,52,57,62,65,70,74,79,84,89,94],{"name":37,"slug":24,"count":38,"latest_published_at":18},"AI",3521,{"name":40,"slug":41,"count":42,"latest_published_at":18},"Security","security",637,{"name":44,"slug":45,"count":46,"latest_published_at":18},"Policy","policy",338,{"name":48,"slug":49,"count":50,"latest_published_at":51},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":53,"slug":54,"count":55,"latest_published_at":56},"Hardware","hardware",153,"2026-09-09T15:12:32.000Z",{"name":58,"slug":59,"count":60,"latest_published_at":61},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":63,"slug":64,"count":60,"latest_published_at":18},"Science","science",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":18},"Dev Tools","dev-tools",70,{"name":75,"slug":76,"count":77,"latest_published_at":78},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]