[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-a-cheaper-way-to-probe-inside-ai-models":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},8128,"a-cheaper-way-to-probe-inside-ai-models","A Cheaper Way to Probe Inside AI Models","A new logistic probe called RAPTOR promises accurate, stable concept vectors for AI activation steering at a fraction of the usual training cost.","A new probe design claims it can map what is happening inside a language model without the usual cost or instability.\n\nResearchers describe RAPTOR (Ridge-Adaptive Logistic Probe), an L2-regularized logistic probe that tunes its ridge penalty on validation data and derives concept vectors from the resulting normalized weights. Tested on instruction-tuned LLMs and human-written concept datasets, it matched or beat established probing baselines on accuracy, held up about as well under directional-stability checks (whether a vector still points the same way after parts of the model are altered), and trained for substantially less compute. The paper backs this up with steering demos showing the extracted vectors actually shift model outputs. It also adds a theory piece: using the Convex Gaussian Min-max Theorem, the authors work out why ridge strength trades off accuracy against stability in a simplified statistical model, then show the same pattern shows up in real LLM embeddings.\n\nProbing is the workhorse behind probe-then-steer interpretability pipelines: find a concept inside a model's activations, then nudge outputs by injecting that vector during inference. Most of that work has leaned on ad hoc probe choices with no real guarantee the resulting vector is stable or worth the compute spent finding it. A cheaper, theoretically grounded probe lowers the barrier for smaller labs to build the same steering pipelines that better-funded interpretability teams already use.\n\nThe framing here is telling: this is pitched as an engineering fix for a known headache, not a claim about what models actually \"know\" - so it is unlikely to settle the field's bigger argument over whether steering vectors capture real concepts at all.","[\"ai-interpretability\",\"activation-steering\",\"llm-research\",\"arxiv\"]","2026-09-28T04:00:00.000Z","2026-09-28T10:33:16.765Z","2026-09-28T10:33:28.128Z","published",null,[],"ai",[26,27,28,29],"ai-interpretability","activation-steering","llm-research","arxiv",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2602.00158",0,{"sections":36},[37,40,44,49,54,59,63,68,73,78,83,88,92,97],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",4794,{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",762,{"name":45,"slug":46,"count":47,"latest_published_at":48},"Policy","policy",399,"2026-09-27T18:39:02.000Z",{"name":50,"slug":51,"count":52,"latest_published_at":53},"Deals","deals",261,"2026-09-27T15:30:35.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Hardware","hardware",188,"2026-09-27T20:46:36.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":18},"Science","science",151,{"name":64,"slug":65,"count":66,"latest_published_at":67},"Consumer Tech","consumer-tech",135,"2026-09-26T14:30:00.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Dev Tools","dev-tools",84,"2026-09-26T04:20:58.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Startups","startups",76,"2026-09-25T18:33:59.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":89,"slug":90,"count":86,"latest_published_at":91},"General","general","2026-09-26T17:02:42.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]