[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-small-ai-models-beat-gpt-5-at-security-query-translation":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},6955,"small-ai-models-beat-gpt-5-at-security-query-translation","Small AI Models Beat GPT-5 at Security Query Translation","A two-stage small-model pipeline for turning plain English into security queries outperforms GPT-5 on accuracy while costing a fraction as much.","Small language models just outperformed GPT-5 at one of the more tedious jobs in a security operations center: turning a plain-English question into a working database query.\n\nResearchers tested a system for converting natural-language questions into Kusto Query Language (KQL), the syntax analysts use to search Microsoft security telemetry. They tried three approaches: prompting a small model with a handful of mined tips about common parsing errors, fine-tuning a small model with chain-of-thought reasoning borrowed from a larger teacher model, and pairing a small model that drafts a query with a separate, cheap large model that checks and fixes it against the schema. That third, two-stage approach worked best. On Microsoft's own NL2KQL Defender evaluation set, it produced syntactically correct queries 98.7% of the time and semantically correct ones 90.6% of the time, topping every other setup the researchers tested, including pipelines built on GPT-4o (87.8%) and GPT-5 (86.1%).\n\nThe bigger news is cost. That two-stage small-model pipeline handled the researchers' 230-query test set for $0.213, versus $2.018 for the GPT-5 version and $2.998 for the GPT-4o version - a 9.5x to 14x reduction, while still beating both on schema accuracy. For security teams drowning in telemetry and short on analysts who can write KQL, a translator that is both cheaper and more accurate is a rare combination, not an incremental one.\n\nThe fine-tuning experiment, notably, went nowhere: teaching the small model to reason like a bigger one never beat prompting it well in the first place - a useful reminder that more training isn't always the fix vendors promise.","[\"ai\",\"security\",\"small-language-models\",\"threat-hunting\"]","2026-09-18T04:00:00.000Z","2026-09-19T00:23:04.244Z","2026-09-19T00:23:16.144Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek says the small-model pipeline 'matches' GPT-5's accuracy, but the body and source show it actually exceeds the GPT-5 baseline (0.906 vs 0.861 schema-valid), which also contradicts the headline's 'Best' framing — reconcile the dek's language with the data and headline.","resolved","security",[32,30,33,34],"ai","small-language-models","threat-hunting",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2512.06660",0,{"sections":41},[42,45,48,53,58,62,66,71,75,80,85,90,95,100],{"name":43,"slug":32,"count":44,"latest_published_at":18},"AI",4082,{"name":46,"slug":30,"count":47,"latest_published_at":18},"Security",661,{"name":49,"slug":50,"count":51,"latest_published_at":52},"Policy","policy",339,"2026-09-17T12:00:00.000Z",{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":18},"Hardware","hardware",155,{"name":63,"slug":64,"count":65,"latest_published_at":18},"Science","science",125,{"name":67,"slug":68,"count":69,"latest_published_at":70},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":18},"Dev Tools","dev-tools",78,{"name":76,"slug":77,"count":78,"latest_published_at":79},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]