AI chatbots are much easier to manipulate when a bad request is dressed up as code instead of plain English.
Researchers built an automated tool called CodeMimicry that wraps harmful requests inside structured, object-oriented code prompts rather than natural-language ones. Tested against eight commercial large language models, it produced a 96.25% attack success rate using just 1.51 queries on average. That beat both older template-based jailbreaks and newer optimization-based ones. The researchers argue this works because safety training is done mostly in natural language, and that training doesn't fully transfer to code.
The bigger problem is that this is a fully automated, black-box method - it needs no access to a model's internals, so anyone could plausibly reproduce it. The researchers also looked inside the models and found that code-formatted prompts measurably push the model's internal representations away from "refusal" directions, pointing to a structural blind spot rather than a one-off clever phrasing trick.
Safety teams keep patching new English phrasings while the code window sits wide open.