Ask a chatbot how to build a weapon and it refuses. Dress the same request up as ASCII art and ask for artistic feedback, and it often just answers.
A new study tests this trick, dubbed the ASCII Attack, across eleven language models and eight harm topics: single-turn, black-box, with the harmful request left fully readable inside the art and the model simply asked to critique it. Researchers compared these framed prompts against a matched direct-question control, scored by a harm-aware classifier. Framed prompts were judged harmful 62% of the time, versus 42% for direct requests. On the most susceptible model, the framed version worked 93% of the time.
That gap matters because it shows safety training still leans on surface phrasing rather than the actual request underneath it. Unlike ArtPrompt, an earlier attack that obscured text to dodge filters, this one hides nothing; it just relabels a request as art criticism, and that alone is often enough. The study also found judges disagreeing with each other on nearly two-thirds of framed prompts, a sign that the tools used to measure harm are shakier than they look.
If a model can be talked out of its own refusal just by being asked to play art critic, the refusal was never really about the content. It was about the wording.