AI/ ai alignment · fine-tuning · llm behavior · ai safety

Teaching an AI to Format Text Nudged Its Willingness to Answer

Post-training a 27B model purely for Korean response style also shifted how often it dodges sensitive questions, researchers found.

Fine-tuning a large language model to sound more Korean also changed how often it agrees to answer questions it should probably refuse.

Researchers post-trained Qwen3.8-27B, a 27-billion-parameter model, purely on Korean response style: verbosity, list and markdown formatting, discourse structure, and register. The training never touched the model's factual knowledge or reasoning targets. Yet two unrelated behaviors moved anyway: how often the model abstains on ambiguous social-bias questions in the KoBBQ benchmark, where the correct answer is 'unknown,' and how often it volunteers unprompted disclosures in securities-guidance scenarios. Using matched controls that held prompts, training recipe, and data volume fixed while swapping only the target text, three 'style' training seeds raised the answer rate by an average of 0.82 percentage points, while three 'neutral' seeds lowered it by an average of 1.53 points, a 2.34-point gap with no overlap between the two seed groups.

That shift comes almost entirely from how often the model chooses to answer at all, not from what it says once it does. A team optimizing purely for tone or formatting could unknowingly make a model more likely to guess on questions it should decline, or to over-disclose in regulated settings like securities guidance. The researchers also found the two automated detectors used to measure this effect agreed with each other anywhere from 44 percent to 99 percent of the time depending on which checkpoint produced the text, a sign the measurement tools themselves are unsteady.

Style is supposed to be the cosmetic layer of a fine-tune. This paper suggests it can quietly rewire judgment calls that have nothing to do with formatting.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →