A new paper says LLMs used as annotators or judges carry baked-in biases that better prompts mostly cannot fix.
Researchers tested nine LLMs across six toxicity datasets and introduced a metric called Definition-Specific Familiarity, or DSF, which measures how closely a model's internal concept of a task matches the definition it is handed. DSF predicted annotation accuracy even after controlling for which dataset was used, a partial correlation of +0.41, while standard text-memorization metrics showed no such link. The team also tested whether extra instructions or examples could correct a model's zero-shot errors, a trait they call decision stickiness, and found only 34.8% of errors got fixed, with the most confident wrong answers the hardest to budge. When the task definition itself was misaligned with the model's internal sense of the category, outputs shifted but the model's stated confidence did not drop.
As more teams lean on LLM-as-a-judge pipelines for content moderation and labeling, this suggests a longer prompt will not necessarily rescue a model whose internal sense of a category, like toxic, does not match your policy. Worse, since confidence scores stay flat even when the definition is off, you cannot just filter out low-confidence answers to catch the mismatch.
It is a quieter but more useful finding than most AI benchmarks: pick the model whose baseline instincts already match your task, because no amount of prompt engineering reliably patches a mismatch after the fact.