A new study argues that data quality rules built for databases don't work once data starts shaping how a model behaves.
Researchers interviewed 16 practitioners across nine organizations about how they define and manage data quality in AI-driven systems, then used reflexive thematic analysis to pull six recurring themes out of the transcripts. Traceability, they found, has shifted from tracing a bug to a line of code toward tracing a model's output back to the data that shaped it. Teams are also using models to judge other models' data, which the paper flags as circular: the judge and the judged share the same blind spots. Agent memory and context are now treated as data objects in their own right, and synthetic or pseudo-labeled data has turned "is this real" into its own quality question.
That matters because most data quality tooling still assumes rows in a database, not context windows or model judgments. As more products bolt an LLM onto their own QA process, the circularity problem the study describes stops being academic and starts being a production risk. And with training-data lawfulness now treated as a gate rather than an afterthought, expect data provenance to become a compliance line item, not just an engineering one.
The paper's proposed fix, "lifecycle assurance," is really just a call for evidence trails a model's influence can be traced back through. It's a sound idea. Whether anyone builds it before the next dataset lawsuit lands is the more interesting question.