A new paper argues that letting AI quietly turn notes, chats, or wearable data into study covariates can change what your experiment is actually measuring.
The researchers propose a causal type discipline for sequential experiments, such as clinical trials, that rely on AI-generated covariates built from notes, conversations, images, or wearable streams. Their argument: a generated feature might represent a treatment version, a pre-action state, a mediator, an outcome proxy, or an intercurrent event, and those roles are not interchangeable. If nobody specifies which role a feature is playing, the AI-generated stand-in can swap the causal question mid-study without anyone noticing. Their proposed fix locks a standardized proximal effect before any generated covariate enters the analysis, then runs each one through a causal-role classifier and a claim-status filter.
That distinction matters because this failure mode does not show up as a biased number on a dashboard. A biased estimate still answers the right question, just badly; a mistyped covariate can answer a different question convincingly, and the output looks just as clean either way. As more trials and A/B tests lean on AI to turn notes, video, or sensor logs into covariates, that is exactly the kind of error a p-value will not catch.
The paper's own simulations are a useful check on claims that AI automatically improves data quality: refinement only helps when it preserves design-relevant information, while erasing that information, leaking post-action information, or selecting markers from the same data used to test them all produced bias or blew past nominal confidence-interval coverage.