dattri-LLM traces which training examples actually shaped a large language model's output.
Training data attribution works by estimating how much each training example contributed to what a model says. The best-known methods do this through per-example gradients, but calculating and storing those gradients across billions of parameters has been slow and brittle. dattri-LLM tackles that by compressing gradients into compact representations and routing gradient computations through a cost model that picks the cheapest viable approach. It hooks into training loops that already call backward(), so it works with distributed setups like Distributed Data Parallel (DDP) and Fully Sharded Data Parallel (FSDP), and with existing pipelines in HuggingFace Transformers, TRL, and OLMo, without requiring any rewrite.
That matters because attribution has mostly lived in research papers, not production pipelines, precisely because it has been too slow and too fiddly to bolt onto a real training run. A tool that plugs into standard distributed training without code changes lowers the bar for labs to actually use attribution for things like finding which documents led to a bad output, or selecting better training data on the fly. The paper reports the library scaling multiple attribution methods to a 110-billion-parameter model across four H200 GPUs, at 3.2 times the throughput of the next-fastest library tested.
Those numbers come from the authors own benchmarks, so treat the speed claim as unverified until other labs test it on their own hardware and models.