A new academic project wants to make sure companies doing legally required AI rights checks actually have evidence to check against.
Under Article 27 of the EU AI Act, companies deploying high-risk AI systems must run Fundamental Rights Impact Assessments before rollout. Researchers say the evidence for those assessments is scattered across incompatible incident databases, mismatched risk vocabularies, and dense legal text. Their fix is a Semantic Web framework that pulls a 150-record corpus into a single SPARQL-queryable knowledge graph of 1,351 RDF triples, tagged along four axes using keyword matching, LLM classification, and a hybrid of both. The framework covers two Annex III high-risk categories (employment and worker management, and access to essential public services), and its five demonstration scenarios surfaced 103 of those records, about 68.7% coverage.
The more interesting number here is 0.045. That is the agreement score when the researchers checked their LLM-based classifier for the employment domain against a 69-record gold standard, barely better than chance. For a tool meant to ground fairness-related legal compliance, that unreliable a result in exactly the domain where workplace-discrimination evidence would turn up is a real problem, not a footnote.
It is a useful case study in why automating AI compliance work is not a simple shortcut: the researchers built the plumbing, then found their own automated tagging isn't trustworthy enough yet to replace someone actually reading the evidence.