Speaker recognition systems can tell voices apart, but nobody could say exactly how - until researchers built a way to read the internal map those systems draw.
The study takes speaker embeddings, the compressed numerical fingerprints these neural networks generate for each voice, and runs them through Single-Linkage Clustering (SLINK) to see if they naturally group into a hierarchy. To check whether those clusters mean anything to a human, the team introduces Hierarchical Cluster-Class Matching (HCCM), a new method that tests how well each cluster lines up with known categories - individual traits like gender, and combined ones like "UK and male." A companion metric, the L-score, quantifies how sloppy or clean each match is, so a middling result is not just waved away as noise. Applied to real embeddings, HCCM found the hierarchy is organized around speaker identity first, then splits along recognizable lines like gender and nationality.
That matters because speaker recognition already sits behind voice unlock, call-center verification, and forensic audio analysis, and nobody auditing those systems could previously say what information the model was actually keying on. A tool that shows a model has silently encoded nationality alongside identity is also a tool for spotting where that encoding could bias outcomes or leak information the system was never asked to infer.
None of this makes the underlying neural network less of a black box - it just gives auditors a better flashlight, and flashlights only illuminate what you think to point them at.