Security/ model-stealing · graph-neural-networks · ai-security · arxiv

Attack Clones Graph AI Models Using a Fraction of the Queries

A new preprint shows an attacker can copy a graph neural network's behavior using only hard labels and a fraction of the queries rivals need.

A new attack can clone a graph neural network's behavior after querying it a fraction of the times earlier methods needed.

According to a preprint posted to arXiv (arXiv:2609.37972, not yet peer-reviewed), researchers built a system called Dagger that steals graph neural networks deployed as black-box APIs. Unlike prior attacks, Dagger works even when the target only returns hard labels, not confidence scores, and when the attacker has no idea what architecture is running underneath. The method trains a surrogate in two phases: first it reconstructs graph structure around sparse or isolated nodes and fills in gaps using a technique called manifold-level node mixup, then it freezes that model and retrains just the classifier with class-balanced sampling to fix uneven label coverage. Tested across four benchmark graphs and four GNN backbones, the paper reports up to 18.16% higher fidelity to the victim model while using 12.23 times fewer queries than the strongest prior attack.

That query efficiency is the real story. Graph neural networks increasingly sit behind machine-learning-as-a-service APIs for things like fraud detection and recommendation systems, and defenders have often assumed hard-label, low-budget, backbone-agnostic attacks were impractical to pull off. This preprint argues that assumption doesn't hold, and it does so under the exact constraints a real-world attacker would actually face, not a lab-friendly best case.

Like most arXiv preprints, this one hasn't been peer-reviewed or replicated outside the authors' own benchmarks, so the headline numbers are a claim worth watching, not a settled verdict.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →