A new generative model called Schema can mimic a giant graph's shape without copying most of its edges.
The paper, titled Scalable Hierarchical Graph Generation via Soft Community Structure, describes a method for generating large attributed graphs - think social networks or citation graphs - that exist as a single object rather than a batch of independent samples. Schema recursively splits a reference graph into a hierarchy of soft communities, giving each node a membership distribution rather than a single label. Generation then runs in three separately trained stages: node attributes conditioned on those memberships, edges within a community based on local structure, and connections between communities through 'bridge' nodes that belong to more than one. Because no stage ever builds the full adjacency matrix, each piece only has to handle a subgraph the size of a community.
That modularity solves a problem other graph generators run into: match the original graph's structure closely and you tend to just memorize it, or get the downstream accuracy right and you either overshoot the reference or fail to run at all on bigger graphs. On four real-world attributed graphs, Schema held onto both the structural balance and the downstream accuracy of the original without reproducing more than a small fraction of its actual edges, and the researchers report it scaling to graphs with up to 10 million nodes.
If the results hold up, that is useful for anyone who wants to share or benchmark on graph-shaped data without handing over the real thing - a quieter, more technical cousin of synthetic data generation for tabular or image datasets. It is also, for now, one arXiv preprint's claims about its own benchmark suite, not an independently verified standard.