AI/ face-detection · computer-vision · ai-research · arxiv

One Model Now Handles Every Face Landmark Dataset at Once

A new face landmark detection method trains one model across multiple datasets and lets it output any number of landmarks on demand.

A new face landmark detection method ditches the one-model-per-dataset approach that has quietly limited the field for years.

Researchers propose Face Part-Anchored Landmark Positions, or FPALPs, which represent every landmark as a progression value between zero and one along a face part's contour. Because any landmark from any "N-point" dataset can be expressed this way, previously incompatible datasets can be merged into a single training set. The team pairs this with FPALP-based queries that get refined through a cross-modality decoder before predicting final coordinates. The resulting system, called Unified Dynamic FLD, trains on any combination of datasets and produces any number of landmark predictions by loading the relevant queries at runtime, and it matched or beat existing state-of-the-art methods across multiple benchmark tests.

Most face landmark models are stuck with whatever point count their training set used, whether that's 68 or 98, forcing developers to maintain separate models for separate needs. Collapsing that into one adaptable model is a real engineering fix, not just a marginal accuracy gain, and it could simplify pipelines for anything from AR filters to expression tracking.

The results are benchmark numbers, not production numbers, so the harder question is whether merging datasets this way holds up once the inputs get messier than curated academic sets.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →