Researchers taught a wearable's motion sensors to spot eating and drinking by first training them on data nobody actually ate or drank to produce.
The team starts with CABiGRU, a neural network that combines convolutional layers, bidirectional GRUs, multi-head attention, and residual connections to read patterns from a smartwatch's accelerometer, gyroscope, and magnetometer. Eating and drinking are rare events compared to everything else a wrist does in a day, so models trained only on real recordings tend to miss them. The fix here is a diffusion model that generates synthetic sensor-data windows, which CABiGRU trains on first before fine-tuning on real recordings from the DEO (drinking/eating/other) dataset. That two-stage approach pushed balanced accuracy to 90.6 percent, beating a standard supervised baseline.
This is a workaround for a problem that shows up across health-tracking tech: the activities worth monitoring, like meals, falls, or medication-taking, are exactly the ones with the least training data. Diffusion models have already remade image and audio generation; using them to manufacture plausible sensor noise is a cheaper substitute for the months of real-world data collection dietary-monitoring research usually requires.
Still, this is one dataset and one activity pair, not a general cure for class imbalance in wearables, so the real test is whether the gains hold up on sensors and bodies the model has never seen.