A new benchmark wants to predict whether a plant-based burger tastes like beef before any human takes a bite.
Researchers released TasteBench, a benchmark and competition for testing whether machine-learning models can predict how foods taste. It combines two tasks: ranking 215 plant-based foods across 24 categories using more than 21,000 human taste evaluations, and classifying 15,000 individual flavor molecules. The team also measured how consistent human tasters are with each other, finding agreement among panelists so low (Krippendorff's alpha of .077) that the theoretical ceiling for aggregated panel rankings tops out at .825. On the same food pairs panelists judged, the best baseline model hit .661 pairwise accuracy - essentially tied with the .650 score of a typical individual panelist.
Sustainable protein companies currently lean on slow, costly human taste panels to check whether a new formulation is close enough to the animal-based product it's replacing. A workable computational proxy could do for food science what molecular docking did for drug discovery - compress months of trial and error into a quick simulation. The catch TasteBench's own numbers reveal is that good here means matching noisy human judgment, not some objective truth about flavor.
A model that ties an average taster isn't cracking taste - it's just as confused as the rest of us, consistently.