AI/ ai · skincare · llm-benchmarks · consumer-tech

Chatbots Flunk Skincare Chemistry, Study Finds

A 14-model benchmark found chatbots answer skincare chemistry questions confidently but often incorrectly, especially on math and chemical structure.

Ask a chatbot to explain your serum's ingredients, and there's a good chance it will sound sure of itself while getting the chemistry wrong.

Researchers benchmarked 14 large language models on cosmetic chemistry, covering the properties of specific skincare ingredients and common consumer scenarios, according to a study posted to arXiv. Web search was disabled, so the models had to rely on what they absorbed during training rather than looking anything up. Overall accuracy was poor, with the weakest results on quantitative reasoning and identifying chemical structures. The models handled general, conversational skincare questions reasonably well, but even those answers lacked the technical depth needed for an informed purchasing decision.

The real risk isn't that a chatbot occasionally garbles a formula. It's that wrong answers sound exactly as confident as right ones, which is precisely what discourages a user from double-checking before mixing two actives. The researchers say closing the gap will take fine-tuning on verified chemical and dermatological data plus real gains in reasoning, not just bigger models.

It's the same pattern showing up in medical and legal chatbot benchmarks: general-purpose models are fluent generalists, not licensed specialists, and the skincare aisle is apparently no exception.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →