A new plug-in lets a cloud AI answer questions about your chart or audio clip without ever seeing the actual chart or audio clip.
Researchers built ReCast, an agentic framework that sits between a user and a remote multimodal model handling chart and speech reasoning tasks. Rather than uploading the raw file, it converts the input locally into a text record describing the task, then has a small 4B-parameter model rewrite the entities and topics inside that record. Numbers get swapped through a locally invertible, role-aware map, and a separate reconstruction agent regenerates a sanitized chart or audio clip to send to the remote solver. When the solver returns its answer as a program, ReCast substitutes the real numbers back in locally before running it.
This fixes a specific hole: plain-text redaction cannot satisfy a model that demands a fixed media interface like an image or audio file, and anonymizing names alone still leaves the sensitive numbers and topics in plain view. Tested on 4,000 held-out ChartQA and speech-QA examples, ReCast scored 75.10% accuracy, retaining 92.43% of what you would get sending the data unprotected, while an automated audit flagged source-content leakage in only 7.95% of requests sent to the solver.
Call it progress, not a solution. An audit that still finds leakage in nearly one of every twelve requests is a reminder that outsourcing reasoning to someone else's model always means trusting someone else's model, no matter how much scrubbing happens first.