AI/ image-editing · ai-research · computer-vision

New Image Editor Lets You Point Instead of Describe

VibeEdit lets users mark edits directly on an image instead of typing prompts, and it outperforms top text-based editors on tricky, lookalike-object edits.

A new research prototype swaps typed edit instructions for marks drawn right on the photo.

Researchers built VibeEdit, an image-editing system that reads spatial marks and short notes placed directly on a canvas instead of a text prompt. The team built on Qwen-Image-Edit, adding layer-decoupled conditioning that processes the source image and the canvas marks separately. To train it, they assembled 1.55 million source-target edit pairs with object masks and structured descriptions, then fine-tuned the model with region-weighted supervision and a rubric-guided reinforcement learning stage aimed at completing edits cleanly while leaving the rest of the image untouched. On a 419-case benchmark built specifically to test picking the right object among lookalikes, VibeEdit scored 79.9 on a VLM rubric and 32.8 dB on outside-region PSNR, versus 67.4 and 24.0 dB for FireRed, the top text-only baseline tested.

Text prompts get clumsy fast when a photo has five identical mugs and you only want to change one. Pointing at the object you mean sidesteps that ambiguity entirely, and the PSNR gap suggests marked edits also bleed less into surrounding pixels - a persistent weak spot for prompt-only editors.

It is still a research benchmark measured against one rival, not a shipped product, so treat the numbers as a promising lead rather than a verdict on how editing tools will actually work next year.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →