AI/ image-fusion · computer-vision · ai-research · benchmarks

AI Model Learns to Judge Infrared Fusion Like a Human

A new learned metric predicts human preferences on infrared-visible image fusion more accurately than 19 traditional formulas.

A team of researchers has trained a machine learning model to judge infrared-visible image fusion the way a person would, instead of relying on decades-old math formulas.

The model, called LPIFM (Learned Perceptual Image Fusion Measure), studies pairs of source images and pairs of fused results together, then predicts which fused image a human would prefer or whether they would call it a tie. It was trained on 6,300 human A/B/Tie comparisons across 25 fusion methods and 21 scenes from the VIFB benchmark, a dataset the researchers built and released publicly. Across four evaluation settings, LPIFM agreed with human judgments 79.2-84.0% of the time and tracked human rankings with Spearman correlations of 0.941-0.977. That beat the best of 19 conventional metrics by 16.3-21.1 percentage points, and on a separate dataset, EVAFusion, it topped every conventional metric after just three epochs of fine-tuning.

Infrared-visible fusion underpins things like night-vision cameras, surveillance systems, and autonomous-vehicle sensors, and grading how good a fusion method looks has meant either expensive human panels or formulas that measure pixel statistics rather than what people actually notice. A metric that predicts human preference without the panel could let researchers compare new fusion methods faster and more cheaply as the field fills with new techniques.

Still, the whole system learned its taste from one benchmark's 21 scenes and 25 methods, so how well that taste generalizes beyond VIFB and EVAFusion remains a question the paper's own transfer test only partially answers.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →