Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program
Paper in proceeding, 2025

Recent advancements in reward signal usage for Large Language Models (LLMs) are remarkable. However, significant challenges exist when transitioning reward signal to the multimodal domain, including labor-intensive annotations, over-reliance on one-step rewards, and inadequate evaluation. To address these issues, we propose SVIP, a novel approach to train a step-level multi-dimensional Chain-of-Thought (CoT) reward model automatically. It generates code for solving visual tasks and transforms the analysis of code blocks into the evaluation of CoT step as training samples. Then, we train SVIP-Reward model using a multi-head attention mechanism called TriAtt-CoT. The advantages of SVIP-Reward are evident throughout the entire process of MLLM. We also introduce a benchmark for CoT reward model training and testing. Experimental results demonstrate that SVIP-Reward improves MLLM performance across training and inference-time scaling, yielding better results on benchmarks while reducing hallucinations and enhancing reasoning ability.

multimodal model

reward model

chain of thought

visual programming

Author

Minghe Gao

Zhejiang University

National University of Singapore (NUS)

Xuqi Liu

Zhejiang University

Yue Nick Zhongqi

Nanyang Technological University

Data Science and AI 3

University of Gothenburg

Yang Wu

Ant group

Shuang Chen

Zhejiang University

Juncheng Li

Zhejiang University

Siliang Tang

Zhejiang University

Fei Wu

Zhejiang University

Tat Seng Chua

National University of Singapore (NUS)

Yueting Zhuang

Zhejiang University

Proceedings of the IEEE International Conference on Computer Vision

15505499 (ISSN) 23807504 (eISSN)

1718-1728
9798331587758 (ISBN)

2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
Honolulu, USA,

Subject Categories (SSIF 2025)

Computer Sciences

DOI

10.1109/ICCV51701.2025.00168

More information

Latest update

8/5/2026 7