TY - GEN
T1 - PRMBENCH
T2 - 63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025
AU - Song, Mingyang
AU - Su, Zhaochen
AU - Qu, Xiaoye
AU - Zhou, Jiawei
AU - Cheng, Yu
N1 - Publisher Copyright:
© 2025 Association for Computational Linguistics.
PY - 2025
Y1 - 2025
N2 - Process-level Reward Models (PRMs) are crucial for complex reasoning and decision-making tasks, where each intermediate step plays an important role in the reasoning process. Since large language models (LLMs) suffer from various types of errors during the reasoning process, PRMs are required to possess nuanced capabilities for detecting various implicit error types in real-world scenarios. However, current benchmarks primarily focus on step correctness, failing to evaluate PRMs' performance systematically. To address this gap, we introduce PRMBENCH, a process-level benchmark specifically designed to assess the fine-grained error detection capabilities of PRMs. PRMBENCH comprises 6,216 carefully designed problems and 83,456 step-level labels, evaluating models across multiple dimensions, including simplicity, soundness, and sensitivity. In our experiments on 25 models, spanning across both open-source PRMs and LLMs prompted as critic models, we uncover significant weaknesses in current PRMs. These findings reveal the challenges inherent in process-level evaluation and highlight key directions for future research, establishing PRMBENCH as a robust testbed for advancing research on PRM evaluation and development.
AB - Process-level Reward Models (PRMs) are crucial for complex reasoning and decision-making tasks, where each intermediate step plays an important role in the reasoning process. Since large language models (LLMs) suffer from various types of errors during the reasoning process, PRMs are required to possess nuanced capabilities for detecting various implicit error types in real-world scenarios. However, current benchmarks primarily focus on step correctness, failing to evaluate PRMs' performance systematically. To address this gap, we introduce PRMBENCH, a process-level benchmark specifically designed to assess the fine-grained error detection capabilities of PRMs. PRMBENCH comprises 6,216 carefully designed problems and 83,456 step-level labels, evaluating models across multiple dimensions, including simplicity, soundness, and sensitivity. In our experiments on 25 models, spanning across both open-source PRMs and LLMs prompted as critic models, we uncover significant weaknesses in current PRMs. These findings reveal the challenges inherent in process-level evaluation and highlight key directions for future research, establishing PRMBENCH as a robust testbed for advancing research on PRM evaluation and development.
UR - https://www.scopus.com/pages/publications/105021054699
U2 - 10.18653/v1/2025.acl-long.1230
DO - 10.18653/v1/2025.acl-long.1230
M3 - Conference contribution
AN - SCOPUS:105021054699
T3 - Proceedings of the Annual Meeting of the Association for Computational Linguistics
SP - 25299
EP - 25346
BT - Long Papers
A2 - Che, Wanxiang
A2 - Nabende, Joyce
A2 - Shutova, Ekaterina
A2 - Pilehvar, Mohammad Taher
PB - Association for Computational Linguistics (ACL)
Y2 - 27 July 2025 through 1 August 2025
ER -