Skip to main navigation Skip to search Skip to main content

PRMBENCH: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

  • Mingyang Song
  • , Zhaochen Su
  • , Xiaoye Qu
  • , Jiawei Zhou
  • , Yu Cheng
  • Fudan University
  • Shanghai Ai Laboratory
  • Soochow University
  • Chinese University of Hong Kong

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

20 Scopus citations

Abstract

Process-level Reward Models (PRMs) are crucial for complex reasoning and decision-making tasks, where each intermediate step plays an important role in the reasoning process. Since large language models (LLMs) suffer from various types of errors during the reasoning process, PRMs are required to possess nuanced capabilities for detecting various implicit error types in real-world scenarios. However, current benchmarks primarily focus on step correctness, failing to evaluate PRMs' performance systematically. To address this gap, we introduce PRMBENCH, a process-level benchmark specifically designed to assess the fine-grained error detection capabilities of PRMs. PRMBENCH comprises 6,216 carefully designed problems and 83,456 step-level labels, evaluating models across multiple dimensions, including simplicity, soundness, and sensitivity. In our experiments on 25 models, spanning across both open-source PRMs and LLMs prompted as critic models, we uncover significant weaknesses in current PRMs. These findings reveal the challenges inherent in process-level evaluation and highlight key directions for future research, establishing PRMBENCH as a robust testbed for advancing research on PRM evaluation and development.

Original languageEnglish
Title of host publicationLong Papers
EditorsWanxiang Che, Joyce Nabende, Ekaterina Shutova, Mohammad Taher Pilehvar
PublisherAssociation for Computational Linguistics (ACL)
Pages25299-25346
Number of pages48
ISBN (Electronic)9798891762510
DOIs
StatePublished - 2025
Event63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025 - Vienna, Austria
Duration: Jul 27 2025Aug 1 2025

Publication series

NameProceedings of the Annual Meeting of the Association for Computational Linguistics
Volume1
ISSN (Print)0736-587X

Conference

Conference63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025
Country/TerritoryAustria
CityVienna
Period07/27/2508/1/25

Fingerprint

Dive into the research topics of 'PRMBENCH: A Fine-grained and Challenging Benchmark for Process-Level Reward Models'. Together they form a unique fingerprint.

Cite this