Skip to main navigation Skip to search Skip to main content

LLMs Performance Evaluation: A Case Study in Climate Change Statements

  • Stony Brook University
  • Brookhaven National Laboratory

Research output: Contribution to journalArticlepeer-review

1 Scopus citations

Abstract

Large Language Models (LLMs) have shown remarkable performance across various natural language processing tasks, but their ability to interpret and classify scientific uncertainty remains underexplored. This study evaluates the performance of four recent LLMs—GPT-4o-mini, GPT-4o, GPT-4-turbo, and GPT-4.5-preview—on the ClimateX dataset, which consists of climate science statements labeled with expert-assessed confidence levels. Using both zero-shot and few-shot prompting strategies, we assess each model’s ability to classify statements into one of four IPCC-defined confidence categories: low, medium, high, and very high. Our findings indicate that while newer models like GPT-4.5-preview outperform earlier versions, overall classification accuracy remains modest (with a maximum of 51.33%). These results highlight the challenges LLMs face in interpreting nuanced scientific language and emphasize the need for further development and domain-specific adaptation to enhance their utility in climate communication and scientific analysis.

Original languageEnglish
Pages (from-to)79-81
Number of pages3
JournalPerformance Evaluation Review
Volume53
Issue number2
DOIs
StatePublished - Aug 27 2025

Fingerprint

Dive into the research topics of 'LLMs Performance Evaluation: A Case Study in Climate Change Statements'. Together they form a unique fingerprint.

Cite this