Skip to main navigation Skip to search Skip to main content

SciAidanBench: Evaluating LLM Scientific Creativity

  • Brookhaven National Laboratory

Research output: Contribution to conferencePaperpeer-review

Abstract

Large language models (LLMs) have demonstrated impressive general creativity, but their capacity for scientific creativity—generating novel, domain-relevant ideas—remains underexplored. We introduce SciAidanBench, a benchmark of 155 open-ended scientific questions designed to measure scientific idea generation in LLMs. We evaluate 19 models across seven vendors and observe a strong positive correlation (r = 0.81) between general and scientific creativity, though models consistently produce fewer scientific ideas per prompt. Vendor-specific variations and performance gaps between non-reasoning and reasoning variants of models are observed. This preliminary investigation provides initial findings on evaluating the scientific creativity of LLMs.

Original languageEnglish
Pages25-28
Number of pages4
DOIs
StatePublished - 2025
EventNew York Scientific Data Summit 2025: Powering the Future of Science with Artificial Intelligence, NYSDS 2025 - New York City, United States
Duration: Sep 11 2025Sep 12 2025

Conference

ConferenceNew York Scientific Data Summit 2025: Powering the Future of Science with Artificial Intelligence, NYSDS 2025
Country/TerritoryUnited States
CityNew York City
Period09/11/2509/12/25

Fingerprint

Dive into the research topics of 'SciAidanBench: Evaluating LLM Scientific Creativity'. Together they form a unique fingerprint.

Cite this