Abstract
Large language models (LLMs) have demonstrated impressive general creativity, but their capacity for scientific creativity—generating novel, domain-relevant ideas—remains underexplored. We introduce SciAidanBench, a benchmark of 155 open-ended scientific questions designed to measure scientific idea generation in LLMs. We evaluate 19 models across seven vendors and observe a strong positive correlation (r = 0.81) between general and scientific creativity, though models consistently produce fewer scientific ideas per prompt. Vendor-specific variations and performance gaps between non-reasoning and reasoning variants of models are observed. This preliminary investigation provides initial findings on evaluating the scientific creativity of LLMs.
| Original language | English |
|---|---|
| Pages | 25-28 |
| Number of pages | 4 |
| DOIs | |
| State | Published - 2025 |
| Event | New York Scientific Data Summit 2025: Powering the Future of Science with Artificial Intelligence, NYSDS 2025 - New York City, United States Duration: Sep 11 2025 → Sep 12 2025 |
Conference
| Conference | New York Scientific Data Summit 2025: Powering the Future of Science with Artificial Intelligence, NYSDS 2025 |
|---|---|
| Country/Territory | United States |
| City | New York City |
| Period | 09/11/25 → 09/12/25 |
Fingerprint
Dive into the research topics of 'SciAidanBench: Evaluating LLM Scientific Creativity'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver