Measuring what matters: construct validity in large language model benchmarks

Bean, A; Kearns, RO; Romanou. A; Hafner, F; Mayne, H; Batzner, J; Foroutan, N; Schmitz, C; Korgul, K; Batra, H; Deb, O; Beharry, E; Emde, C; Foster, T; Gausen, A; Grandury, M; Han, S; Hofmann, V; Ibrahim, L; Kim, H; Kirk, H; Lin, F; Liu, GK-M; Luettgau, L; Magomere, J; Rystrom, J; Sotnikova, A; Yang, Y; Zhao, Y; Bibi, A; Bosselut, A; Clark, R; Cohan, A; Foerster, J; Gal, Y; Hale, S; Raji, ID; Summerfield, C; Torr, P; Ududec, C; Rocher, L; Mahdi, A

Conference item

Measuring what matters: construct validity in large language model benchmarks

Abstract:: Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as ‘safety’ and ‘robustness’ requires strong construct validity, that is, having measures that represent what matters to the phenomenon. With a team of 29 expert reviewers, we conduct a systematic review of 445 LLM benchmarks from leading conferences in natural language processing and machine learning. Across the reviewed articles, we find patterns related to the measured phenomena, tasks, and scoring metrics which undermine the validity of the resulting claims. To address these shortcomings, we provide eight key recommendations and detailed actionable guidance to researchers and practitioners in developing LLM benchmarks.

Publication status:: Published

Peer review status:: Peer reviewed

Actions

Email

Email this record

Send the bibliographic details of this record to your email address.

Your Email
Please enter the email address that the record information will be sent to.

-
Your message (optional)
Please add any additional information to be included within the email.
Cite

Cite this record

APA Style

Bean, A., Kearns, R. O., A, R., Hafner, F., Mayne, H., Batzner, J., Foroutan, N., Schmitz, C., Korgul, K., Batra, H., Deb, O., Beharry, E., Emde, C., Foster, T., Gausen, A., Grandury, M., Han, S., Hofmann, V., Ibrahim, L., … Mahdi, A. (2025). Measuring what matters: construct validity in large language model benchmarks.

MLA Style

Bean, A., et al. Measuring What Matters: Construct Validity in Large Language Model Benchmarks. NeurIPS, 2025.

Chicago Style

Bean, A, RO Kearns, Romanou. A, F Hafner, H Mayne, J Batzner, N Foroutan, et al. 2025. “Measuring What Matters: Construct Validity in Large Language Model Benchmarks.” In . NeurIPS.
Share
Print

Access Document

Files:: Rocher_et_al_2025_Measuring_what_matters.pdf

(Preview, Accepted manuscript, pdf, 1.1MB, Terms of use)

Publication website:: https://neurips.cc/virtual/2025/loc/san-diego/poster/121477

Authors

+ Bean, A More by this author

Institution:: University of Oxford
Division:: SSD
Department:: Oxford Internet Institute
Role:: Author
ORCID:: 0000-0001-8439-5975

+ Kearns, RO More by this author

Institution:: University of Oxford
Division:: SSD
Department:: Oxford Internet Institute
Role:: Author

+ Romanou. A More by this author

Role:: Author

+ Hafner, F More by this author

Institution:: University of Oxford
Division:: SSD
Department:: Oxford Internet Institute
Role:: Author

+ Mayne, H More by this author

Institution:: University of Oxford
Division:: SSD
Department:: Oxford Internet Institute
Role:: Author
ORCID:: 0000-0001-5506-3509

More authors...

+ UK Research and Innovation More from this funder

Funder identifier:: https://ror.org/001aqnf71
Grant:: MR/Y015711/1

Publisher:: NeurIPS
Publication date:: 2025-12-04
Acceptance date:: 2025-09-18
Event title:: 39th Annual Conference on Neural Information Processing Systems (NeurIPS 2025)
Event location:: San Diego, CA, USA and New Mexico, Mexico
Event website:: https://neurips.cc/
Event start date:: 2025-12-02
Event end date:: 2025-12-07

Language:: English
Pubs id:: 2346381
Local pid:: pubs:2346381
Deposit date:: 2025-12-05

Terms of use

Copyright holder:: Bean et al
Notes:: This paper was presented at the 39th Annual Conference on Neural Information Processing Systems (NeurIPS 2025), 2nd-7th December 2025, San Diego, CA, USA and New Mexico, Mexico.
The author accepted manuscript (AAM) of this paper has been made available under the University of Oxford's Open Access Publications Policy, and a CC BY public copyright licence has been applied.

Licence:: CC Attribution (CC BY)

Views and Downloads

About views and downloads

If you are the owner of this record, you can report an update to it here: Report update to this record

Conference item

Measuring what matters: construct validity in large language model benchmarks

Actions

Access Document

Authors

Terms of use

Views and Downloads

Altmetrics

Dimensions

Conference item

Measuring what matters: construct validity in large language model benchmarks

Actions

Access Document

Authors

Funding

Bibliographic Details

Item Description

Terms of use

Metrics

Views and Downloads

Altmetrics

Dimensions