Article
Mental Health Benchmarks for Large Language Models: A Systematic Scoping Review
2026-08-01
Abstract excerpt
<p>Large language models (LLMs) are increasingly used for mental health support, yet standards for evaluating their safety and competence remain unsettled. Benchmarks, standardized tests with predetermined scoring criteria, are the most common evaluation approach, but this landscape has not been systematically mapped. In this systematic scoping review, we searched four databases, machine-learning repositories, and...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 40062bc4-6b15-56e7-a22c-9d8e2f916907
- DOI
- 10.31234/osf.io/fahuk_v1
