Article
BLM-Hum1800: A Human-Validated Benchmark for Linguistic Competence in English and French
2026-08-21
Abstract excerpt
<p>Evaluating the fine-grained linguistic competence of Large Language Models (LLMs) requires diagnosticbenchmarks anchored by controlled syntactic paradigms and human behavioral baselines. Here, wepresent a bilingual dataset comprising 1,800 human-validated linguistic problem instances across Englishand French. Based on the Blackbird Language Matrices (BLM) framework, the dataset targets three coregrammatical phe...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- f7612927-cbef-5ad3-a3f0-4675a0f6e745
- DOI
- 10.31234/osf.io/r3egu_v1
