Article
A Multi-Model, Multi-Domain Benchmark of Large Language Model Agreement with Humans and with Each Other in Sentiment Classification
2026-08-18
Abstract excerpt
<title>Abstract</title> <p>Large language models (LLMs) are widely used to classify sentiment, yet cross-provider comparisons on the same data are rare and inter-model agreement is largely unmeasured. We evaluate five LLMs from four providers (Claude Opus 4.7, GPT-4o, GPT-5.5, Llama 3.1 8B, and Gemini 2.5 Pro), with VADER (a rule-based lexicon) and RoBERTa (a supervised transformer) as a baseline, on five human-a...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 04c54625-49a3-526e-8cbc-a7c400b551a6
- DOI
- 10.21203/rs.3.rs-10415268/v1
