Article
How Well Do Large Language Models Agree with Humans — and Each Other? A Multi-Model, Multi-Domain Sentiment Benchmark Study
2026-06-23
Abstract excerpt
<title>Abstract</title> <p>Large language models (LLMs) are now widely used to classify sentiment. Yet we rarely compare models from different providers on the same data, and we know little about how often they agree with each other. We evaluate five LLMs from four providers (Claude Opus 4.7, GPT-4o, GPT-5.5, Llama 3.1 8B, and Gemini 2.5 Pro) and add VADER, a rule based lexicon, as a baseline. The test set has fi...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 51d74941-aabb-5521-8c4f-a44a7fef32eb
- DOI
- 10.21203/rs.3.rs-9960161/v1
