Back to search

Article

How Well Do Large Language Models Agree with Humans — and Each Other? A Multi-Model, Multi-Domain Sentiment Benchmark Study

2026-06-23

Abstract excerpt

<title>Abstract</title> <p>Large language models (LLMs) are now widely used to classify sentiment. Yet we rarely compare models from different providers on the same data, and we know little about how often they agree with each other. We evaluate five LLMs from four providers (Claude Opus 4.7, GPT-4o, GPT-5.5, Llama 3.1 8B, and Gemini 2.5 Pro) and add VADER, a rule based lexicon, as a baseline. The test set has fi...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
51d74941-aabb-5521-8c4f-a44a7fef32eb
DOI
10.21203/rs.3.rs-9960161/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
How Well Do Large Language Models Agree with Humans — and Each Other? A Multi-Model, Multi-Domain Sentiment Benchmark StudyDOI 10.21203/rs.3.rs-9960161/v1
Select a neighboring publication to make it the new centre.