Back to search

Article

A Multi-Model, Multi-Domain Benchmark of Large Language Model Agreement with Humans and with Each Other in Sentiment Classification

2026-08-18

Abstract excerpt

<title>Abstract</title> <p>Large language models (LLMs) are widely used to classify sentiment, yet cross-provider comparisons on the same data are rare and inter-model agreement is largely unmeasured. We evaluate five LLMs from four providers (Claude Opus 4.7, GPT-4o, GPT-5.5, Llama 3.1 8B, and Gemini 2.5 Pro), with VADER (a rule-based lexicon) and RoBERTa (a supervised transformer) as a baseline, on five human-a...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
04c54625-49a3-526e-8cbc-a7c400b551a6
DOI
10.21203/rs.3.rs-10415268/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
A Multi-Model, Multi-Domain Benchmark of Large Language Model Agreement with Humans and with Each Other in Sentiment ClassificationDOI 10.21203/rs.3.rs-10415268/v1
Select a neighboring publication to make it the new centre.