Article
Mutation Testing of Task-Scoped State Oracles in Software-Agent Benchmarks: A Cross-Benchmark Empirical Study
2026-08-18
Abstract excerpt
<title>Abstract</title> <p>Stateful software-agent benchmarks increasingly use executable evaluators to decide whether an agent produced the requested persistent effects. Yet a high task score is informative only if the evaluator rejects harmful extra effects without rejecting equivalent representations. We introduce a deterministic mutation-testing protocol for calibrating such task-scoped state oracles. Startin...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- a2eff2ad-8323-52d2-b63c-f902d9ecc15c
- DOI
- 10.21203/rs.3.rs-10665114/v1
