Back to search

Article

Mutation Testing of Task-Scoped State Oracles in Software-Agent Benchmarks: A Cross-Benchmark Empirical Study

2026-08-18

Abstract excerpt

<title>Abstract</title> <p>Stateful software-agent benchmarks increasingly use executable evaluators to decide whether an agent produced the requested persistent effects. Yet a high task score is informative only if the evaluator rejects harmful extra effects without rejecting equivalent representations. We introduce a deterministic mutation-testing protocol for calibrating such task-scoped state oracles. Startin...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
a2eff2ad-8323-52d2-b63c-f902d9ecc15c
DOI
10.21203/rs.3.rs-10665114/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Mutation Testing of Task-Scoped State Oracles in Software-Agent Benchmarks: A Cross-Benchmark Empirical StudyDOI 10.21203/rs.3.rs-10665114/v1
Select a neighboring publication to make it the new centre.