Article
Benchmarking large language models for ACMG/AMP variant interpretation and variant calling
2026-07-05
Abstract excerpt
Agentic large language models are increasingly used across the genomic workflow, from variant calling to clinical interpretation, yet they are evaluated by accuracy alone, a single figure that cannot say whether a system is safe or where in the workflow a failure originates. We present ClawBench, a framework that attributes each outcome to the architectural layer that produced it across both halves of the canonica...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 27006277-8463-5c5b-9f7b-3b71b4d42169
- DOI
- 10.64898/2026.06.30.735646
