Back to search

Article

Benchmarking large language models for ACMG/AMP variant interpretation and variant calling

2026-07-05

Abstract excerpt

Agentic large language models are increasingly used across the genomic workflow, from variant calling to clinical interpretation, yet they are evaluated by accuracy alone, a single figure that cannot say whether a system is safe or where in the workflow a failure originates. We present ClawBench, a framework that attributes each outcome to the architectural layer that produced it across both halves of the canonica...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
27006277-8463-5c5b-9f7b-3b71b4d42169
DOI
10.64898/2026.06.30.735646
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Benchmarking large language models for ACMG/AMP variant interpretation and variant callingDOI 10.64898/2026.06.30.735646
Select a neighboring publication to make it the new centre.