Article
FHIR-AgentEval: A Modular Sandbox for Benchmarking Clinical LLM Agents with an Evaluation of Memory-Augmented Configurations
2026-02-09
Abstract excerpt
<title>Abstract</title> <p>Healthcare data exchange increasingly relies on HL7 FHIR, but FHIR's implementation complexity creates barriers for clinical workflows. Large language model (LLM) agents could bridge this gap by translating natural language requests into structured FHIR operations, yet their reliability remains unproven. We present FHIR-AgentEval, an extensible evaluation sandbox comprising 43 modular t...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- d2dad0fa-06cd-5e8b-ad65-bb8f6439e0f7
- DOI
- 10.21203/rs.3.rs-8746188/v1
