Back to search

Article

FHIR-AgentEval: A Modular Sandbox for Benchmarking Clinical LLM Agents with an Evaluation of Memory-Augmented Configurations

2026-02-09

Abstract excerpt

<title>Abstract</title> <p>Healthcare data exchange increasingly relies on HL7 FHIR, but FHIR's implementation complexity creates barriers for clinical workflows. Large language model (LLM) agents could bridge this gap by translating natural language requests into structured FHIR operations, yet their reliability remains unproven. We present FHIR-AgentEval, an extensible evaluation sandbox comprising 43 modular t...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
d2dad0fa-06cd-5e8b-ad65-bb8f6439e0f7
DOI
10.21203/rs.3.rs-8746188/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
FHIR-AgentEval: A Modular Sandbox for Benchmarking Clinical LLM Agents with an Evaluation of Memory-Augmented ConfigurationsDOI 10.21203/rs.3.rs-8746188/v1
Select a neighboring publication to make it the new centre.