Article
Benchmarking System Configurations for Rubric-Constrained Personalized Exercise Planning
2026-08-20
Abstract excerpt
Agentic Large Language Model (LLM) systems that turn wearable streams into daily workout plans need matched-backbone tests with shared post-processing. We compare four system configurations for daily workout generation—Baseline-LLM (raw chart, no precompute), Single Agent, Multi-Agent, and ReAct—on N = 50 matched user-days (5 users × 10 days) from longitudinal wearable data. Four blind LLM judges (claude-opus-4-5,...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 4b70988b-b846-55fd-a240-a9290898f689
- DOI
- 10.20944/preprints202608.1437.v1
