Article
HAR-Agent: Multilingual Multimodal Activity Recognition via Knowledge-Distilled LLM Reasoning
2026-03-31
Abstract excerpt
We present HAR-Agent, a multilingual multimodal human activity recognition system that unifies visual, audio, and text perception through a common textual representation, with a large language model serving as the central decision-maker. We build a multimodal agent architecture and systematically compare 17 model configurations-varying parameter count from 1.5 billion to 72 billion and training method between inst...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 35e1e52f-a3ff-566c-b409-0f004085e952
- DOI
- 10.22541/au.177499021.11495152/v1
