Back to search

Article

HAR-Agent: Multilingual Multimodal Activity Recognition via Knowledge-Distilled LLM Reasoning

2026-03-31

Abstract excerpt

We present HAR-Agent, a multilingual multimodal human activity recognition system that unifies visual, audio, and text perception through a common textual representation, with a large language model serving as the central decision-maker. We build a multimodal agent architecture and systematically compare 17 model configurations-varying parameter count from 1.5 billion to 72 billion and training method between inst...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
35e1e52f-a3ff-566c-b409-0f004085e952
DOI
10.22541/au.177499021.11495152/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
HAR-Agent: Multilingual Multimodal Activity Recognition via Knowledge-Distilled LLM ReasoningDOI 10.22541/au.177499021.11495152/v1
Select a neighboring publication to make it the new centre.