Article
TransMODAL: A Dual-Stream Transformer with Adaptive Co-Attention for Efficient Human Action Recognition
2025-07-29
Abstract excerpt
Human Action Recognition has seen significant advances through Transformer-based architectures, yet achieving nuanced understanding often requires fusing multiple data modalities. Standard models relying solely on RGB video can struggle with actions defined by subtle motion cues rather than appearance. This paper introduces TransMODAL, a novel dual-stream Transformer that synergistically fuses spatiotemporal appea...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 9084ab35-6f62-50b9-983d-0ee4cb87f1f2
- DOI
- 10.20944/preprints202507.2386.v1
