Article
VT-BPAN: Vision Transformer-based bilinear pooling and attention network fusion of RGB and skeleton features for human action recognition
2023-03-06
Abstract excerpt
Recent generation Microsoft Kinect Camera captures a series of multimodal signals that provide RGB video, depth sequences, and skeleton information, thus it becomes an option to achieve enhanced human action recognition performance by fusing different data modalities. However, most existing fusion methods simply fuse different features, which ignores the underlying semantics between different models, leading to a...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 066a37fc-abb2-53fc-8748-523be8faba66
- DOI
- 10.21203/rs.3.rs-2627627/v1
