Back to search

Article

VT-BPAN: Vision Transformer-based bilinear pooling and attention network fusion of RGB and skeleton features for human action recognition

2023-03-06

Abstract excerpt

Recent generation Microsoft Kinect Camera captures a series of multimodal signals that provide RGB video, depth sequences, and skeleton information, thus it becomes an option to achieve enhanced human action recognition performance by fusing different data modalities. However, most existing fusion methods simply fuse different features, which ignores the underlying semantics between different models, leading to a...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
066a37fc-abb2-53fc-8748-523be8faba66
DOI
10.21203/rs.3.rs-2627627/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
VT-BPAN: Vision Transformer-based bilinear pooling and attention network fusion of RGB and skeleton features for human action recognitionDOI 10.21203/rs.3.rs-2627627/v1
Select a neighboring publication to make it the new centre.