A comprehensive survey on RGB-D-based human action recognition: algorithms, datasets, and popular applications
This survey reviews CNN, RNN, and Transformer algorithms, datasets, and applications for RGB-D human action recognition.
Published Aug 13, 202516 citationsPaper ↗
Only vote on papers you've read. Sign in with GitHub to vote.
A comprehensive survey that uniquely maps CNN-RNN-Transformer RGB-D action recognition and datasets, though its taxonomy offers little synthesis and boilerplate future directions.
Abstract
Due to the rapid advances in computer vision and deep learning, human action recognition has become one of the most important representative tasks for video understanding. Especially for human action recognition based on RGB-D data, a promising research direction, there has been a number of researchers to work on. In particular, convolutional neural networks (CNNs) are capable of image classification tasks, recurrent neural networks (RNNs) are skilled in sequence-based problems, and Transformer is good at global modeling. In this survey, we introduce a number of algorithms based on CNNs, RNNs and Transformer for RGB-D based human action recognition, which could be categorized into four parts: RGB-based, depth-based, skeleton-based and RGB-D based. As a survey focusing on the RGB-D based human action recognition, we thoroughly represent the algorithms, datasets and popular applications for it. What’s more, we give some possible future research directions for this field in the last part.