awesome-action-recognition

Action Recognition Resources

A curated collection of resources and research papers on action recognition and video understanding techniques.

A curated list of action recognition and related area resources

GitHub

4k stars
207 watching
724 forks
last commit: over 3 years ago
Linked from 3 awesome lists

action-classificationaction-detectionaction-recognitionactivity-recognitionactivity-understandingawesomeawesome-listobject-recognitionpose-estimationvideo-processingvideo-recognitionvideo-understanding

Awesome Action Recognition: / Action Recognition and Video Understanding / Summary posts

Deep Learning for Videos: A 2018 Guide to Action RecognitionSummary of major landmark action recognition research papers till 2018
Literature Survey: Human Action RecognitionBrief human action recognition literature survey of work published between 2014 and 2019

Awesome Action Recognition: / Action Recognition and Video Understanding / Video Representation

Why Can't I Dance in the Mall? Learning to Mitigate Scene Bias in Action RecognitionJ. Choi et al., NeurIPS2019
SlowFast Networks for Video RecognitionC. Feichtenhofer et al., ICCV2019
Large-scale weakly-supervised pre-training for video action recognitionD. Ghadiyaram et al., arXiv2019
Video Classification with Channel-Separated Convolutional NetworksD. Tran et al., arXiv2019
DistInit: Learning Video Representations without a Single Labeled VideoR. Girdhar et al., arXiv2019
SCSampler: Sampling Salient Clips from Video for Efficient Action RecognitionB. Korbar et al., arXiv2019
Video Action Transformer NetworkR. Girdhar et al., CVPR2019
Learning Correspondence from the Cycle-consistency of TimeX. Wang et al., CVPR2019
Representation Flow for Action RecognitionAJ. Piergiovanni and M. S. Ryoo et al., CVPR2019
Collaborative Spatiotemporal Feature Learning for Video Action RecognitionC. Li et al., CVPR2019
Learning Video Representations from Correspondence ProposalsX. Liu et al., CVPR2019
Timeception for Complex Action RecognitionN. Hussein et al., CVPR2019
The Visual Centrifuge: Model-Free Layered Video RepresentationsJ.-B. Alayrac et al., CVPR2019
Long-Term Feature Banks for Detailed Video UnderstandingC.-Y. Wu. et al., CVPR2019
Temporal Relational Reasoning in VideosB. Zhou et al., ECCV2018
Action Recognition Zoo244over 7 years ago- Codes for popular action recognition models, written based on pytorch, verified on the something-something dataset
Videos as Space-Time Region GraphsX. Wang and A. Gupta, ECCV2018
Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?K. Hara et al., CVPR2019
A Closer Look at Spatiotemporal Convolutions for Action RecognitionD. Tran et al., CVPR2018
Attend and Interact: Higher-Order Object Interactions for Video UnderstandingCY. Ma et al., CVPR 2018
Non-Local Neural NetworksX. Wang et al., CVPR2018
Rethinking Spatiotemporal Feature Learning For Video UnderstandingS. Xie et al., arXiv2017
ConvNet Architecture Search for Spatiotemporal Feature LearningD. Tran et al, arXiv2017. Note: Aka Res3D. : In the repository, C3D-v1.1 is the Res3D implementation
Learning Spatio-Temporal Representation with Pseudo-3D Residual NetworksZ. Qui et al, ICCV2017
Quo Vadis, Action Recognition? A New Model and the Kinetics DatasetJ. Carreira et al, CVPR2017. ,
Learning Spatiotemporal Features with 3D Convolutional NetworksD. Tran et al, ICCV2015. Note: Aka C3D. Note that the official caffe does not support python wrapper. , , , : ,
Deep Temporal Linear Encoding NetworksA. Diba et al, CVPR2017
Temporal Convolutional Networks: A Unified Approach to Action Segmentation and DetectionC. Lea et al, CVPR 2017
Long-term Temporal ConvolutionsG. Varol et al, TPAMI2017
Temporal Segment Networks: Towards Good Practices for Deep Action RecognitionL. Wang et al, arXiv 2016
Convolutional Two-Stream Network Fusion for Video Action RecognitionC. Feichtenhofer et al, CVPR2016
Two-Stream Convolutional Networks for Action Recognition in VideosK. Simonyan and A. Zisserman, NIPS2014
Temporal Recurrent Networks for Online Action DetectionM. Xu et al, ICCV2019
Long Short-Term Transformer for Online Action DetectionM. Xu et al, Neurips2021
[3D ResNet PyTorch]3,920over 5 years ago
[PyTorch Video Research]533over 7 years ago
[M-PACT: Michigan Platform for Activity Classification in Tensorflow]107over 7 years ago
[Inflated models on PyTorch]148over 5 years ago
[I3D models transfered from Tensorflow to PyTorch]532over 2 years ago
[A Two Stream Baseline on Kinectics dataset]42over 7 years ago
[MMAction]1,863over 4 years ago
[MMAction2]4,360about 2 years ago
[PySlowFast]6,680almost 2 years ago
[Decord]1,923about 2 years agoEfficient video reader for python
[I3D models converted from Tensorflow to Core ML]24about 6 years ago
[Extract frame and optical-flow from videos, #docker]133about 4 years ago
[NVIDIA-DALI, video loading pipelines]
[NVIDIA optical-flow SDK]

Awesome Action Recognition: / Action Recognition and Video Understanding / Action Classification

Guided Weak Supervision for Action Recognition with Scarce Data to Assess Skills of Children with AutismP. Pandey et al, AAAI 2020
Neural Graph Matching Networks for Fewshot 3D Action RecognitionM. Guo et al., ECCV2018
Temporal 3D ConvNets using Temporal Transition LayerA. Diba et al., CVPRW2018
Temporal 3D ConvNets: New Architecture and Transfer Learning for Video ClassificationA. Diba et al., arXiv2017
Attentional Pooling for Action RecognitionR. Girdhar and D. Ramanan, NIPS2017
Fully Context-Aware Video PredictionByeon et al, arXiv2017
Hidden Two-Stream Convolutional Networks for Action RecognitionY. Zhu et al, arXiv2017
Dynamic Image Networks for Action RecognitionH. Bilen et al, CVPR2016
Long-term Recurrent Convolutional Networks for Visual Recognition and DescriptionJ. Donahue et al, CVPR2015
Describing Videos by Exploiting Temporal StructureL. Yao et al, ICCV2015. note: from the same group of RCN paper “Delving Deeper into Convolutional Networks for Learning Video Representations"
Two-Stream SR-CNNs for Action Recognition in VideosL. Wang et al, BMVC2016
Real-time Action Recognition with Enhanced Motion Vector CNNsB. Zhang et al, CVPR2016
Action Recognition with Trajectory-Pooled Deep-Convolutional DescriptorsL. Wang et al, CVPR2015

Awesome Action Recognition: / Action Recognition and Video Understanding / Skeleton-Based Action Classification

Actional-Structural Graph Convolutional Networks for Skeleton-Based Action RecognitionM. Li et al., CVPR2019
An Attention Enhanced Graph Convolutional LSTM Network for Skeleton-Based Action RecognitionC. Si et al., CVPR2019
View Adaptive Neural Networks for High Performance Skeleton-Based Human Action RecognitionP. Zhang et al., TPAMI2019
Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action RecognitionS. Yan et al., AAAI2018
Deep Progressive Reinforcement Learning for Skeleton-Based Action RecognitionY. Tang et al., CVPR2018
Co-occurrence Feature Learning from Skeleton Data for Action Recognition and Detection with Hierarchical AggregationC. Li et al., IJCAI2018
Part-based Graph Convolutional Network for Action RecognitionK. Thakkar et al., BMVC2018

Awesome Action Recognition: / Action Recognition and Video Understanding / Temporal Action Detection

Rethinking the Faster R-CNN Architecture for Temporal Action LocalizationYu-Wei Chao et al., CVPR2018
Weakly Supervised Action Localization by Sparse Temporal Pooling NetworkPhuc Nguyen et al., CVPR 2018
Temporal Deformable Residual Networks for Action Segmentation in VideosP. Lei and S. Todrovic., CVPR2018
End-to-End, Single-Stream Temporal Action Detection in Untrimmed VideosShayamal Buch et al., BMVC 2017
Cascaded Boundary Regression for Temporal Action DetectionJiyang Gao et al., BMVC 2017 [ ]
Temporal Tessellation: A Unified Approach for Video AnalysisKaufman et al., ICCV2017
Temporal Action Detection with Structured Segment NetworksY. Zhao et al., ICCV2017
Temporal Context Network for Activity Localization in VideosX. Dai et al., ICCV2017
Detecting the Moment of Completion: Temporal Models for Localising Action CompletionF. Heidarivincheh et al., arXiv2017
CDC: Convolutional-De-Convolutional Networks for Precise Temporal Action Localization in Untrimmed VideosZ. Shou et al, CVPR2017
SST: Single-Stream Temporal Action ProposalsS. Buch et al, CVPR2017
R-C3D: Region Convolutional 3D Network for Temporal Activity DetectionH. Xu et al, arXiv2017
DAPs: Deep Action Proposals for Action UnderstandingV. Escorcia et al, ECCV2016
Online Action Detection using Joint Classification-Regression Recurrent Neural NetworksY. Li et al, ECCV2016. Noe: RGB-D Action Detection
Temporal Action Localization in Untrimmed Videos via Multi-stage CNNsZ. Shou et al, CVPR2016. Note: Aka S-CNN
Fast Temporal Activity Proposals for Efficient Detection of Human Actions in Untrimmed VideosF. Heilbron et al, CVPR2016. Note: Depends on , aka SparseProp
Actionness Estimation Using Hybrid Fully Convolutional NetworksL. Wang et al, CVPR2016. Note: The code is not a complete verision. It only contains a demo, not training
Learning Activity Progression in LSTMs for Activity Detection and Early DetectionS. Ma et al, CVPR2016
End-to-end Learning of Action Detection from Frame Glimpses in VideosS. Yeung et al, CVPR2016. Note: This method uses reinforcement learning
Fast Action Proposals for Human Action Detection and SearchG. Yu and J. Yuan, CVPR2015. Note: code for FAP is NOT available online. Note: Aka FAP
Bag-of-fragments: Selecting and encoding video fragments for event detection and recountingP. Mettes et al, ICMR2015
Action localization in videos through context walkK. Soomro et al, ICCV2015

Awesome Action Recognition: / Action Recognition and Video Understanding / Spatio-Temporal Action Detection

A Better Baseline for AVAR. Girdhar et al., ActivityNet Workshop, CVPR2018
Real-Time End-to-End Action Detection with Two-Stream NetworksA. El-Nouby and G. Taylor, arXiv2018
Human Action Localization with Sparse Spatial SupervisionP. Weinzaepfel et al., arXiv2017
Unsupervised Action Discovery and Localization in VideosK. Soomro and M. Shah, ICCV2017
Spatial-Aware Object Embeddings for Zero-Shot Localization and Classification of ActionsP. Mettes and C. G. M. Snoek, ICCV2017
Action Tubelet Detector for Spatio-Temporal Action LocalizationV. Kalogeiton et al, ICCV2017
Tube Convolutional Neural Network (T-CNN) for Action Detection in Videoset al, ICCV2017
Chained Multi-stream Networks Exploiting Pose, Motion, and Appearance for Action Classification and DetectionM. Zolfaghari et al, ICCV2017
TORNADO: A Spatio-Temporal Convolutional Regression Network for Video Action ProposalH. Zhu et al., ICCV2017
Online Real time Multiple Spatiotemporal Action Localisation and Predictionet al, ICCV2017
AMTnet: Action-Micro-Tube regression by end-to-end trainable deep architectureS. Saha et al, ICCV2017
Am I Done? Predicting Action Progress in VideosF. Becattini et al, BMVC2017
Generic Tubelet Proposals for Action LocalizationJ. He et al, arXiv2017
Incremental Tube Construction for Human Action DetectionH. S. Behl et al, arXiv2017
Multi-region two-stream R-CNN for action detectionand C. Schmid. ECCV2016
Spot On: Action Localization from Pointly-Supervised ProposalsP. Mettes et al, ECCV2016
Deep Learning for Detecting Multiple Space-Time Action Tubes in VideosS. Saha et al, BMVC2016
Learning to track for spatio-temporal action localizationP. Weinzaepfel et al. ICCV2015
Action detection by implicit intentional motion clusteringW. Chen and J. Corso, ICCV2015
Finding Action TubesG. Gkioxari and J. Malik CVPR2015
APT: Action localization proposals from dense trajectoriesJ. Gemert et al, BMVC2015
Spatio-Temporal Object Detection ProposalsD. Oneata et al, ECCV2014
Action localization with tubelets from motionM. Jain et al, CVPR2014
Spatiotemporal deformable part models for action detectionet al, CVPR2013
Action localization in videos through context walkK. Soomro et al, ICCV2015
Fast Action Proposals for Human Action Detection and SearchG. Yu and J. Yuan, CVPR2015. Note: code for FAP is NOT available online. Note: Aka FAP

Awesome Action Recognition: / Action Recognition and Video Understanding / Ego-Centric Action Recognition

Actor and Observer: Joint Modeling of First and Third-Person VideosG. Sigurdsson et al., CVPR2018

Awesome Action Recognition: / Action Recognition and Video Understanding / Miscellaneous

What and How Well You Performed? A Multitask Learning Approach to Action Quality AssessmentP. Parma and B. T. Morris. CVPR2019
PathTrack: Fast Trajectory Annotation with Path SupervisionS. Manen et al., ICCV2017
CortexNet: a Generic Network Family for Robust Visual Temporal RepresentationsA. Canziani and E. Culurciello - arXiv2017
Slicing Convolutional Neural Network for Crowd Video UnderstandingJ. Shao et al, CVPR2016
Two-Stream (RGB and Flow) pretrained model weights26almost 10 years ago

Awesome Action Recognition: / Action Recognition and Video Understanding / Action Recognition Datasets

Video Dataset Overview from Antoine Miech
HACS
Moments in Time,
AVA, , for missing videos
Kinetics, ,
OOPSA dataset of unintentional action,
COINa large-scale dataset for comprehensive instructional video analysis,
YouTube-8M,
YouTube-BB,
DALYDaily Action Localization in Youtube videos. Note: Weakly supervised action detection dataset. Annotations consist of start and end time of each action, one bounding box per each action per video
20BN-JESTER,
ActivityNetNote: They provide a download script and evaluation code
Charades
Charades-Ego, - First person and third person video aligned dataset
EPIC-Kitchens, - First person videos recorded in kitchens. Note they provide download scripts and a python library
Sports-1MLarge scale action recognition dataset
THUMOS14Note: It overlaps with dataset
THUMOS15Note: It overlaps with dataset
HOLLYWOOD2:
UCF-101, , and , and . And there are also some pre-computed spatiotemporal action detection
UCF-50
UCF-Sports, note: the train/test split link in the official website is broken. Instead, you can download it from
HMDB
J-HMDB
LIRIS-HARL
KTH
MSR ActionNote: It overlaps with datset
Sports Videos in the Wild
NTU RGB+D763over 4 years ago
Mixamo Mocap Dataset
UWA3D Multiview Activity II Dataset
Northwestern-UCLA Dataset
SYSU 3D Human-Object Interaction Dataset
MEVA (Multiview Extended Video with Activities) Dataset

Awesome Action Recognition: / Action Recognition and Video Understanding / Video Annotation

Efficiently scaling up crowdsourced video annotationC. Vondrick et. al, IJCV2013
The Design and Implementation of ViPERD. Mihalcik and D. Doermann, Technical report
VTT: Visual Object Tagging Tool4,331almost 5 years ago. Modern app to annotate objects in videos and images. It facilitates the development of an end-to-end machine learning pipeline encompassing the annotation/export/import of assets. Moreover, it could run as a native app or via web
VIA: VGG Image Annotator. Simple and standalone manual annotation web-app for image, audio and video. It runs in the web browser and does not require any installation or setup

Awesome Action Recognition: / Object Recognition / Object Detection

Deformable Convolutional NetworksJ. Dai et al., ICCV2017
Detectron26,295almost 3 years agoOpen Source Object Detection Framework from Facebook AI Research. Includes Mask R-CNN, FPN, and etc. Caffe2 implementation
Mask R-CNNK. He et al, , , , , - State-of-the-art object detection/instance segmentation algorithm
Faster R-CNNS. Ren et al, NIPS2015. , , , - State-of-the-art object detector
YOLOJ. Redmon et al, CVPR2016. , - Fast object detector
YOLO9000J. Redmon and A. Farhadi, CVPR2017. - State-of-the-art object detector which can detect 9000 objects in realtime
SSDW. Liu et al, ECCV2016. , , - State-of-the-art object detector with realtime processing speed
RetinaNetTsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He and Piotr Dollár, Facebook AI Research FAIR & ICCV 2017. - State-of-the-art object detector with realtime processing speed

Awesome Action Recognition: / Object Recognition / Video Object Detection

[code]553about 8 years ago[Detect to Track and Track to Detect] - C. Feichtenhofer et al., ICCV2017. ,
[code]724almost 5 years ago[Flow-Guided Feature Aggregation for Video Object Detection] - X. Zhu et al., ICCV2017. , aka FGFA

Awesome Action Recognition: / Object Recognition / Video Object Detection Datasets

ImageNet VID
YouTube-8M,
YouTube-BB,

Awesome Action Recognition: / Pose Estimation / Pose Estimation

AlphaPose8,084over 2 years agoPyTorch based realtime and accurate pose estimation and tracking tool from SJTU
Detect-and-Track: Efficient Pose Estimation in VideosR. Girdhar et al., arXiv2017
OpenPose Library31,463about 2 years agoCaffe based realtime pose estimation library from CMU
Realtime Multi-Person 2D Pose Estimation using Part Affinity FieldsZ. Cao et al, CVPR2017. depends on the - Earlier version of OpenPose from CMU
DensePoseDense pose human estimation in the wild implemented in the Detectron framework
MultiPoseNet: Fast Multi-Person Pose Estimation using Pose Residual NetworkM. Kocabas et al, ECCV2018
DeepLabCut: markerless pose estimation of user-defined body parts with deep learningA. Mathis et al, Nature Neuroscience 2018

Awesome Action Recognition: / Competitions / Competitions

ActEV (Activities in Extended VideoActivity detection in security camera videos. Runs through 2021. Hosted by NIST

Backlinks from these awesome lists:

More related projects: