InternVideo
by OpenGVLab
Pythonpushed almost 2 years ago
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
AI summary
Video foundation models
Develops general video foundation models and related datasets for multimodal understanding and generation through generative and discriminative learning.
- stars
- 1.5K
- forks
- 91
- watching
- 27
- #action-recognition
- #benchmark
- #contrastive-learning
- #foundation-models
- #instruction-tuning
- #masked-autoencoder
- #multimodal
- #open-set-recognition
- #self-supervised
- #spatio-temporal-action-localization
- #temporal-action-localization
- #video-clip
- #video-data
- #video-dataset
- #video-question-answering
- #video-retrieval
- #video-understanding
- #vision-transformer
- #zero-shot-classification
- #zero-shot-retrieval