VILA
by NVlabs
Pythonpushed almost 2 years ago
VILA - a multi-image visual language model with training, inference and evaluation recipe, deployable from cloud to edge (Jetson Orin and laptops)
AI summary
Video understanding framework
A visual language model that leverages pre-trained models and large-scale training to understand video and images, enabling applications like video reasoning and in-context learning.
- stars
- 2.1K
- forks
- 168
- watching
- 32