Video-LLaVA
【EMNLP 2024🔥】Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
AI summary
Video generator
A deep learning framework for generating videos from text inputs and visual features.
- stars
- 3.1K
- forks
- 219
- watching
- 29
Similar projects
Found by comparing what the projects do, not just their names.
haotian-liu/llava20.7K
Visual Instruction System
A system that uses large language and vision models to generate and process visual instructions
Multimodal model developer
Develops large multimodal models for various computer vision tasks including image and video analysis
Multimodal alignment model
Extending pretraining models to handle multiple modalities by aligning language and video representations
Video understanding model
An audio-visual language model designed to understand and respond to video content with improved instruction-following capabilities
Mixture of Experts Model
A large vision-language model using a mixture-of-experts architecture to improve performance on multi-modal learning tasks
Model debiasing
Debiasing techniques to minimize hallucinations in large visual language models
Video analysis toolkit
A comprehensive video understanding toolbox and benchmark with modular design, supporting various tasks such as action recognition, localization, and retrieval.
Vision-LM Framework
An open-source framework for training large language models with vision capabilities.
Instruction-following model tuner
An implementation of a method for fine-tuning language models to follow instructions with high efficiency and accuracy
LLaMA Tuner
A tool for efficiently fine-tuning large language models across multiple architectures and methods.
luodian/otter3.6K
Multi-modal AI model
A multi-modal AI model developed for improved instruction-following and in-context learning, utilizing large-scale architectures and various training datasets.
x-plug/mplug-owl2.4K
Visual AI model
Develops large language models that can understand and generate human-like visual and video content
Video benchmarking toolkit
Evaluates and benchmarks large language models' video understanding capabilities
Server framework
A fast serving framework for large language models and vision language models.
Visual unification framework
A framework for unified visual representation in image and video understanding models, enabling efficient training of large language models on multimodal data.