VideoLLaMA2

Video processor

An audio-visual language model designed to advance spatial-temporal modeling and audio understanding in video processing.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

GitHub

957 stars
11 watching
62 forks
Language: Python
last commit: almost 2 years ago

Related projects:

RepositoryDescriptionStars
showlab/vlogTransforms video content into a long document containing visual and audio information that can be used for chat or other applications.545
aspiers/ly2videoConverts music represented by a GNU LilyPond file into a video containing a horizontally scrolling music staff synchronized with audio rendering.158
dvlab-research/llama-vidAn image-based language model that uses large language models to generate visual and text features from videos748
dcdmllm/momentorA video Large Language Model designed for fine-grained comprehension and localization in videos with a custom Temporal Perception Module for improved temporal modeling58
showlab/show-1This project enables text-to-video generation using a combination of pixel and latent diffusion models.1,110
damo-nlp-sg/llm-zooA collection of information about various large language models used in natural language processing272
mbzuai-oryx/video-chatgptA video conversation model that generates meaningful conversations about videos using large vision and language models1,246
vpgtrans/vpgtransTransfers visual prompt generators across large language models to reduce training costs and enable customization of multimodal LLMs270
singularity42/vgan-tensorflowAn implementation of a deep learning model to generate videos with dynamic scenes15
nus-hpc-ai-lab/videosysA comprehensive toolkit for high-performance video generation and processing1,819
rupertluo/valleyAn offline video assistant system powered by large language models and computer vision techniques.210
lxtgh/omg-segDevelops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.1,336
boheumd/ma-lmmThis project develops an AI model for long-term video understanding254
bryandlee/tune-a-videoUnofficial implementation of a deep learning model to generate or modify video content191
radi-cho/datasetgptA command-line interface to generate textual datasets with Large Language Models293