LLaMA-VID
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models (ECCV 2024)
AI summary
Video image processor
An image-based language model that uses large language models to generate visual and text features from videos
- stars
- 748
- forks
- 45
- watching
- 14
Similar projects
Found by comparing what the projects do, not just their names.
Image processor
An all-in-one demo for interactive image processing and generation
Image segmentation tool
A system that uses large language models to generate segmentation masks for images based on complex queries and world knowledge.
Visual Prompt Model
A system designed to enable large multimodal models to understand arbitrary visual prompts
Image editing assistant
An implementation of a multimodal generation assistant using large language models and various image editing techniques.
Visual decoder
A large language model designed to process and generate visual information
Image processor
A system for scaling large language models to process and understand visual information from multiple images efficiently.
Image understanding model
A large multi-modal model developed using the Llama3 language model, designed to improve image understanding capabilities.
Video processor
An audio-visual language model designed to advance spatial-temporal modeling and audio understanding in video processing.
Multimodal LLM
An implementation of a multimodal language model with capabilities for comprehension and generation
VQA model
A multimodal LLM designed to handle text-rich visual questions
libav/libav1.1K
Multimedia processor
A collection of libraries and tools for processing multimedia content
Video processor
A web-based video processing tool that uses AI to facilitate cultural and linguistic tasks such as transcription, translation, and audio synthesis.
Image processor
A Lua binding for a fast image processing library with low memory needs.
Image processing library
A library of fast computer vision algorithms implemented in C++ for speed, operating over numpy arrays.
Text controller
An interactive control system for text generation in multi-modal language models