VisionLLM
by OpenGVLab
VisionLLM Series
AI summary
Visual decoder
A large language model designed to process and generate visual information
- stars
- 956
- forks
- 29
- watching
- 45
Similar projects
Found by comparing what the projects do, not just their names.
LLM trainer
Transfers visual prompt generators across large language models to reduce training costs and enable customization of multimodal LLMs
Conversational interface
An interactive tool that connects multiple visual models and an LLM to facilitate text-based conversations.
Image segmentation tool
A system that uses large language models to generate segmentation masks for images based on complex queries and world knowledge.
Data comprehension tool
A research project that develops tools and models for understanding visual data in the open world, enabling applications such as image-text retrieval and relation comprehension.
lxtgh/omg-seg1.3K
Visual Model
Develops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.
Video image processor
An image-based language model that uses large language models to generate visual and text features from videos
nvlabs/prismer1.3K
Vision-Language Model
A deep learning framework for training multi-modal models with vision and language capabilities.
LLM
An open-source implementation of a vision-language instructed large language model
Long context transfer
An open-source project that enables the transfer of language understanding to vision capabilities through long context processing.
Task solver
An open-source framework for augmenting large language models with tools by searching on graphs to solve complex real-world tasks.
Language model toolkit
A guide to using pre-trained large language models in source code analysis and generation
Image understanding model
A large multi-modal model developed using the Llama3 language model, designed to improve image understanding capabilities.
Visual encoder
A vision-language model that uses a query transformer to encode images as visual tokens and allows flexible choice of the number of visual tokens.
openbmb/viscpm1.1K
Multimodal Models
A family of large multimodal models supporting multimodal conversational capabilities and text-to-image generation in multiple languages
Multimodal LLM
An implementation of a multimodal language model with capabilities for comprehension and generation