360VL
by 360CVGroup
AI summary
Image understanding model
A large multi-modal model developed using the Llama3 language model, designed to improve image understanding capabilities.
- stars
- 32
- forks
- 2
- watching
- 0
Similar projects
Found by comparing what the projects do, not just their names.
Video image processor
An image-based language model that uses large language models to generate visual and text features from videos
Identity-aware video model
An open-source project that aims to improve large vision-language models by integrating identity-aware capabilities and utilizing visual instruction tuning data for movie understanding
lxtgh/omg-seg1.3K
Visual Model
Develops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.
nvlabs/prismer1.3K
Vision-Language Model
A deep learning framework for training multi-modal models with vision and language capabilities.
Visual Prompt Model
A system designed to enable large multimodal models to understand arbitrary visual prompts
Visual decoder
A large language model designed to process and generate visual information
Model optimizer
This project presents an optimization technique for large-scale image models to reduce computational requirements while maintaining performance.
Document comprehension model
An implementation of a vision vocabulary model for large language models to improve document understanding and recognition capabilities
ML image recognition model
An implementation of a multimodal learning approach to improve language models' ability to recognize unseen images and understand novel concepts.
Image segmentation tool
A system that uses large language models to generate segmentation masks for images based on complex queries and world knowledge.
Multilingual Model
Develops and publishes large multilingual language models with advanced mixing-of-experts architecture.
Video understanding model
This project develops an AI model for long-term video understanding
Vision-Language Model
A multimodal AI model that enables real-world vision-language understanding applications
openbmb/viscpm1.1K
Multimodal Models
A family of large multimodal models supporting multimodal conversational capabilities and text-to-image generation in multiple languages
Visual reasoning engine
A framework for training multi-modal language models with a focus on visual inputs and providing interpretable thoughts.