MobileVLM
Strong and Open Vision Language Assistant for Mobile Devices
AI summary
Vision Language Model
An implementation of a vision language model designed for mobile devices, utilizing a lightweight downsample projector and pre-trained language models.
- stars
- 1.1K
- forks
- 69
- watching
- 21
Similar projects
Found by comparing what the projects do, not just their names.
Image detector assistant
An AI-powered image detection system with language-based reasoning capabilities
nvlabs/prismer1.3K
Vision-Language Model
A deep learning framework for training multi-modal models with vision and language capabilities.
Vision-Language Learning Model
Develops and trains models for vision-language learning with decoupled language pre-training
Vision-Language Model
A multimodal AI model that enables real-world vision-language understanding applications
Vision-Language Model Framework
Implementing a unified modal learning framework for generative vision-language models
Visual Knowledge Model
This project integrates visual knowledge into large language models to improve their capabilities and reduce hallucinations.
Multimodal evaluation framework
Develops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.
Multimodal LLM
A multi-modal language model that integrates image, video, audio, and text data to improve language understanding and generation
Vision language model trainer
An annotated preference dataset and training framework for improving large vision language models.
Long context transfer
An open-source project that enables the transfer of language understanding to vision capabilities through long context processing.
Visual decoder
A large language model designed to process and generate visual information
Image captioner
An end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.
Vision-Language Bridge
Pre-trains a multilingual model to bridge vision and language modalities for various downstream applications
Multimodal LLM
A multi-modal large language model that integrates natural language and visual capabilities with fine-tuning for various tasks
Visual Model
Develops a multimodal Chinese language model with visual capabilities