DeepSeek-VL

Vision-Language Model

A multimodal AI model that enables real-world vision-language understanding applications

DeepSeek-VL: Towards Real-World Vision-Language Understanding

GitHub

2k stars
20 watching
202 forks
Language: Python
last commit: over 2 years ago
foundation-modelsvision-language-modelvision-language-pretraining

Related projects:

RepositoryDescriptionStars
deepseek-ai/deepseek-llmA large language model trained on a massive dataset for various applications1,512
deepseek-ai/deepseek-moeA large language model with improved efficiency and performance compared to similar models1,024
nvlabs/prismerA deep learning framework for training multi-modal models with vision and language capabilities.1,299
darshandeshpande/jax-modelsProvides a collection of deep learning models and utilities in JAX/Flax for research purposes.151
vishaal27/sus-xThis is an open-source project that proposes a novel method to train large-scale vision-language models with minimal resources and no fine-tuning required.94
baaivision/eveA PyTorch implementation of an encoder-free vision-language model that can be fine-tuned for various tasks and modalities246
abbypa/nnproject_deepmaskA deep learning implementation of an object segmentation algorithm.187
vhellendoorn/code-lmsA guide to using pre-trained large language models in source code analysis and generation1,789
yiren-jian/blitextDevelops and trains models for vision-language learning with decoupled language pre-training24
meituan-automl/mobilevlmAn implementation of a vision language model designed for mobile devices, utilizing a lightweight downsample projector and pre-trained language models.1,076
lackel/aglaImproves large vision-language models' ability to accurately describe images by combining global and local attention mechanisms.18
deepset-ai/farmAn open-source framework for adapting representation models to various tasks and industries1,743
ailab-cvc/seedAn implementation of a multimodal language model with capabilities for comprehension and generation585
vlf-silkie/vlfeedbackAn annotated preference dataset and training framework for improving large vision language models.88
jiutian-vl/jiutian-lionThis project integrates visual knowledge into large language models to improve their capabilities and reduce hallucinations.124