360VL

Image understanding model

A large multi-modal model developed using the Llama3 language model, designed to improve image understanding capabilities.

GitHub

32 stars
0 watching
2 forks
Language: Python
last commit: over 2 years ago

Related projects:

RepositoryDescriptionStars
dvlab-research/llama-vidAn image-based language model that uses large language models to generate visual and text features from videos748
jiyt17/ida-vlmAn open-source project that aims to improve large vision-language models by integrating identity-aware capabilities and utilizing visual instruction tuning data for movie understanding26
lxtgh/omg-segDevelops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.1,336
nvlabs/prismerA deep learning framework for training multi-modal models with vision and language capabilities.1,299
wisconsinaivision/vip-llavaA system designed to enable large multimodal models to understand arbitrary visual prompts302
opengvlab/visionllmA large language model designed to process and generate visual information956
alibaba/conv-llavaThis project presents an optimization technique for large-scale image models to reduce computational requirements while maintaining performance.106
ucas-haoranwei/varyAn implementation of a vision vocabulary model for large language models to improve document understanding and recognition capabilities1,831
isekai-portal/link-context-learningAn implementation of a multimodal learning approach to improve language models' ability to recognize unseen images and understand novel concepts.91
dvlab-research/lisaA system that uses large language models to generate segmentation masks for images based on complex queries and world knowledge.1,923
xverse-ai/xverse-moe-a36bDevelops and publishes large multilingual language models with advanced mixing-of-experts architecture.37
boheumd/ma-lmmThis project develops an AI model for long-term video understanding254
deepseek-ai/deepseek-vlA multimodal AI model that enables real-world vision-language understanding applications2,145
openbmb/viscpmA family of large multimodal models supporting multimodal conversational capabilities and text-to-image generation in multiple languages1,098
deepcs233/visual-cotA framework for training multi-modal language models with a focus on visual inputs and providing interpretable thoughts.162