LongVA

Long context transfer

An open-source project that enables the transfer of language understanding to vision capabilities through long context processing.

Long Context Transfer from Language to Vision

GitHub

347 stars
7 watching
18 forks
Language: Python
last commit: almost 2 years ago

Related projects:

RepositoryDescriptionStars
luogen1996/lavinAn open-source implementation of a vision-language instructed large language model513
vpgtrans/vpgtransTransfers visual prompt generators across large language models to reduce training costs and enable customization of multimodal LLMs270
vhellendoorn/code-lmsA guide to using pre-trained large language models in source code analysis and generation1,789
evolvinglmms-lab/lmms-evalTools and evaluation framework for accelerating the development of large multimodal models by providing an efficient way to assess their performance2,164
opengvlab/visionllmA large language model designed to process and generate visual information956
byungkwanlee/collavoDevelops a PyTorch implementation of an enhanced vision language model93
umass-foundation-model/3d-llmDeveloping a Large Language Model capable of processing 3D representations as inputs979
boheumd/ma-lmmThis project develops an AI model for long-term video understanding254
freedomintelligence/longllavaA system for scaling large language models to process and understand visual information from multiple images efficiently.183
nvlabs/prismerA deep learning framework for training multi-modal models with vision and language capabilities.1,299
dvlab-research/lisaA system that uses large language models to generate segmentation masks for images based on complex queries and world knowledge.1,923
lxtgh/omg-segDevelops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.1,336
yiren-jian/blitextDevelops and trains models for vision-language learning with decoupled language pre-training24
vivo-ai-lab/bluelmDevelops and releases large language models trained on vast amounts of data for various applications, including natural language understanding, text generation, and more.864
deepseek-ai/deepseek-vlA multimodal AI model that enables real-world vision-language understanding applications2,145