LongVA
Long Context Transfer from Language to Vision
AI summary
Long context transfer
An open-source project that enables the transfer of language understanding to vision capabilities through long context processing.
- stars
- 347
- forks
- 18
- watching
- 7
Similar projects
Found by comparing what the projects do, not just their names.
LLM
An open-source implementation of a vision-language instructed large language model
LLM trainer
Transfers visual prompt generators across large language models to reduce training costs and enable customization of multimodal LLMs
Language model toolkit
A guide to using pre-trained large language models in source code analysis and generation
Model evaluation toolkit
Tools and evaluation framework for accelerating the development of large multimodal models by providing an efficient way to assess their performance
Visual decoder
A large language model designed to process and generate visual information
Vision Language Model
Develops a PyTorch implementation of an enhanced vision language model
3D LLM
Developing a Large Language Model capable of processing 3D representations as inputs
Video understanding model
This project develops an AI model for long-term video understanding
Image processor
A system for scaling large language models to process and understand visual information from multiple images efficiently.
nvlabs/prismer1.3K
Vision-Language Model
A deep learning framework for training multi-modal models with vision and language capabilities.
Image segmentation tool
A system that uses large language models to generate segmentation masks for images based on complex queries and world knowledge.
lxtgh/omg-seg1.3K
Visual Model
Develops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.
Vision-Language Learning Model
Develops and trains models for vision-language learning with decoupled language pre-training
Language model development platform
Develops and releases large language models trained on vast amounts of data for various applications, including natural language understanding, text generation, and more.
Vision-Language Model
A multimodal AI model that enables real-world vision-language understanding applications