LaVIT
by jy0205
LaVIT: Empower the Large Language Model to Understand and Generate Visual Content
AI summary
Visual understanding and generation framework
A unified framework for training large language models to understand and generate visual content
- stars
- 544
- forks
- 30
- watching
- 14
Similar projects
Found by comparing what the projects do, not just their names.
Visual representation generator
A project that generates visual representations tailored for general visual reasoning, leveraging hierarchical scene descriptions and instance-level world knowledge.
Visualization framework
An approach to designing and building visualizations through literate programming with Elm, Markdown, and Vega.
LLM
An open-source implementation of a vision-language instructed large language model
Visual reasoning tool
A deep learning framework designed to improve visual reasoning capabilities by utilizing concepts and semantic relations.
Region-of-Interest Training
Training and deploying large language models on computer vision tasks using region-of-interest inputs
Visual unification framework
A framework for unified visual representation in image and video understanding models, enabling efficient training of large language models on multimodal data.
nvlabs/prismer1.3K
Vision-Language Model
A deep learning framework for training multi-modal models with vision and language capabilities.
LLM trainer
Transfers visual prompt generators across large language models to reduce training costs and enable customization of multimodal LLMs
Visual decoder
A large language model designed to process and generate visual information
Vision Reasoning Framework
A deep learning framework for iteratively decomposing vision and language reasoning via large language models.
Visual QA framework
A framework for training Hierarchical Co-Attention models for Visual Question Answering using preprocessed data and a specific image model.
Visualization toolkit
A software framework that provides a widget model approach to create interactive visualizations in Common Lisp for Jupyter notebooks.
lxtgh/omg-seg1.3K
Visual Model
Develops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.
vega/vega11.3K
Visualization grammar
A declarative format for creating interactive visualization designs
Visual Instructions
A dataset of fine-grained visual instructions generated by prompting a large language model with images from another dataset