LaVIT

Visual understanding and generation framework

A unified framework for training large language models to understand and generate visual content

LaVIT: Empower the Large Language Model to Understand and Generate Visual Content

GitHub

544 stars
14 watching
30 forks
Language: Jupyter Notebook
last commit: almost 2 years ago

Related projects:

RepositoryDescriptionStars
lavi-lab/visual-tableA project that generates visual representations tailored for general visual reasoning, leveraging hierarchical scene descriptions and instance-level world knowledge.14
gicentre/litvisAn approach to designing and building visualizations through literate programming with Elm, Markdown, and Vega.382
luogen1996/lavinAn open-source implementation of a vision-language instructed large language model513
nvlabs/relvitA deep learning framework designed to improve visual reasoning capabilities by utilizing concepts and semantic relations.64
jshilong/gpt4roiTraining and deploying large language models on computer vision tasks using region-of-interest inputs517
pku-yuangroup/chat-univiA framework for unified visual representation in image and video understanding models, enabling efficient training of large language models on multimodal data.895
nvlabs/prismerA deep learning framework for training multi-modal models with vision and language capabilities.1,299
vpgtrans/vpgtransTransfers visual prompt generators across large language models to reduce training costs and enable customization of multimodal LLMs270
opengvlab/visionllmA large language model designed to process and generate visual information956
hxyou/idealgptA deep learning framework for iteratively decomposing vision and language reasoning via large language models.32
jiasenlu/hiecoattenvqaA framework for training Hierarchical Co-Attention models for Visual Question Answering using preprocessed data and a specific image model.349
yitzchak/ngl-cljA software framework that provides a widget model approach to create interactive visualizations in Common Lisp for Jupyter notebooks.2
lxtgh/omg-segDevelops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.1,336
vega/vegaA declarative format for creating interactive visualization designs11,276
x2fd/lvis-instruct4vA dataset of fine-grained visual instructions generated by prompting a large language model with images from another dataset131