Chat-UniVi

Visual unification framework

A framework for unified visual representation in image and video understanding models, enabling efficient training of large language models on multimodal data.

[CVPR 2024 HighlightšŸ”„] Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

GitHub

895 stars
7 watching
43 forks
Language: Python
last commit: almost 2 years ago
image-understandinglarge-language-modelsvideo-understandingvision-language-model

Related projects:

RepositoryDescriptionStars
pku-yuangroup/languagebindExtending pretraining models to handle multiple modalities by aligning language and video representations751
pku-yuangroup/video-benchEvaluates and benchmarks large language models' video understanding capabilities121
byungkwanlee/moaiImproves performance of vision language tasks by integrating computer vision capabilities into large language models314
jy0205/lavitA unified framework for training large language models to understand and generate visual content544
nvlabs/prismerA deep learning framework for training multi-modal models with vision and language capabilities.1,299
wisconsinaivision/vip-llavaA system designed to enable large multimodal models to understand arbitrary visual prompts302
pzzhang/vinvlA project aimed at improving visual representations in vision-language models by developing an object detection model for richer visual object and concept representations.350
zhourax/vegaDevelops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.33
pku-yuangroup/moe-llavaA large vision-language model using a mixture-of-experts architecture to improve performance on multi-modal learning tasks2,023
hxyou/idealgptA deep learning framework for iteratively decomposing vision and language reasoning via large language models.32
shizhediao/davinciImplementing a unified modal learning framework for generative vision-language models43
jiutian-vl/jiutian-lionThis project integrates visual knowledge into large language models to improve their capabilities and reduce hallucinations.124
yuliang-liu/monkeyAn end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.1,849
penghao-wu/vstarPyTorch implementation of guided visual search mechanism for multimodal LLMs541
mingyuliutw/unitAn unsupervised deep learning framework for translating images between different modalities1,994